The platform underneath

The AI operating system for your GPU cluster

Every model, every app —
cheaper and governed.

NurOS is a turnkey AI operating system that lets organisations run state-of-the-art models entirely on their own premises — with the convenience of a cloud service and none of the data-exposure, cost, or dependency risks. It ships pre-installed on GPU-accelerated appliances, scales across multiple nodes to serve the largest models, and operates indefinitely without an internet connection.

Capabilities at a glance

What NurOS delivers.

Capability What it means for you
Turnkey, pre-installed Powered on and serving models out of the box — no integration project.
Active-active high availability Continuous service with no single point of failure. All nodes live, all sharing load.
Multi-node SOTA clustering Run the largest frontier-class models across multiple appliances with tensor and pipeline parallelism.
Privacy-aware intelligent routing Sensitive work stays local; only safe, suitable work can reach external services — enforced automatically.
Centralized cost & credential governance One vault, enforced budgets, full audit trail — no shadow IT, no runaway spend.
Natural-language administration Manage the platform by chatting with NurOps, an autonomous operations agent, alongside a full web console.
Air-gapped with on-demand OTA Total isolation by default; controlled updates only when you choose to connect.

Two problems every AI team hits at scale.

Runaway cost
Intelligent routing sends each call to the cheapest model that still meets your policy and latency budget — so spend tracks value, not volume. Local models handle the majority of traffic at near-zero marginal cost.
Credential sprawl
Every API key for every external provider held in a centrally managed vault — issued, scoped, rotated, and audited from one place. No keys pasted into apps, no shadow IT, no uncontrolled spending.
System architecture

Three principles. One coherent platform.

NurOS is organised as cooperating layers that sit entirely inside your security perimeter. Every request passes through an intelligent router; all external connectivity is mediated by a centralized governance layer.

USER LAYER NL Interface / Chat Web Console Programmatic clients / Coding agents INTELLIGENT WORKLOAD ROUTER privacy class · cost · latency · policy · provider selection ON-PREM GPUs Sensitive & regulated data Near-zero marginal cost · sovereign OPEN-SOURCE MODELS Internal & non-sensitive workloads Self-hosted · cost-efficient FRONTIER CLOUD Non-sensitive · burst · newest weights OpenAI · Anthropic · Google GOVERNANCE LAYER credential vault · per-user / dept / use-case limits · audit trail · compliance reporting
Local autonomy
The platform runs fully inside your perimeter and never depends on outside connectivity to serve requests. Air-gapped is the default mode, not a degraded one.
Governed boundaries
Any interaction with external AI services is brokered through a single controlled gateway that enforces privacy, budget, and audit rules — never bypassed.
Autonomous operation
Routine management is handled conversationally by the built-in NurOps agent, reducing the specialized expertise needed to run advanced AI infrastructure.
Turnkey & always-on

Active-active high availability.

NurOS arrives as a complete, pre-installed AI runtime. The operating system, model-serving engine, management services, and administrative tooling are pre-integrated and validated together — so an appliance is productive from first power-on, not after a lengthy assembly project.

The runtime is active-active: critical platform services run as redundant peers across the cluster, all live and sharing load. If any node becomes unavailable, surviving nodes continue serving without interruption. There is no passive standby to fail over to — availability is a property of the cluster, not of any one machine.

  • Pre-installed runtimeValidated software stack pre-loaded on every appliance. Productive from first power-on.
  • Active-active redundancyAll control services live simultaneously. Load is shared, not parked on standby. No single point of failure.
  • Self-healingThe platform automatically redistributes work and restores service as nodes recover — without operator intervention.
  • Hardware-acceleratedTuned to extract full performance from multi-GPU appliances out of the box.
Multi-node clustering

Run the largest frontier-class models.

Frontier-class models can exceed the capacity of any single machine. NurOS clusters multiple appliances so that one model — and even a single request — can be served across several machines working in concert.

Tensor parallelism
Splits the heavy mathematical work of a model across the accelerators within an appliance for maximum throughput on tightly coupled hardware.
Pipeline parallelism
Distributes a model in stages across multiple appliances, allowing very large models to run that would not fit on a single machine. Scales as you add nodes.

Built-in efficiency techniques.

Used alongside distribution, these techniques ensure each appliance delivers maximum value — so organisations can host larger models and serve more users on the same hardware.

TurboQuant™
Advanced quantization that compresses models to low-bit precision with minimal accuracy loss — without lengthy calibration. Models occupy far less accelerator memory, so larger SOTA models fit on a given appliance and more requests can run concurrently.
Multi-Token Prediction (MTP)
Generates several tokens per step instead of one at a time, substantially raising throughput and lowering response latency for models that support it. Same hardware, materially higher output.
Intelligent routing

Route every request to the cheapest compliant model.

Every request entering the platform is evaluated by a workload-aware router that decides where it should be served — weighing task complexity, latency demand, and privacy class.

Routing by privacy class

Privacy class Typical content Where it runs
Sensitive / regulated Personal data, confidential records, secrets, proprietary source code Always local — never leaves the perimeter
Internal Proprietary but non-personal material Local-preferred
Non-sensitive General knowledge, public information Eligible for external SOTA when complexity or latency warrants

Cost- and latency-aware provider selection.

When a request is eligible to use an external service, the router does not pick a provider arbitrarily. It continuously tracks the price and observed response time of every connected external provider and automatically dispatches each request to the option that best satisfies the active policy.

External provider Current cost Observed latency Router decision
Provider A Higher Higher Not selected
Provider B Lower Lower Selected automatically
If Provider B later becomes more expensive or slower than Provider A, the router reverses the decision on its own — always steering traffic to the provider that best meets the configured cost and latency goals. No operator intervention required.
Native coding-agent integration
Development tools and AI coding assistants connect through standard interfaces and benefit from the same routing, privacy, and governance policies as every other workload — so engineering teams get frontier coding assistance without sending proprietary source code outside the organisation unless policy explicitly allows it.
Centralized governance

One vault. Enforced budgets. Full audit trail.

When the platform uses external AI services, every credential and every dollar is centrally controlled. API keys for external providers are held in a centrally managed vault rather than scattered across individual users, applications, or scripts.

This eliminates the credential sprawl and "shadow IT" that lead to uncontrolled spending and security exposure — turning external AI consumption from an unpredictable expense into a governed, fully accountable line item.

Governance control What it does
Per-user limits Caps external spend for each individual.
Per-department limits Enforces budgets per team or cost centre.
Per-use-case limits Bounds spend for specific applications or workflows.
Audit trail Records every external request for compliance and review.
Consumption reporting Dashboards and reports show usage and cost in real time.
NurOps™ autonomous administration

Manage your cluster in plain English.

NurOS includes NurOps™, a built-in autonomous operations agent that understands the platform's complete operational state. Administrators manage the cluster simply by describing what they want — through a chat interface or the web console.

  • ProvisioningBring new models or services online and add appliances to the cluster — described in plain language, executed by NurOps.
  • Quota adjustmentChange user, department, or use-case limits on request. NurOps applies the change and confirms completion.
  • Status inspectionAsk questions about health, utilisation, and capacity at any time. NurOps answers from live platform telemetry.
  • DiagnosticsNurOps investigates issues, explains root causes, and recommends or applies fixes — reducing time-to-resolution and specialist dependency.
Why NurOps lowers the bar
Because the agent already knows the system's operational details, it lowers the specialized skill required to operate advanced AI infrastructure and shortens the time from question to answer — while the web console remains available for teams who prefer point-and-click control or need detailed visual dashboards. A lean team can do the work of a large one.
Example conversation with NurOps: "We're onboarding the legal team next Monday. Add 50 users, set their monthly budget at $200 each, and bring up a private Llama-3 instance for their contract review workflow." → NurOps confirms the plan, executes all three tasks, and reports back when complete.
Air-gapped operation

Total isolation. Controlled updates.

NurOS is designed to run completely isolated from the public internet. In air-gapped mode the platform serves models, authenticates users, and is administered entirely within your perimeter — no outbound connectivity required for normal operation.

Isolation is the default state, not a degraded one — the platform behaves identically whether or not a network path to the outside world exists.

  • Isolation by defaultFull functionality with no internet dependency. Every feature available in a fully isolated environment.
  • Administrator-controlled OTA updatesAn administrator opens a controlled update window when convenient. Connectivity occurs only when, and as long as, an operator chooses.
  • Verified packagesEvery update is integrity-checked before it is applied. No unsigned or unverified packages accepted.
  • Offline update pathSustainable maintenance for fully air-gapped sites — update packages staged and applied entirely offline, preserving complete isolation end to end.
Security & privacy

Policy-first, by design.

Identity-aware routing
Sensitive data never reaches the wrong model. Every call is evaluated against the requester's identity and the content's privacy class.
Leak & injection detection
Continuous monitoring on every call for data leakage and prompt-injection attempts — not just at the perimeter.
Shadow-AI discovery
Surfaces unsanctioned agents and AI tools your team didn't approve. Visibility before exposure.
Air-gapped breaker
Last-line isolation when something slips. Hard cutoff from external networks, fully audited.
Representative use cases

Where NurOS delivers the most value.

The capabilities above combine to address a range of demanding real-world scenarios.

Regulated & privacy-sensitive workloads
Healthcare, financial services, legal, and public-sector organisations need frontier AI but cannot expose personal or regulated data to outside services. Sensitive requests are recognized and kept on locally hosted models — protected data never leaves the perimeter, while the centralized audit trail provides the evidence compliance teams require.
Air-gapped & high-security environments
Defence, intelligence, critical-infrastructure, and industrial sites often operate with no connection to the public internet. NurOS runs fully air-gapped as its normal mode — serving models, authenticating users, and being administered entirely on-site — with updates only through controlled, administrator-initiated windows or an entirely offline update path.
Software engineering with proprietary code
Engineering teams want frontier coding assistance without sending proprietary source code to third parties. Coding agents and development tools connect natively and are subject to the same privacy-aware routing as every other workload — routine and sensitive coding tasks are served locally, giving teams modern AI assistance without leaking intellectual property.
Cost-governed hybrid AI consumption
Organisations that have experienced unpredictable external-AI bills can make local models the default destination for the majority of traffic, reserving external services for the minority of requests that truly warrant them. Per-user, per-department, and per-use-case spending limits convert a volatile expense into a predictable, fully accountable one.
Hosting the largest frontier-class models
Some models are too large for any single machine. By combining tensor parallelism within an appliance and pipeline parallelism across appliances, NurOS lets a single workload span the cluster — so organisations can self-host the largest open SOTA models. As model sizes and demand grow, additional appliances extend both capacity and scale.
Mixed & evolving hardware fleets
Hardware is rarely uniform, and it changes over time. NurOS allows appliances of different types and accelerator generations to operate together as one cluster, so organisations can adopt new hardware incrementally and protect prior investments without disruptive forklift upgrades. Workloads are placed automatically on the most suitable hardware.
Distributed & multi-site deployments
Enterprises with multiple facilities, branches, or edge locations can run a self-contained NurOS cluster at each site. Every site keeps operating through intermittent or absent connectivity, while administrators retain consistent, centralised governance of identity, policy, and spend across the fleet when connectivity is available.
Self-service operations for lean teams
Advanced AI infrastructure traditionally demands scarce specialist skills. Because NurOps™ understands the platform's full operational state and responds to plain-English instructions, smaller teams can provision models, adjust quotas, inspect status, and diagnose issues conversationally — lowering the expertise barrier and freeing specialists for higher-value work.

Live in days. Grow on your terms.

Cloud
D1
Spin up
Policy & cost caps applied
D3
First agent live
Your team using NurOS
Elastic & frontier
Scale without limits
On-prem
W1
Appliance ships
Pre-configured for your stack
W2
First agent live
100% within your perimeter
Sovereign & controlled
Zero bytes leave your network
Hybrid
D1
Cloud live instantly
Agents running day one
W2
On-prem joins
Sensitive work moves on-prem
Best of both
Policy routes every workload
MAX
performance per
$1 spent
3 days
from sign-up to
first agent live
0 byte
leaves your perimeter
once on-prem is live
1
control plane
across all deployments
NurOS™ delivers the capability of frontier AI with the control of on-premises infrastructure. Turnkey and highly available. Scales across appliances to host the largest models. Routes work intelligently to keep sensitive data in-house. Governed centrally. Operated conversationally. Fully air-gapped.
NurOS™ Technical White Paper · Nurol, Inc. · 2026
Power on your AI

Cloud, on-prem, or hybrid — one platform.

NurOS runs state-of-the-art AI on infrastructure you own, govern, and trust. Talk to us about your deployment.