The AI operating system for your GPU cluster
Every model, every app —
cheaper and governed.
NurOS is a turnkey AI operating system that lets organisations run state-of-the-art models entirely on their own premises — with the convenience of a cloud service and none of the data-exposure, cost, or dependency risks. It ships pre-installed on GPU-accelerated appliances, scales across multiple nodes to serve the largest models, and operates indefinitely without an internet connection.
What NurOS delivers.
| Capability | What it means for you |
|---|---|
| Turnkey, pre-installed | Powered on and serving models out of the box — no integration project. |
| Active-active high availability | Continuous service with no single point of failure. All nodes live, all sharing load. |
| Multi-node SOTA clustering | Run the largest frontier-class models across multiple appliances with tensor and pipeline parallelism. |
| Privacy-aware intelligent routing | Sensitive work stays local; only safe, suitable work can reach external services — enforced automatically. |
| Centralized cost & credential governance | One vault, enforced budgets, full audit trail — no shadow IT, no runaway spend. |
| Natural-language administration | Manage the platform by chatting with NurOps, an autonomous operations agent, alongside a full web console. |
| Air-gapped with on-demand OTA | Total isolation by default; controlled updates only when you choose to connect. |
Two problems every AI team hits at scale.
Three principles. One coherent platform.
NurOS is organised as cooperating layers that sit entirely inside your security perimeter. Every request passes through an intelligent router; all external connectivity is mediated by a centralized governance layer.
Active-active high availability.
NurOS arrives as a complete, pre-installed AI runtime. The operating system, model-serving engine, management services, and administrative tooling are pre-integrated and validated together — so an appliance is productive from first power-on, not after a lengthy assembly project.
The runtime is active-active: critical platform services run as redundant peers across the cluster, all live and sharing load. If any node becomes unavailable, surviving nodes continue serving without interruption. There is no passive standby to fail over to — availability is a property of the cluster, not of any one machine.
-
Pre-installed runtimeValidated software stack pre-loaded on every appliance. Productive from first power-on.
-
Active-active redundancyAll control services live simultaneously. Load is shared, not parked on standby. No single point of failure.
-
Self-healingThe platform automatically redistributes work and restores service as nodes recover — without operator intervention.
-
Hardware-acceleratedTuned to extract full performance from multi-GPU appliances out of the box.
Run the largest frontier-class models.
Frontier-class models can exceed the capacity of any single machine. NurOS clusters multiple appliances so that one model — and even a single request — can be served across several machines working in concert.
Built-in efficiency techniques.
Used alongside distribution, these techniques ensure each appliance delivers maximum value — so organisations can host larger models and serve more users on the same hardware.
Route every request to the cheapest compliant model.
Every request entering the platform is evaluated by a workload-aware router that decides where it should be served — weighing task complexity, latency demand, and privacy class.
Routing by privacy class
| Privacy class | Typical content | Where it runs |
|---|---|---|
| Sensitive / regulated | Personal data, confidential records, secrets, proprietary source code | Always local — never leaves the perimeter |
| Internal | Proprietary but non-personal material | Local-preferred |
| Non-sensitive | General knowledge, public information | Eligible for external SOTA when complexity or latency warrants |
Cost- and latency-aware provider selection.
When a request is eligible to use an external service, the router does not pick a provider arbitrarily. It continuously tracks the price and observed response time of every connected external provider and automatically dispatches each request to the option that best satisfies the active policy.
| External provider | Current cost | Observed latency | Router decision |
|---|---|---|---|
| Provider A | Higher | Higher | Not selected |
| Provider B | Lower | Lower | Selected automatically |
One vault. Enforced budgets. Full audit trail.
When the platform uses external AI services, every credential and every dollar is centrally controlled. API keys for external providers are held in a centrally managed vault rather than scattered across individual users, applications, or scripts.
This eliminates the credential sprawl and "shadow IT" that lead to uncontrolled spending and security exposure — turning external AI consumption from an unpredictable expense into a governed, fully accountable line item.
| Governance control | What it does |
|---|---|
| Per-user limits | Caps external spend for each individual. |
| Per-department limits | Enforces budgets per team or cost centre. |
| Per-use-case limits | Bounds spend for specific applications or workflows. |
| Audit trail | Records every external request for compliance and review. |
| Consumption reporting | Dashboards and reports show usage and cost in real time. |
Manage your cluster in plain English.
NurOS includes NurOps™, a built-in autonomous operations agent that understands the platform's complete operational state. Administrators manage the cluster simply by describing what they want — through a chat interface or the web console.
-
ProvisioningBring new models or services online and add appliances to the cluster — described in plain language, executed by NurOps.
-
Quota adjustmentChange user, department, or use-case limits on request. NurOps applies the change and confirms completion.
-
Status inspectionAsk questions about health, utilisation, and capacity at any time. NurOps answers from live platform telemetry.
-
DiagnosticsNurOps investigates issues, explains root causes, and recommends or applies fixes — reducing time-to-resolution and specialist dependency.
Total isolation. Controlled updates.
NurOS is designed to run completely isolated from the public internet. In air-gapped mode the platform serves models, authenticates users, and is administered entirely within your perimeter — no outbound connectivity required for normal operation.
Isolation is the default state, not a degraded one — the platform behaves identically whether or not a network path to the outside world exists.
-
Isolation by defaultFull functionality with no internet dependency. Every feature available in a fully isolated environment.
-
Administrator-controlled OTA updatesAn administrator opens a controlled update window when convenient. Connectivity occurs only when, and as long as, an operator chooses.
-
Verified packagesEvery update is integrity-checked before it is applied. No unsigned or unverified packages accepted.
-
Offline update pathSustainable maintenance for fully air-gapped sites — update packages staged and applied entirely offline, preserving complete isolation end to end.
Policy-first, by design.
Where NurOS delivers the most value.
The capabilities above combine to address a range of demanding real-world scenarios.
Live in days. Grow on your terms.
$1 spent
first agent live
once on-prem is live
across all deployments
Read the technical detail.
NurOS™ delivers the capability of frontier AI with the control of on-premises infrastructure. Turnkey and highly available. Scales across appliances to host the largest models. Routes work intelligently to keep sensitive data in-house. Governed centrally. Operated conversationally. Fully air-gapped.
Cloud, on-prem, or hybrid — one platform.
NurOS runs state-of-the-art AI on infrastructure you own, govern, and trust. Talk to us about your deployment.