Every 100,000 GPUs you run burns ~$60M a year in power.
We hand you 10%+ of it back. Permanently.
Modelled at 0.85 kW/GPU incl. cooling · $0.09/kWh US industrial average · 90% utilization
Power is your largest operating expense — we turn it into immediate net margin. Validated across a series of bare-metal evaluations on third-party NVIDIA Blackwell and Hopper nodes — measured board-level power and thermal reductions in the 10–35% range under sustained high-utilisation load, with the live runtime policy path exercised end-to-end.

A series of controlled validations on third-party bare-metal NVIDIA Blackwell and Hopper nodes — measured reductions of 10–35% under sustained load across controlled phases, with the full remote policy path exercised live on each.
Range across controlled bare-metal validations on Blackwell and Hopper nodes, vs unconstrained baselines under sustained full-utilisation load.
Secured remote policy with on-node agent — power limits actuated across the valid hardware range, under load.
Power, utilisation, limits and thermals recorded across every controlled phase.
Results from a series of controlled validations on third-party bare-metal infrastructure. Fleet-level projections elsewhere on this page use a deliberately conservative 10% rate. Proprietary decision logic is not disclosed.
Read the technical validation noteUnlock 10%+ energy efficiency in every NVIDIA GPU cluster. Models and serving stack untouched.
That efficiency is not an abstraction — it is roughly $60 per GPU, per year, returned at any scale. Here is what it is worth to a fleet the size of yours.
Modelled at 0.85 kW per GPU (600 W board power × 1.4 PUE), $0.09/kWh (US industrial average) and 90% cluster utilization. Projections at fleet scale, from a control path measured on bare metal — not a claim that a fleet of this size has been instrumented. Adjust to your own fleet in the calculator below.
No code changes. No stack changes. Your models run untouched.
Patent-pending proprietary architecture, validated on bare-metal NVIDIA Blackwell RTX PRO 6000 hardware (controlled single-node run). H100 compatibility confirmed in engineering environments.
CO₂ reduction tracked and reported as standard, every hour of runtime.
AI is getting
expensive to run.
Large inference clusters run thousands of GPUs at near-constant draw. Every watt that isn't doing useful compute is capital — evaporating continuously, around the clock, across every node in the fleet.
The stack beneath the model — schedulers, runtimes, memory pathways — was never designed for energy-precision. It leaves margins wide, and those margins compound into the single largest line item on the infrastructure bill.
The gap between what a cluster could consume and what it does consume is where the money lives. Nobody has been able to reach it without touching performance.
PitCrew
A continuous runtime optimiser that sits alongside your inference stack and closes the energy gap — silently, permanently, and engineered to extend across fleets.
Built as a sealed, patent-pending system and protected by trade secrets — engineered for environments where security and intellectual property matter.
What it is: a lightweight on-node agent that continuously actuates GPU power limits under a live policy — served locally or from a secured remote policy service — without modifying your models, your scheduler, or your serving layer.
What it is not: a model compression tool, a quantisation pass, a scheduling change, or anything that requires modifying your existing stack. It does not alter models, weights, or application code. It does not require changes to your serving stack. Energy policy is applied only at the GPU management layer.
Energy, not output
Optimises power draw at the GPU management layer. Models, frameworks and serving code remain untouched.
Continuous Optimisation
A persistent observe → decide → actuate loop that adapts to live load in real time — no batch jobs, no redeploys, no intervention.
Black-Box Protected
Patent-pending methods and trade-secret core. Your models never leave your environment. We never see your data.
Instant Visibility
A live telemetry dashboard the moment you deploy — power, cost, and carbon savings accumulating in real time.
Your power bill is unrealised margin.
We don't just save energy — we turn your largest operating expense into immediate net margin. Model your own fleet below, at our conservative 10% saving rate.
Board power per profile: RTX PRO 6000 Blackwell 600 W · H100 SXM 700 W ·
B200 1000 W · GB300 1400 W — each scaled by a 1.4 PUE for cooling and facility
overhead, at the US industrial electricity average (~$0.09/kWh, EIA).
Energy spend is scaled by your cluster utilization (default 90% — few fleets run truly continuous). Illustrative model at a 10% saving rate — a projection for planning, not a guarantee for a specific deployment.
Watch the savings
accumulate in real time.
From the moment PitCrew engages, your engineers see money and carbon being returned to the business — continuously, while every workload runs at full performance. No degradation. No intervention. Just the line going up.
Representative interface · sample data, not a live customer deployment
Figures shown are illustrative of potential impact at scale — modelled on a 10,000-GPU cluster at a 10%+ saving — not live results from a specific customer deployment.
Built by people who
understand the data centre.
Traditional clusters are tuned for peak power draw. PitCrew-enabled clusters are tuned for performance-per-watt. That shift — from the physics of the rack to the software of the stack — is why operators see PitCrew as the natural upgrade path for their infrastructure, not another tool bolted on top.
Thermal & power density
Every watt not drawn is heat your cooling plant never has to reject. PitCrew actively shapes the rack’s thermal profile, returning headroom to power- and cooling-constrained halls.
Modern GPU racks are provisioned against peak draw, not useful work. Tuning for performance-per-watt lets the same power envelope carry more inference — without touching the utility contract.
Clusters that run hot spend their lives skirting thermal and power limits. A calmer energy profile means fewer excursions, steadier clocks, and more predictable latency at the tail.
Stack-native integration
Built for the stacks you already run — TensorRT-LLM, vLLM and custom serving layers. No model changes, no re-quantisation, no rewrites to your inference logic.
PitCrew operates as a telemetry and control-logic layer at the GPU management plane. It does not modify models, frameworks, schedulers or application code.
The same per-node control loop is designed to roll out cluster by cluster, like the infrastructure software your teams already trust — observable end-to-end, and removable without a trace if you ever choose to.
Your racks were engineered for more useful work per watt than they deliver today. PitCrew unlocks the performance density they were always designed for.
Engineered to interfere
with nothing.
GPU Management Layer Only
Observes runtime state and applies energy policy at the GPU management layer. Never modifies models, frameworks, schedulers, or application code.
Works with TensorRT-LLM / vLLM / Custom Stacks
Compatible with all major inference serving layers. No integration lock-in.
Lightweight Sidecar
Deploys as a single container per node. Designed for sub-1% overhead. No kernel patches.
Bare-Metal Blackwell Validated
Controlled validation on a third-party bare-metal NVIDIA Blackwell node — sustained full-utilisation load, live policy actuation, full telemetry logged. Not simulation.
No Model Changes Required
Your weights, your quantisation, your precision. Nothing is touched or retrained.
No Stack Modifications
No scheduler changes, no driver updates, no downtime. Typical deployment in minutes.
Written for the engineer who doesn't believe us.
A series of controlled evaluations on third-party bare-metal hosts spanning NVIDIA Blackwell (RTX PRO 6000) and Hopper (H100) architectures — no hypervisor on the evaluation path. Measured savings across the series fall in the 10–35% range. Each run: phased operation under sustained full-utilisation stress load, with thousands of telemetry samples captured throughout. Power and thermals were measured at board level; application-level QoS was not the measurement target of these runs and is scoped explicitly in the note.
| Phase | Role | Limit | Aggregate GPU Power | Mean Temp |
|---|---|---|---|---|
| A | Baseline | 600 W class | ~4.4 kW | ~66 °C |
| B | Static policy | 420 W class | ~3.0–3.2 kW | ~57 °C |
| P | PitCrew path | Live / variable | Reduced vs A | ~49 °C |
Phase detail from one representative Blackwell-class evaluation. Busy-sample figures (GPU utilisation ≥ 80%) under sustained load. Phase P mixes live limit changes and varied utilisation windows, so its aggregate is deliberately not quoted as a headline savings percentage — its result is successful remote policy → on-node actuation under load, at lower power and temperature than baseline.
Representative result: on the Blackwell run detailed above, the static power-policy phase delivered on the order of ~30%+ lower aggregate GPU power versus baseline under sustained high utilisation, with a clear reduction in operating temperature. Across the full validation series — Blackwell and Hopper — measured reductions ranged from 10% to 35%.
PitCrew Runtime Energy Control
Technical Validation Note — Blackwell- and Hopper-class bare-metal evaluations
- —Method at a high level — phase design, load profile, and telemetry cadence.
- —Measured per-GPU and aggregate power across all three controlled phases.
- —Thermal results on busy devices, and how Phase P was assessed.
- —Scope and non-goals, stated plainly — including what the run does not establish.
- —Interpretation for operators, and how a pilot extends measurement to your fleet.
Your models never leave your cluster.
We know the first question every infrastructure team asks. PitCrew is built for environments where data sovereignty is non-negotiable.
Runs inside your perimeter
The agent deploys on your hardware. Policy can be served locally or from a secured Lux Sempiterna endpoint — your workloads stay put either way.
Zero data egress
No data ever leaves your cluster. We cannot see your models, weights, prompts, or customer data — and your environments are never used to train public models.
GPU management layer only
Observes runtime state and applies energy policy at the GPU management layer. No access to models, weights, prompts, or customer data.
Sealed & patent-pending
Patent-pending technology with a trade-secret core. The policy logic remains Lux Sempiterna's. Removable in minutes, leaving your stack exactly as it was.
Five nodes. Seven days. Nothing to lose.
You should not have to take an efficiency claim on faith, and you should never have to risk a production cluster to test one. So we made the decision free.
Five nodes
We deploy on a small slice of your cluster. No changes to your models, kernels, or application code.
Seven days
PitCrew runs alongside your live workload. We measure energy on the instrumented nodes against your own baseline. Throughput and latency are reviewed against SLOs agreed before day one.
The verdict
You receive a full telemetry report. If we have not shown at least 5% lower energy on the pilot nodes over the agreed windows, you pay nothing and we uninstall. Material breach of the pre-agreed latency or throughput SLOs is treated the same way.
Aligned pricing. After a successful pilot, enterprise terms are typically 15–25% of verified energy-cost savings on covered nodes, alongside a modest platform fee — with measurement rules fixed in writing: joint baseline, same meters, same windows. If the agreed energy reduction isn't there, the commercial obligation isn't either.
The largest unlocked cost in AI
is the energy you don't need.
At hyperscale, fractions of a percent translate into hundreds of millions. Lux Sempiterna makes large-scale AI inference meaningfully more efficient — not at the model layer, not at the scheduler, but at the runtime where the waste actually lives.
Cut the line item nobody can reach
Recover meaningful energy savings — fleet projections use a conservative 10%+ rate derived from controlled bare-metal Blackwell validation — without modifying models, schedulers or your serving stack.
Efficiency at the runtime layer
AI infrastructure spend is scaling faster than compute itself. PitCrew addresses the largest unlocked cost in inference — the energy gap — with a defensible, patent-pending core.
More margin per GPU-hour
Every credit hour you sell becomes more profitable when the underlying draw drops. PitCrew lets you offer greener, cheaper compute without repricing.
A founder-led systems engineering practice.
The core team is small, high-calibre, and operates under NDA. Proprietary methods remain sealed; measurement boundaries and operational behaviour are stated plainly.
His systems work began with early personal computers and Amiga hardware, including close study of its custom graphics architecture, and continued through building machines around NVIDIA GPUs. He later completed official workflow training at RED Studios Hollywood and spent years managing NVIDIA GPU and CUDA pipelines in feature-film and high-end post-production environments — real production workloads, real deadline pressure, real thermal and performance constraints.
Hardware-native intuition
Long hands-on history with graphics architectures, from Amiga custom chips through modern NVIDIA GPUs.
Production-systems discipline
Experience running and managing GPU pipelines where stability and predictability were non-negotiable.
Non-interference design
The control plane is built to observe and actuate at the GPU management layer only. Models, schedulers and application code remain untouched.
Sealed core, transparent boundaries
Proprietary decision logic stays protected; what is measured, what is guaranteed, and what remains out of scope are stated without ambiguity.
Additional first-class technical collaborators work under NDA. The team's shared standard is simple: measure carefully, act narrowly, and never surprise the stack.
Stop paying for
wasted energy.
Schedule a cluster audit or request sandbox access. Designed for rapid deployment. Most evaluation clusters begin seeing measurable efficiency signals within days of the pilot window.
Request a Detailed Pilot
For infrastructure and platform teams ready to evaluate PitCrew on their own clusters. Tell us a little about your environment and we'll follow up with next steps.
Scaling the global
infrastructure for AI.
Lux Sempiterna is a Delaware C-Corp building the proprietary runtime architecture for the next decade of AI compute — with particular strength in large-scale LLM inference. By optimising clusters at the hardware–software interface, we enable a fundamental shift: reducing the carbon footprint of massive GPU fleets while simultaneously expanding net margins for infrastructure operators.
Our architecture is field-validated, patent-pending, and developed in consultation with leading experts in high-performance computing. We are currently scaling operations and selectively engaging with strategic investment partners who understand that the future of compute is defined by efficiency, not just raw power.
All inquiries are treated as confidential.
Margin expansion via efficiency
PitCrew improves the underlying economics of every GPU-hour. Lower energy draw per unit of inference means structurally higher net margins for infrastructure operators — a permanent architectural gain, not a temporary optimisation.
High-margin, low-carbon compute
The same architecture that expands margins compounds environmental impact at fleet scale. Every deployed cluster becomes measurably cleaner — sustainability delivered as a competitive advantage, not a cost centre.