06.10.2026

Announcing our $52 million funding

Turning Existing AI Infrastructure Into More Compute

Today we're announcing $52M in funding, led by Creandum and Cusp Capital, to double the world's compute without building a single new data center.

Capability is more than capacity

The AI industry measures infrastructure by what it contains: GPUs, racks, megawatts. We measure it by what it delivers.

A cluster is not valuable because it is large. It is valuable because of the models it can serve, the users it can support and the training it can complete, at the performance, reliability and cost the business needs. That is capability. Capacity is only the raw material.

Turba Labs exists to close the gap between the two. Our vision is to double useful compute without building a single new data center. Turba Labs is an AI performance platform for teams who operate AI infrastructure.

What we see

Running AI infrastructure is a complexity problem, and the industry has answered it by simplifying. Models, serving engines, schedulers, accelerators, memory and networks are built and tuned as separate layers. Each layer can be well engineered while the whole system stays inefficient.

The waste is physical, not theoretical. Accelerators wait on memory and communication. Capacity sits stranded in isolated pools while jobs queue elsewhere. A larger batch raises total throughput while every user waits longer. High utilization does not mean high useful output.

Predicting what AI infrastructure will deliver, or how much more the hardware already installed can produce, is getting harder. Different chips and generations now run side by side, serving more models and workloads than ever, and every one changes how the others perform.

For years, the answer was more hardware. That worked while the capacity was available. Today, GPUs, power, capital and data-center space are scarce, and more hardware simply reproduces the same bottlenecks at greater expense.

Open-weight models raise the stakes. More models, chip architectures and deployment options create freedom, and far more combinations to get right. Every mix of resources, models and user behavior is a different system with different limits.

The root cause is not a lack of capacity. It is that AI infrastructure is never understood and optimized as one system.

What we believe

  1. Useful work is the only metric that matters. For inference, that is output that meets its service objectives. For training, it is progress toward an agreed target. Tokens per second and utilization mean little on their own.
  2. The workload decides which hardware matters. A model name and parameter count don't define how it runs. Precision, context length, batch composition and concurrency move the bottleneck. Weights fitting on a device does not mean the service fits at its intended load.
  3. The system is the unit of optimization. Performance is set by the critical path across compute, memory and network. A training step with 60 ms of compute and 40 ms of exposed communication gets only 1.43× faster when compute speed doubles. Local gains count only when they move the whole.
  4. Predict before you act. Every change to models, placement or capacity has consequences, including the cost of the transition itself. They should be known before the change is made, not discovered in production.
  5. Complexity should be made actionable, not hidden. Abstractions that hide the dependencies that decide performance create the inefficiency they promise to remove. Operators need clear choices and the reasons behind them.

What we build

Turba Labs treats workloads and infrastructure as one execution system. We connect analytics, a calibrated digital twin and orchestration in a continuous feedback loop.

Turba Labs predicts what AI infrastructure will deliver and makes sure it does, workload by workload, in real time. Operators get more output from the hardware they already have, and control behind every decision. The same understanding serves every stage of the lifecycle, from planning and deployment to continuous operation.

  • Before a decision: predictive scenario planning shows what a new model, more users or different hardware will do to performance, reliability and cost.
  • In live operation: the platform steers routing, admission, replicas, placement, batching and parallelism as resources, models and demand shift. It stays within memory limits, service requirements and power budgets.
  • For every decision: an explanation of the binding constraint, the expected improvement and the observation that would show the forecast was wrong.

Because the system is optimized as a whole, gains stop coming at each other's expense. In a modeled mixed-inference example across five workload profiles and six optimization measures, blended cost per token fell by a factor of 9.4, about 89%, with service objectives maintained. That is a result for the stated example, not a universal guarantee. It shows the size of what is being left on the table.

Our commitments

Our strength has to come from the quality of our models, the usefulness of our decisions and outcomes customers can reproduce. This is how we expect to be judged.

  • We own the application outcome. We optimize for accepted output and real training progress. We make service requirements, workloads and constraints explicit, and report latency, throughput, cost and energy separately.
  • We earn trust through evidence. We compare against a well-tuned baseline on representative workloads, including bursts. We always separate modeled results from measured deployment outcomes.
  • We make heterogeneity usable. We build one decision layer across hardware and software environments, without pretending devices are interchangeable. Each platform keeps its strengths and is judged by the same application objective.
  • We work with the ecosystem, not against it. Compilers, serving engines, cluster managers and observability tools each do essential work. We connect decisions across their boundaries so that a gain in one layer becomes a gain for the application.
  • We improve continuously and operate responsibly. Our feedback loop adapts to change while respecting hard limits and the cost of disruption. A forecast is only valuable if it can be tested, explained and corrected.

Our call

The next phase of AI will not be won by whoever builds the most capacity. It will be won by whoever gets the most out of it.

We ask every operator, from neoclouds to enterprises running their own AI infrastructure, to change the question. Not "How many GPUs do I have?" but "What can my infrastructure actually deliver?"

It took us more than two years of research and engineering to build a system that answers that question and keeps answering it as everything changes. We did it because we believe every unit of compute should deliver more useful work across the entire lifecycle of AI. Capability is more than capacity.

Dr. Patrick Jahnke and Dr. Hans-Juergen Schmidtke
Founders, Turba Labs | September 2026