10 September 2026

Scaling Compute: from new chips to working systems

A new piece of AI hardware can look shiny and exciting on a bench, but once you put it into a system, sometimes the picture changes.

There is memory to feed it, networking to connect it, software to orchestrate it, power to supply it and heat to remove. An improvement at the level of one component may disappear, or become much more valuable, when you assess it across the whole rack.

That gap between a promising technology and a working system has become one of the defining questions for our Scaling Compute programme.

Sarah Suraj Daniel

ARIA Board Member, Sarah Hunter, sat down with Suraj Bramhavar, Programme Director for Scaling Compute, and Danyal Akarca, co-founder of Callosum, to talk about ARIA's Scaling Compute programme, how the AI hardware field has changed in recent years, and what excites them about the future.

Watch now

Finding the white space

Scaling Compute began with a broad question: could radically different hardware approaches make large-scale AI computation significantly cheaper?

As Programme Director Suraj Bramhavar explored the landscape, he found an underdeveloped space at the intersection of computing and physics, with unconventional approaches that existing funding and commercial models struggled to support. So when the Scaling Compute programme launched in 2024, we set an intentionally ambitious target: could new hardware approaches reduce the cost of running large-scale AI by more than 1,000 times?

The programme initially funded 12 R&D projects spanning new computational primitives, advanced networking and interconnects, and software for modelling the performance, power and cost of technologies at system scale.

Some projects question the computational primitives themselves. Natural systems process information efficiently, often making use of noise, probability and physical dynamics that conventional digital computers work hard to suppress. Could some of those properties become useful computational resources?

Others are focussed on how data moves, how chips communicate and how future systems could be modelled before committing the time and cost required to manufacture them.

Aiming at a moving target

Over the course of the programme, investment in AI infrastructure grew rapidly and the technology itself evolved. 

By early 2025, we described working with “moving targets”. As models become more interactive and agentic, we introduced continuous benchmarking to compare emerging technologies against a fast-moving commercial baseline.

That evolution led to a new question – even if a new piece of hardware looks promising on the bench, how do you find out whether it still works once it becomes part of a complete system? This is the gap that the Scaling Inference Lab is designed to explore.

SIL Build

Backed by £50m from ARIA and delivered by CommonAI CIC, the Lab is building an open testbed where new AI technologies can be inserted into rack-level systems, run against real workloads and compared with existing commercial systems. Rather than waiting years between demonstrations, the model is deliberately iterative: a new experimental rack every six months, each acting as a structured pilot around an unproven technology.

As Sir Andy Hopper, Chairman of CommonAI CIC, describes the aim of the Lab as an environment where “new AI infrastructure can be tested and proven at system scale.”

Building the experiment at rack scale

The first Lab clusters make this approach concrete:

  • Cluster 1 is testing whether pipeline parallelism can extract useful inference performance from lower-cost commodity and older GPU hardware.
  • Cluster 2 explores whether different kinds of accelerators can be combined and intelligently assigned different work.
  • Cluster 3 focuses on optical circuit switching, examining whether new networking approaches can remove bottlenecks in increasingly complex AI systems.

The common idea is that future AI infrastructure may not be one homogeneous stack containing ever larger numbers of the same accelerator – it could be far more heterogeneous: different compute architectures, memory technologies, networks and software assembled around particular workloads.

From experiment to capability

This issue is also becoming a wider UK infrastructure question.

In June 2026, the UK AI Hardware Plan committed at least £20m in additional funding to expand the Scaling Inference Lab, building on ARIA's £50m commitment, with the aim to strengthen the UK's ability to test new compute systems on real workloads and reduce technical risk before wider deployment. The Lab has also partnered with SovereignAI’s £100 million R&D Procurement Scheme to support technologies that improve the efficiency of AI computing infrastructure.

This is important – because the difficult space between research and deployment is often where promising technologies stall. The Lab is designed to close this gap by testing these approaches in real systems, and generating the evidence needed to progress them.

To build on this momentum, we are recruiting for a Scaling Inference Lab Director with deep expertise in AI infrastructure and large-scale compute to drive the Lab’s path to scale and sustainability.

Find out more and apply