
For more than four decades, AMI has written the firmware that boots, manages, and secures the world's servers. That kind of longevity in a business long defined by proprietary code makes the company's latest strategic shift notable. On a recent episode of The Control Plane, Allyson Klein, CEO of TechArena, sat down with Zachary Bobroff, VP of product management at AMI, a Lattice Company, to talk about why open source has become central to AMI's firmware strategy, and how the company is untangling what he calls firmware's "five-headed hydra."
As VP of product management, Zachary spends much of his time gathering requirements from AMI's two core customer bases: system manufacturers, including OEMs and ODMs, and the hyperscalers that deploy that hardware at scale. Neoclouds and other large-scale consumers of systems are increasingly entering the conversation too, he said, reshaping how the industry defines its firmware needs.
Historically, OEMs owned platform architecture, Zachary explained. As hyperscalers rose to prominence, influence moved first to ODMs and then to the hyperscalers themselves. Now, with end customers looking to take control of firmware and vertically integrate it, open source has become the best way to reach that goal. In Zachary’s words, "Open source is the best path to innovation, ecosystem adoption, and having everybody on the same page."
AMI has increased its contributions to OpenBMC, the Linux Foundation project that originated at IBM as a lighter-weight alternative to full-featured BMCs. AMI has also built MegaRAC OneTree Community Edition, its own distribution of OpenBMC. "Our OneTree product is OpenBMC at its core, and you can kind of view it as a distribution of OpenBMC. This is a tagged, well-validated version of OpenBMC. We're testing it across multiple silicon vendors, multiple ODM platforms and OEM platforms,” Zachary said.
Customers who need more can still license AMI's proprietary version, which provides long-term direct customer support and SLAs, plus IP packages for advanced needs.
The name OneTree comes from a direct response to a problem Zachary discussed at a recent OpenBMC meetup AMI co-hosted with Meta: the proliferation of firmware forks across silicon vendors and manufacturers, a challenge he calls a five-headed hydra.
Given finite engineering resources, it’s not scalable to continue to work across these multiple trees. “What we've done is take information from those different forks and merge them back into a common tree today,” he said.
Security threads through all of it. "We've taken a strong approach to security by putting out advisories to our customers and going through the effort of having third-party audits of our code base," he said, pointing to AMI's annual OCP SAFE audits of its community edition release through the Open Compute Project.
Sovereign data requirements are adding another layer of complexity. As new regional cloud providers emerge to meet local data rules, local players are staring to emerge within these regions, and many are encountering firmware at scale for the first time.
AMI backs its open source strategy with a long-term customer relationship that spans the entire hardware life cycle. "We work pre power on, well before silicon is available, all the way through 10 years to 15 years after launch," Zachary said, positioning AMI as an economy-of-scale partner that spares customers the cost of building their own in-house firmware teams.
For developers who want in, Zachary pointed to AMI's developer portal for MegaRAC OneTree Community Edition, its community portal, upcoming OpenBMC meetups, and the Open Compute Project as entry points. "We were very open to having discussions with customers and industry experts," he added.
AMI's open source pivot is a direct response to real fragmentation risk, as multiple forks across silicon paths and rising sovereign requirements threaten to make firmware unmanageable at data center scale. By feeding its OneTree CE distribution back into OpenBMC and layering proprietary IP packages on top where customers need them, AMI is betting that a single, secure, well-audited tree beats a forest of one-off forks. With AI infrastructure evolving from servers to racks to pods to rows to rooms, that bet on convergence looks well timed.

Modern AI infrastructure runs on components pulled from dozens of vendors, and each arrives with its own silicon, its own firmware and its own idea of what security should look like. That patchwork is the problem I set out to unpack on the latest episode of “The Control Plane,” an AMI-sponsored TechArena podcast. I sat down with Alex Williams, founder of The New Stack, and Stefano Righi, security and industry advisor at AMI, a Lattice company. We discussed what’s needed to trust a chip nobody can fully see inside.
Alex told us that elements of our conversation felt familiar. He spent a decade watching cloud infrastructure mature after founding The New Stack in 2014, when early cloud teams treated workloads as interchangeable since so little real state touched the infrastructure itself. At a recent KubeCon in Atlanta, Alex got an idea of the challenges AI infrastructure presents while sitting in on a session with a senior Uber engineer who described standing up a multicloud GPU environment.
“They had to configure every GPU from every different vendor,” Alex said. “And that became a nightmare for them. And it still is.”
Alex drew a direct line to software’s own fragmentation problem three decades earlier. When the 1990s produced a dozen competing versions of Unix, the industry converged on Linux kernel. By 2014, containers ushered in the age of Kubernetes and open source tools such as Cilium, eBPF and Open Policy Agent emerged to manage workloads. All were built on the Linux kernel. No comparable tool ecosystem exists yet for AI hardware, he said. For example, Alex cited “the attestation issue,” one that becomes paramount when a chip does not get verified. The instinct is to blame a hallucination, he said, when the problem could be caught early with early detection of rogue behaviors.
Stefano said that modern AI systems assemble components from vendors that each bring their own silicon, firmware, update process and security model, and that variety is what creates the visibility gaps Alex had described. The industry’s identity conversation used to stop at the user logging into a system, Stefano said, but machines now need an identity of their own. His prescription is a genuine root of trust: hardware that carries a cryptographic identity, supports secure boot, and can produce its own measurement and attestation, all resting on a foundation that can be independently verified.
“We used to say that nobody can be the cop of himself,” he said, adding that an open, community-verified standard such as Caliptra, the Open Compute Project-backed silicon root of trust project, is the shared foundation the ecosystem needs.
Alex said that firmware is already the closest thing the industry has to a software control plane, a shift he has watched play out inside the platform engineering teams that emerged once DevOps proved necessary but not sufficient on its own. He said those teams turned their attention to the cloud and to firmware, the layer that orchestrates data movement, manages power and configures interconnects, and that firmware’s reach has extended past traditional x86 and ARM CPUs over the past two or three years. Stefano tied that shift back to AMI’s own history: four decades of boot firmware experience now feed a combined offering spanning the data plane, control plane and security plane, aimed at building the chip-to-cloud chain of trust.
Our conversation made the case for treating chip trust as a shared industry problem rather than a single company’s responsibility. The language of cloud native computing is already finding a second life in hardware: Pets vs. cattle, control planes, attestation and identity are concepts platform teams once fought to standardize, and now they have to relearn them one layer down, closer to the metal, where data, Alex said, has been promoted from second-class citizen to first-class citizen with AI infrastructure. Stefano’s argument that no single vendor can police its own hardware reframes chip trust as a problem the industry must solve through open standards, not a feature any single company can attach and sell. The vendors that get there first will be the ones already fluent in both worlds, which is exactly the bet AMI is making with decades of firmware history behind it.
To hear the full conversation, listen to the podcast episode or visit AMI.com.

Investment in AI infrastructure is now measured in the trillions of dollars, and power demand at the largest sites rivals the electricity needs of entire cities. According to Ken Sun, Corporate Vice President AMI, a Lattice company, these numbers indicate why today’s AI buildout is one of the most consequential engineering undertakings in modern history. Ken joined host Allyson Klein on The Control Plane, a TechArena podcast, to talk about why firmware, for long the quietest layer of the data center, has become central to building AI data centers.
Ken came to AMI after roles spanning service providers and cloud infrastructure, and he described the opportunity as a rare moment to help the entire ecosystem, from silicon to OEMs, ODMs, and hyperscalers, collaborate and innovate in new ways. As Corporate Vice President, his focus is go-to-market strategy, ecosystem alignment, and customer success across that effort.
The industry commonly describes the modern tech stack in five layers: energy, silicon, cloud, models, and applications. What rarely gets discussed is the layer connecting all five.
“I really see firmware, and platform software to some extent, as critical pieces to help secure, scale, manage and also enable sustainability of your entire AI infrastructure state. Firmware works with all five layers to ensure there is a coherent connection across the layers and across the ecosystem,” said Ken. That positioning also explains why AMI sits across so much of the value chain, working with power and cooling vendors, silicon providers, OEMs, ODMs, hyperscalers, neo-clouds, and open source communities alike.
Ken credited Sanjoy Maity, SVP and Business Unit Leader, at AMI, a Lattice Company, with a phrase he has embraced: the five-headed Hydra.
“You can’t tackle any one of these challenges on its own, it seems like whenever you have control over one of them, another problem arises,” he said, referring to challenges posed by fragmentation, security, power and thermal management, and fleet management. AMI’s answer leans on firmware, transparency, and open standards.
The company is among the leading contributors to the Open Compute Project and to OpenBMC, work Ken sees as double-edged: open source accelerates innovation on a common stack, but can lead to more fragmentation if left unmanaged. AMI’s response is its MegaRAC OneTree Community Edition, an effort to streamline and upstream disparate open source contributions into a more unified base, paired with the services and support customers need across design, deployment, and maintenance.
Ken’s team recently returned from Computex in Taipei, and he pointed to several trends that stood out. He was impressed at seeing the pace of innovation in action, as systems became more modular, more compact, and easier to build, operate, and deploy than just a year earlier. Rack-level boot and management is a growing challenge as operators now boot and manage hundreds of nodes per rack, and power and cooling have become dynamic rather than static, which Ken called one of the toughest parts of the equation the industry is solving for. He also highlighted predictive failure as an emerging goal: managing fleets, including aging ones, with less human intervention through self-awareness and self-healing capabilities. Modular, plug-and-play design at both the hardware and firmware level, he added, is increasingly aimed at accelerating time to market.
“Security has shifted from being a design check box to just board level mandate,” Ken said, showing up in every part of the stack, from hardware up through the models. In his view, every component now needs hardware root of trust attestation, with crypto agility increasingly required to prepare for post-quantum threats. Regulation adds another layer of complexity: the EU Cyber Resilience Act extends requirements well beyond secure coding practices into vulnerability management, firmware update processes, documentation, and incident response across a product’s lifecycle.
The customer conversation has shifted as a result. “Rather than debating which technology or which route to trust to pick, customers are asking if they have a trusted partner to work with and not just a supplier,” he said. That distinction is the role AMI wants to play across its 40 years of relationships with silicon partners, ODMs, and operators.
Ken Sun’s conversation makes a clear case that firmware has moved from background utility to strategic layer. As AI infrastructure investment climbs into the trillions and power needs rival small cities, the five-headed Hydra of fragmentation, security, power and thermal management, and fleet management can no longer be solved piecemeal. AMI’s bet is that customers will find value in their open source contribution paired with services, security IP, and lifecycle support, over point solutions. For data center leaders navigating rising regulatory demands and dynamic power and cooling requirements, the real question Ken poses is not which technology to pick, but which partner can be trusted with the full lifecycle.

MLCommons released results Tuesday for MLPerf Storage v3.0, the industry benchmark that measures how storage systems handle machine learning workloads. Version 3.0 adds two tests aimed at AI inference and opens the suite to S3 object storage for the first time, extending a benchmark that had measured only training and checkpointing.
The update pushes MLPerf Storage past training data delivery and into how storage supports AI systems already serving users. Nineteen organizations submitted 144 performance results this round, including 11 first-time entrants such as Azure, NVIDIA and Nebius. The results also reveal a wide spread in power efficiency among competing systems, evidence that AI storage architecture remains unsettled even as adoption grows.
Version 3.0 adds a KV cache test, which measures how storage systems handle the read/write operations behind LLM inference. KV caching lets a model reuse key-value vectors it already computed instead of recalculating them on every conversation turn, a common technique in transformer-based AI inference. The benchmark simulates multi-turn conversations that write a context once and read it repeatedly, with transfers ranging from 64 MiB to about 3 GiB and a median workload near 24,561 contexts totaling 13 TiB of data.
The suite also adds a vector database test, which measures performance for the indexing and query workloads behind RAG pipelines. The test uses the Milvus database with 1 million vectors at 1,536 dimensions and a DiskANN index, generating a stream of small, random read-only queries whose accuracy is checked against brute-force ground truth.
“These new additions to the benchmark suite round out the test collection, covering a larger range of AI inference workloads that drive storage needs,” said Brian Belgodere, MLPerf Storage working group co-chair.
He added: “Including tests that decompose monolithic AI systems and focus on specific storage uses and patterns, such as checkpointing, KV caching and vector databases, gives stakeholders a much clearer idea of how to engineer and provision AI systems to minimize storage performance bottlenecks.”
David Kanter, founder of MLCommons and the head of MLPerf, said the KV cache test differs from the suite's earlier training and checkpointing tests in one respect. Those benchmarks were derived from MLPerf's own industry-standard training workloads, while KV cache runs on an emulation the working group built itself, since no comparable industry-standard agentic workload existed in time for this release. He said the working group wants to align future KV cache rounds with MLPerf's own agentic inference benchmark once that work matures, replacing the purpose-built emulation with a workload derived from a documented standard.
Version 3.0 also adds support for S3 object storage as an access layer alongside the existing POSIX file system standard, letting submitters run training and checkpointing workloads against object storage instead of through a file system. About one-sixth of this round's submissions used the S3 layer.
“As the scale of AI contexts reaches into the trillions, we expect object-based storage systems to emerge as a viable, and possibly preferred, alternative to filesystem-based storage,” said Curtis Anderson, working group co-chair. “By enabling S3 support now, we are ensuring that stakeholders will have the performance information they need to make smart decisions.”
In a news briefing, Anderson put that scale in concrete terms. “If you look at the math for a KV cache environment, a billion iPhones or 5 million iPhones, every iPhone user has 1,000 contexts. Now you’re talking trillions of contexts that need to be stored. That’s an object problem, not a file system problem.”
Submitters including Nebius, NVIDIA and OpenLake used S3-compatible object storage for training and checkpointing workloads in this round.
The results also gave MLCommons its first broad look at power efficiency across submitted systems. On-premises submissions for the checkpointing write test posted a median of 14 GB/second per watt, with the top result reaching 201. The UNet3D read test showed a median of 34 GB/second per watt and a top result of 277.
“There is a wide range of power efficiencies represented in the results,” Anderson said. “It also shows that there is ample room for further improvement, and we encourage all suppliers to optimize for that metric.”
The working group also reframed how it wants benchmark customers to read the numbers. Rather than raw bandwidth, the training benchmark scores how many accelerators a system can keep above 90% utilization. Checkpointing scores duration. KV cache scores the number of conversations a system supports. Anderson said the comparisons that matter to data centers are performance per rack unit and performance per watt of provisioned power, since a working data center's space and power budgets are the constraints operators cannot expand on demand.
This round marks an expansion of what MLPerf Storage measures, from a single moment in a model's life to its full arc. Training and checkpointing belong in a model’s build phase. KV cache and vector database capture how it performs once deployed and serving real conversations and queries. That shift to covering a model's full life cycle tracks where AI investment is going as more organizations move models into production.
The economics are shifting to match. Anderson's accelerator-hours framing, storage judged by GPU time saved as well as dollars per terabyte, reflects a broader repricing of infrastructure happening across the AI stack, not just storage. Object storage's arrival alongside parallel file systems tells a similar story: As AI context volumes climb toward the trillions, storage architectures built for a smaller era are being tested by systems built for hyperscale conversational AI.
As AI infrastructure keeps changing shape, from training clusters to inference fleets to the agentic systems MLCommons is only beginning to benchmark, MLPerf Storage remains one of the few places buyers can compare vendors on equal footing.

Engineering teams building with generative AI are facing a challenge: How do you move faster while making sure customers can trust your technology? Recently, Solidigm’s Jeniece Wnorowski and I sat down with Priya Sawant, senior vice president of engineering at ASAPP, to talk through that tension. ASAPP builds the chat and voice AI agents that power contact centers for large enterprises, and Priya’s teams sit at the center of a shift redefining how engineers work and what they build.
Priya described the change as twofold: Her engineers now write code with a different set of tools, and the products they ship look different, too. Generative AI is probabilistic by nature, she said, and you have to make sure that your customers believe that your technology is doing exactly what you’re saying it’s doing. Building evaluation outputs and visibility into agent behavior became a priority for her teams. Speed matters at ASAPP, she said, but not at the cost of the hygiene that earned the company its customers in the first place.
Priya said her platform teams design tools that make engineers’ jobs easier, creating “golden pathways that everyone else can follow.”
She breaks the discipline into three phases:
A platform built without attention to developer experience creates friction, Priya said, and friction kills adoption.
Priya said that when addressing internal engineering problems, you have to identify “the common denominator of the challenges that you want to solve for” and provide flexibility on top of that. Security and reliability, for example, stay fixed, while other components are more pliable. She pointed out that bottom-up adoption, spread through engineer-champions who find real value early and convey that enthusiasm to their peers, works better than top-down mandates. ASAPP borrows a page from open-source communities here, running user groups where engineers weigh in on different tools and the build plan.
Platform teams still do what Priya calls “glue work”: stitching disparate systems into a coherent experience for internal customers, and AI tools are giving these teams the same productivity boost that product engineers already enjoy. But with this boost comes more complexity: The surface area that platform teams need to cover has grown fast, spanning new models, inference platforms, frameworks and experimentation tools that shift by the week. ASAPP’s own tooling shows that shift in miniature. Engineers used to build language-based frameworks that baked in the company’s best practices, making it easier to spin up a new service. Now the team builds a base agent framework so that anyone plugging a new agent into ASAPP’s platform starts from a shared foundation rather than building one from scratch. That same instinct shows up in the product itself, as an observability suite letting enterprise customers verify that their deployed agents hold to business policy, paired with what Priya calls “the agent flywheel”: a system that mines usage data to surface the next automation opportunity. Human-in-the-loop design rounds out the approach, keeping a person in the workflow wherever regulation or customer preference calls for one.
Priya’s conversation offers a lesson for engineering leaders navigating the AI transition. Successful teams will build evaluation, observability and developer experience into the product at the start instead of patching them on after something breaks. Underneath all of it sits a storage and compute layer that has to keep pace with what these agents demand. ASAPP’s golden-pathway approach to platform engineering, paired with attention to what internal engineers need, offers a working model for enterprises that want to move fast on AI without losing the trust that took years to build.
To hear the full conversation, listen to the podcast episode or visit asapp.com.

Global data center power demand is projected to rise 27% in 2026 to 132 GW, en route to roughly 290 GW by 2030, per recent industry forecasts.
In the latest edition of The Control Plane, guest Zane Ball, chief technology officer of the Open Compute Project Foundation (OCP), joined hosts Allyson Klein, CEO of TechArena, and Colin Brix, Vice President of Marketing at AMI, a Lattice company, to discuss the growth of open hardware and what Zane calls the upcoming “golden age of firmware.”
Zane spent almost 30 years at Intel, most recently leading its data center and AI engineering team, before retiring and joining OCP. OCP was founded 15 years ago after Facebook and other hyperscalers started designing their own servers and going directly to Taiwan’s ODM ecosystem for manufacturing. What began as a gathering of a few hundred people has grown into a community of more than 500 member companies collaborating across over 150 technical projects, grown to more than 10,000 attendees at OCP’s most recent global summit.
The conversation centered on why open hardware matters more, not less, as AI scales. Zane explained that AI’s tightly coupled nature, where chip design, model design, software, the data center, and the power grid all influence one another, creates real pressure toward proprietary, vertically integrated stacks. That approach pays off short term, he said, but grows costly once an operator needs to adapt to a different geography, climate, or energy profile.
As Zane put it, “Fragility is the reason proprietary solutions can be challenged.” Finding the right interface points is what allows competition and innovation to flourish both above and below them, he added.
That thinking carries directly into what Zane sees as a defining shift for the industry: firmware’s rising importance. He said the industry is entering “a golden age of firmware,” driven by AI data centers that bring together battery energy storage systems, low-voltage DC power, liquid cooling units, and IT racks, all of which need to read the same telemetry and act in concert.
A major benefit of interoperable systems is improved reliability. “AI systems are not super reliable. With cloud systems, you can isolate every piece of the machine from every other piece. AI systems aren't like that. If one GPU goes down over here, it affects everything else. If you have better manageability, better firmware, and better collaboration across all these systems, guess what? You can build more reliable systems,” said Zane.
According to Zane, citing studies from the Duke University Nicholas Institute, if the power grid could curtail demand for just half a percent of the year, roughly 44 hours, close to 100 gigawatts of capacity that already exists in North America would become available today. This statistic illustrates the other big opportunity for open source. If data centers could flex their consumption during rare grid stress events, the industry could unlock enormous latent capacity without new infrastructure.
Colin brought the firmware vendor’s perspective to the discussion. AMI has spent 40 years managing the disaggregated hardware that makes platforms work, and the company is stepping out of a purely behind-the-scenes role.
“We’ve been this quiet player that just makes things work,” he said, adding that AMI wants to bring the voice of the end customer upstream, so reliability and security get designed in early rather than addressed after deployment.
Asked whether operators will start demanding open source firmware in their RFPs, Zane said many already do. Most cloud companies use open BMCs rather than proprietary manageability firmware, largely because they want visibility into the code running in their own data centers.
Not every layer needs to be open, he added. A BIOS configuring registers close to the silicon can reasonably stay proprietary, while the bigger opportunity lies in the layer above it, where firmware delivers security and manageability features across the whole data center. Taking inspiration from Robert Frost’s famous line ‘good fences make good neighbours’ which his old colleague Jim Keller often quoted, Zane said the goal is to build clear interfaces that let each layer of the stack innovate independently.
Zane’s conversation makes a clear case for open innovation in firmware. The bigger opportunity extends past the data center walls: if operators, chipmakers, and firmware vendors like AMI agree on the right open interfaces, the industry can unlock existing grid capacity, improve reliability, and avoid the fragility of fully proprietary stacks.

With AI deployments fundamentally changing the data center infrastructure conversation, firmware has become the trusted source of telemetry driving automated decisions across workload placement, power balancing, cooling optimization, and predictive maintenance.
For our last 5 Fast Facts Q&A in our summer series on AI infrastructure requirements, we sat down with Colin Brix, vice president of marketing at AMI, a Lattice Company, to discuss the critical role firmware plays in managing the components of the AI stack and AMI's emphasis on an open, unified codebase. Here's what we learned.
Q1: AMI's firmware runs from the BIOS and BMC on a single server up to fleet-level management across a data center. What are operators asking that firmware to do today that they weren't a year ago?
A: A year ago, operators were primarily focused on server health, uptime, and traditional infrastructure monitoring. Today, AI deployments have fundamentally changed the conversation.
Operators are demanding much deeper visibility into the resources that directly impact token production and infrastructure efficiency. That means granular telemetry around GPU performance, accelerator utilization, power consumption, thermal behavior, interconnect performance, and rack-level power dynamics. They are also asking firmware to provide increasingly sophisticated power management capabilities, allowing them to optimize performance-per-watt without sacrificing workload throughput.
Just as importantly, operators are looking for real-time data that can feed higher-level AI factory orchestration systems. Firmware is no longer just reporting health status. It's becoming the trusted source of telemetry that drives automated decisions around workload placement, power balancing, cooling optimization, and predictive maintenance.
In short, firmware is evolving from a management layer into a data and control layer for the AI factory.
Q2: The boot-to-workload chain is invisible when it works and everything when it doesn't. Where in that chain do AI deployments most often run into trouble?
A: The biggest challenges emerge at the intersection between rapidly evolving hardware and increasingly complex software stacks.
AI infrastructure combines CPUs, GPUs, accelerators, networking fabrics, storage, power systems, and orchestration software that often come from multiple vendors. Any mismatch in firmware versions, configuration settings, security policies, device initialization, or hardware inventory can create failures that are difficult to diagnose and expensive to resolve.
What makes AI environments unique is that issues that might only affect a single server in a traditional environment can impact an entire training cluster. A node that is misconfigured, running an inconsistent firmware level, or failing attestation checks can prevent thousands of GPUs from operating at full efficiency.
This is why operators place such a premium on consistency and trust throughout the boot chain. Secure provisioning, firmware integrity validation, hardware attestation, and automated fleet-wide lifecycle management have become critical because the cost of a single misbehaving node is dramatically higher in an AI environment.
From AMI's perspective, success comes from creating a trusted and repeatable path from power-on to productive workload execution across the entire fleet.
Q3: MegaRAC OneTree pulls management of compute, power, cooling, and networking into one open codebase across different silicon. What changes for an operator when those sit in one place instead of separate tools?
A: MegaRAC OneTree's unified codebase enables operators to focus on maximizing the token output of their AI factory instead of spending time maintaining dozens of independent management stacks.
With separate management domains, every platform, subsystem, and vendor often comes with its own codebase, update cycle, security process, telemetry model, and operational workflow. That creates tremendous operational complexity and introduces risk every time an update or security patch must be deployed.
A unified codebase changes that equation. When a vulnerability is discovered or a feature enhancement is required, it can be addressed once and propagated consistently across the heterogeneous fleet. Operators gain a common operational model regardless of whether they're managing compute nodes, accelerators, power infrastructure, cooling systems, or networking equipment.
Equally important, every component speaks the same language. Telemetry becomes normalized, automation becomes simpler, and fleet-wide optimization becomes practical. By creating uniformity across a heterogeneous AI factory, OneTree removes friction from daily operations and allows teams to focus on efficiency, performance, and scale rather than infrastructure complexity.
Q4: As fleets scale into AI factories, what has to be true about telemetry and automation for one team to actually run tens of thousands of nodes?
A: At AI factory scale, telemetry must be accurate, consistent, and synchronized.
Decisions about workload placement, power allocation, thermal management, and capacity planning are only as good as the data feeding those decisions. Accurate telemetry combined with precise timestamps is essential because operators are increasingly correlating events across thousands of systems simultaneously.
Beyond accuracy, standardization and unification become absolute requirements. AI factories are inherently heterogeneous environments containing compute platforms, accelerators, networking fabrics, power infrastructure, cooling systems, and storage resources. If each subsystem exposes different management interfaces and telemetry formats, automation becomes fragile and operational overhead grows exponentially.
Successful operators require systems that expose common telemetry models, common APIs, and common lifecycle management processes. Only then can automation safely aggregate fleet-wide data, identify anomalies, trigger remediation actions, and maintain optimal performance at scale.
The reality is that nobody manually operates a 50,000-node AI factory. The telemetry and automation architecture must be designed so the fleet effectively manages itself, with humans focusing on policy and optimization rather than individual device administration.
Q5: What does an operator need from the control plane to be able to promise a predictable and trustworthy cost per token?
A: A predictable cost per token begins with complete visibility and control across the entire AI factory.
Operators must continuously optimize compute utilization, networking efficiency, power consumption, and cooling performance on a second-by-second basis. Any blind spot or inconsistency directly affects infrastructure efficiency and drives up token costs.
To achieve this, operators need a control plane that is fast, reliable, and unified. They need a single operational framework that provides trusted telemetry, consistent automation, and coordinated management across all infrastructure domains. Their engineering teams should spend their time optimizing AI production rather than reconciling conflicting data sources or maintaining multiple management stacks.
Trust is equally critical. The control plane must be built on a common hardware root of trust that validates systems from initial power-on through workload execution. Operators need assurance that every system in the fleet is running authorized firmware, trusted software, and verified hardware.
Ultimately, predictable cost per token requires three things: trusted data, automated optimization, and strong security. AMI's role is to provide the foundational management infrastructure that enables all three at AI factory scale.

As AI inference workloads shift toward agentic AI, memory capacity increasingly determines the cost to serve them.
This summer, TechArena has been asking the companies that build AI infrastructure how their requirements are changing as deployments scale. We sat down with Randy Kreiser, Field CTO at Graid Technology, which has developed a portfolio of storage performance and resilience products. Using GPU-accelerated RAID for KV cache offloading, he said, makes time to first token (TTFT) roughly three times faster than running with no cache offload at all.
We also discussed what a storage tier has to deliver to keep pace with the GPU and where storage fits into the industry's shift toward rack-scale AI systems designed as a whole. Here's what we learned.
Q1: AI infrastructure conversations used to start and end with FLOPS, but inference has moved the constraint to memory. From your vantage point, how did memory become the gating factor for inference economics, and why can't GPU memory alone solve it?
A: Training is limited by how fast you can process data; inference is limited by how much conversational state you can keep available. Every active request carries a KV cache, the model's working memory of the conversation, and cache demand grows with context length and concurrent users, not simply with model size.
That matters especially for agentic AI. A long-running multistep agent can create substantial cache demand even when GPU compute is not the limiting resource. The constraint is not the GPU; it is the memory available to hold and reuse state.
GPU memory alone cannot solve the problem economically. HBM comes attached to an accelerator, so adding memory often means buying compute you do not need. When cache capacity is exhausted, the system evicts state and recomputes it later. GPU utilization may still look healthy, but useful work declines because the system is regenerating context it already processed.
Q2: KV cache offloading essentially turns NVMe storage into an extension of the memory hierarchy. What has to be true of the storage layer for that tiering to work at inference speed rather than becoming the next bottleneck?
The cache must be faster to retrieve than it is to recompute. If it is not, offloading simply adds latency to the inference path.
A: The threshold is higher than many teams expect. In a controlled vLLM and LMCache benchmark using a 235 billion-parameter MoE model across four NVIDIA H200 GPUs, Linux MD RAID5 increased TTFT from 29.4 seconds with no offload to 36.6 seconds. Adding protection to an insufficiently fast storage path can make inference worse, not better.
For NVMe to serve as an effective inference-memory tier, it needs:
Miss any one of these, and the storage tier becomes the next bottleneck. Graid Technology's innovative volume management capabilities deliver all of these.
Q3: Traditional RAID architectures were built for a CPU-centric world and burn the very cycles and PCIe lanes AI servers can't spare. What convinced Graid Technology that RAID logic belonged on the GPU, and what does that unlock in an AI data center that a hardware RAID card or software RAID can't?
A: RAID is fundamentally parallel parity math: XOR and Galois-field operations. Traditional architectures placed that work on a RAID-controller ASIC because CPUs were poorly suited to it and there was no better parallel processor in the server.
That is no longer true. Modern AI servers already contain massively parallel GPUs, while NVMe has exposed the limits of conventional RAID designs. A hardware controller can become a bandwidth ceiling as Gen5 NVMe arrays scale up. Software RAID avoids that controller bottleneck but consumes host CPU cores and PCIe resources that AI workloads need for feeding accelerators.
GPU-accelerated RAID changes the tradeoff: SupremeRAID AE uses a small portion (approximately 4% or 6 SMs) of an installed GPU rather than a dedicated RAID card, preserves host CPU resources, and enables a direct storage-to-GPU-memory path. The result is protected, high-throughput storage designed around the AI server rather than bolted onto it.
Q4: You've positioned SupremeRAID AE around KV cache offloading, citing roughly 3x improvement in TTFT. Walk us through how GPU-accelerated storage changes the inference pipeline in practice. Where do those gains actually come from?
A: Correct, we measured 3.26x faster TTFT than no offload and 4x faster than Linux MD RAID5, reducing mean TTFT from 29.4 seconds to 9.0 seconds.
The gain is not merely faster storage. It comes from avoiding unnecessary GPU computation. TTFT is heavily influenced by prefill, the work required to process the full prior context before the model can generate its first token. In long-context agentic workflows, that prefill work is expensive and often repeated.
When the KV cache already exists, the system can retrieve it instead of recomputing it. A sufficiently fast storage read is far cheaper than rerunning attention across a large context window. That releases GPU cycles for decoding and serving additional requests. The business case is simple: Fetch must beat recompute. If it does not, the offload tier works against you. SupremeRAID is proven to be faster than recompute.
Q5: The industry is moving from components bought off spec sheets to rack-scale systems designed as a whole. As storage becomes a design partner to the GPU rather than a peripheral, how does Graid Technology's roadmap fit into that integration story, and what does storage look like in the AI data center two years out?
A: Our Agentic AI Storage Portfolio is organized by deployment scale, KV Cache Server, KV Cache Rack and KV Cache Platform, rather than by individual SKUs. That reflects a shift in AI infrastructure design: Once the rack is the unit of deployment, storage cannot be treated as a component to integrate afterward.
KV Cache Platform aligns with NVIDIA's STX reference architecture. The roadmap includes native BlueField-4 DPU execution in H2 2026 and expanded drive-count support, allowing one SupremeRAID instance to span multiple CMX chassis and present a virtualized pool to an entire rack of STX nodes.
Over the next two years, storage will increasingly be evaluated in inference outcomes rather than raw capacity: cost per million tokens served, cache-hit rate, and TTFT. Terabytes remain necessary, but they stop being the headline metric. RAID and I/O processing will follow the available parallel compute, from GPUs today to DPUs in the next phase.
.jpg)
Co-packaged optics change manufacturing and integration requirements for electro-optic polymers, said Robert Blum, senior vice president of sales and marketing at Lightwave Logic, which develops EO polymers. That means much closer integration with the switch ASIC, CPU or GPU, he said, along with hybrid bonding processes that run at higher temperatures than optical assemblies for pluggable transceivers.
This summer, TechArena has been asking the companies that build AI infrastructure how their requirements are changing as deployments scale. We sat down with Robert, who discussed where interconnect materials science needs to go next, what separates EO polymers from competing approaches, and what's standing in the way of faster industry-wide adoption of optical solutions over copper. Here's what we learned.
A: There has been tremendous progress in both the materials used for lasers and modulators. III-V materials have improved, enabling higher-power lasers and electro-absorption modulators up to 200 Gb/s per lane. Silicon photonics has improved, achieving 200 Gb/s with micro-ring modulators and highly doped p-n junctions.
But new materials are required for the next modulator generation where 400 Gb/s speeds are needed. That's where thin-film lithium niobate and electro-optic polymers come into play. We see strong momentum behind EO polymers, because they can be easily integrated into standard silicon photonics foundry processes.
A: EO polymers have really improved in performance and reliability, thanks in part to the lessons learned from the OLED industry, and are now ready for deployment. They integrate much more easily with silicon photonics than lithium niobate, so the foundries love these materials. And customers like them because they can enable much more compact modulators with lower drive voltages than lithium niobate. The real race is about ramping production capacity and getting the 400G ecosystem in place.
A: Optics tends to be more complicated than copper. High-fiber-count detachable optical connectors are one of the bottlenecks. Form factors for optical engines are not standardized, so many solutions are proprietary or require custom designs and packages.
Silicon photonics foundries have ramped capacity at astonishing rates, but lead times can still be quite long. On the other hand, there is a large payback for going to optical at these higher data rates, so many suppliers are eager to make the transition.
A: Optical assemblies that go into pluggable transceivers tend to have more forgiving process requirements. Going through standard solder reflow or wire bonding is one of the main requirements.
For CPO, you need to integrate much more tightly with the switch ASIC, CPU or GPU. So you potentially start to look at hybrid bonding, which has higher processing temperatures and also much more complex assemblies in general. All this requires much closer collaboration with system integrators.
A: There are several items that are required for true high-volume production: Silicon photonics foundries need to mature their slot waveguide processes, which means that the doping profiles and slot dimensions, for example, are within their required tolerances.
The back end of line process transfer needs to be completed and qualified. And customers need to complete the design and qualification of their own module.
Finally, the rest of the 400G per lane ecosystem needs to be ready, which typically is focused on DSPs and SerDes but may also include new optical and electrical connector assemblies and higher-power laser sources.

This summer, TechArena has been talking to the companies building the power and cooling infrastructure behind AI data centers.
We caught up with Rich Whitmore, president and CEO of Motivair by Schneider Electric, which designs and manufactures liquid cooling systems, including coolant distribution units, for high-density AI and HPC deployments.
As rack densities climb and heat removal becomes as critical to AI deployment as power delivery, liquid cooling has become a foundational part of data center design from the outset. We talked about how Motivair builds cooling into AI projects from the beginning, how far liquid cooling can scale as rack power keeps rising, and what the company is engineering next as AI infrastructure demands continue to grow. Here’s what we learned.
Q1: Cooling has moved from the back of the house to the center of the design. How does Motivair help operators build cooling in from the start rather than add it at the end?
A: Cooling is now a foundational design decision for AI infrastructure. Motivair works alongside Schneider Electric’s power and digital infrastructure portfolio to help customers design liquid cooling into the project from day one. That integrated approach supports higher-density computing, simplifies deployment and helps customers reduce time to power by avoiding costly redesigns later in the project.
Q2: As rack power keeps climbing, how far can liquid cooling scale, and what changes along the way?
A: Liquid cooling is designed to scale alongside the growing demands of AI, from today’s high-density racks to tomorrow’s AI factories. As rack power increases, the challenge shifts from simply removing heat to delivering reliable, efficient and scalable thermal management across an entire facility. Motivair by Schneider Electric is advancing this next generation of liquid cooling with solutions such as coolant distribution units designed to support multimegawatt AI deployments and beyond. The future of cooling requires deeper integration with power, controls and digital infrastructure so operators can deploy faster, optimize performance and support the next wave of AI workloads.
Q3: What separates cooling that works for a supercomputer from cooling that works across thousands of racks in production?
A: Production environments require more than excellent thermal performance. They require repeatability, uptime, ease of service and seamless integration with the rest of the infrastructure. Motivair combines decades of liquid cooling expertise with Schneider Electric’s end-to-end infrastructure capabilities to help customers deploy AI at scale while reducing operational complexity and accelerating deployment.
Q4: When an operator chooses Motivair over another cooling option, what usually tips the decision in your favor?
A: Customers are looking for proven technology backed by deep engineering expertise and the ability to deploy quickly. Motivair offers high-performance liquid cooling that is part of Schneider Electric’s broader AI infrastructure portfolio, giving customers confidence that power, cooling, controls and services are engineered to work together. That integration helps reduce project risk and improve time to deployment.
Q5: What are you engineering for next, beyond today’s deployments?
A: The industry is moving toward even higher-density AI infrastructure with greater demands for efficiency, flexibility and speed of deployment. We’re focused on advancing liquid cooling technologies, expanding manufacturing capacity, and developing integrated solutions that help customers deploy next-generation AI infrastructure with greater confidence, scalability and operational efficiency.

The cost of serving a model increasingly comes down to how fast accelerators can move data between one another as AI inference workloads shift toward multistep reasoning and mixture-of-experts models.
Faster chips are not the answer, said Aanchal Sharma, senior director of product management at Astera Labs, because they buy more compute, not less waiting. Put the fastest chip in the world behind a slow fabric, and you have built what she calls “an expensive space heater.”
We sat down with Aanchal as part of our series about how companies that build AI infrastructure are seeing requirements change as deployments scale. Astera Labs builds fabric switches and connectivity products that let AI accelerators from different vendors work together inside a rack.
She talked about what it takes to prove multi-vendor interoperability before deployment, why Astera Labs builds Scorpio and Taurus as open rather than proprietary interconnects, and where the fabric has to scale next as clusters grow toward hundreds of thousands of accelerators. Here’s what we learned.
A: At rack scale, physics doesn't care whose logo is on the silicon. A GPU from one vendor, a CPU from another, memory from a third: They all have to hold a stable low-latency link under real production traffic every hour of every day for years. That's the bar.
Most failures show up exactly where you'd expect: signal integrity breaking down across longer or noisier channels, firmware that was never tested together choking on link training, and faults that only appear once hundreds of accelerators fire the same collective operation at the same instant.
We stress-test these exact combinations in our Cloud-Scale Interop Lab before a customer ever racks a single unit so customers can deploy with confidence and focus their engineering resources on AI innovation rather than infrastructure integration challenges.
A: It involves validating the rack as a complete system, not just each component in isolation.
We re-create the customer's topology with the specific CPUs, GPUs, memory devices, fabric switches and retimers, then test across operating systems, software stacks, workloads and protocols. Working with ecosystem partners, we use co-emulation, test automation and continuous regression testing to identify compatibility issues and confirm seamless integration before deployment.
The result is a validated configuration that reduces integration risk and accelerates time to deployment. That's the entire point of an interop lab: catch every integration risk on our bench months before a customer ever has to.
A: An open fabric means an operator never has to bet the entire rack on one company's roadmap. Mix the best GPU with the best CPU with the best memory; swap a supplier when lead times blow out; carry hardware forward across generations instead of tearing it out.
Taurus supports this model across Ethernet, UALink and ESUN, while Scorpio provides an open software-defined fabric architecture that supports diverse accelerators, system topologies, and both open and platform-specific protocols.
Compared with a closed interconnect, this keeps architecture and supplier choices open as performance, availability and platform requirements change.
A: Faster chips buy more compute. They don't buy less waiting.
As inference workloads move toward multistep reasoning and mixture-of-experts architectures, more of the total job becomes accelerators talking to each other rather than accelerators doing math. Put the fastest chip in the world behind a slow, high-latency fabric and you've built an expensive space heater, with GPUs sitting idle waiting on data instead of generating tokens.
That's exactly why Scorpio builds acceleration for collective operations, like Hypercast, directly into the fabric; why Taurus keeps those links clean and low power across Ethernet, UALink and ESUN as clusters scale; and why COSMOS gives operators the visibility to find and kill communication bottlenecks in real time. Tokens per watt and tokens per dollar get decided in the fabric long before anyone reads a chip's spec sheet.
A: The fabric has to scale in three directions at once: scale up inside the rack, scale out across racks and rows, and increasingly scale across between clusters and sites.
Scale-up keeps pushing toward higher radix and lower latency inside the rack. Scale-out has to connect racks and rows without the distance tax of added latency and cost eating the gains. Scale-across is the newest of the three, holding performance together across distances that scale-up and scale-out were never built to cover.
That means more optical connectivity, higher-density switching, and fabric that carries intelligence, not just bits, accelerating collective operations instead of passively relaying them. Open standards like CXL, Ethernet, PCIe, UALink and ESUN give customers one way to build that path without betting the company, but that's only part of the story. Some of the largest deployments we support run on NVLink Fusion, and others need a fully custom link built around one customer's architecture. Products like Taurus and Scorpio are built to make all three paths work.

As new AI cloud providers appear, many compete on a single number: how many GPUs they can rent by the hour. Dan Brown, who leads hardware engineering at DigitalOcean, sees the real contest elsewhere. In a recent TechArena Data Insights episode, Solidigm’s Jeniece Wnorowski and I talked with Dan and Solidigm account executive Ty MacAdam about building and delivering holistic data center infrastructure solutions for AI workloads. Dan's argument was consistent throughout: raw GPU capacity means little if data cannot reach it quickly, and that is where DigitalOcean is placing its bet.
Dan frames the core AI infrastructure challenge around three factors: where and how data sets are captured, how quickly that data can travel from its source into an AI cluster, and how fast systems can exchange tokens between systems in the AI cluster to reduce the data down into a meaningful set that can be used for inference. When any of those steps drags, the entire pipeline slows.
“The AI solution designs are limited by what we call data velocity,” Dan said. “How quickly can you get your data in, transform it, and get it back out for next useful step in the pipeline.”
Much of the friction, he noted, comes from customers trying to run modern AI against aging storage. Many arrive with legacy arrays that have been in service for a decade or more, then discover that infrastructure cannot keep pace. Moving that data into DigitalOcean’s high-performance block and object storage, or directly onto the AI machines, is where the gains show up. Solidigm supplies the high-speed NVMes, including PCIe Gen5 drives, that sit close to the GPUs and keep the read and write cycle tight.
Dan is direct about how DigitalOcean differs from the wave of providers rushing into AI hosting. Many, he argues, do one thing only.
“They’re taking GPUs and servers, shoving them into boxes and then selling them by the hour to customers,” Dan said. “We call it a GPU landlord.” With those providers, he explained, customers must bring their own data scientists, hardware engineers, and expertise to make the hardware useful.
DigitalOcean’s answer is a full suite of compute and storage products that surround the GPU and share the same backbone network and data center, tightly coupled to improve velocity. Rather than selling raw capacity by the hour, Dan said, the company is building a unified cloud where every resource around the AI cluster exists to make that cluster faster and more cost effective to run.
A clear expression of that approach is DigitalOcean Inference Router, which reads the intent behind each request and matches it to the model that best fits the developer’s task and priorities, whether they are optimizing for quality, cost, or latency, without any application-side routing logic. The approach is grounded in years of preference-aware routing research from the team behind the open-source Arch-Router work, now part of DigitalOcean. Dan described the result as token as a service, and he said the productivity effect has been striking.
Dan said the productivity gain has been dramatic: teams can now prototype a working product for a few hundred dollars in minutes, rather than standing up entire teams of administrators and data scientists to build the back end first.
Asked what organizations still overlook, Dan pointed to density and reached for a piece of computing history. Admiral Grace Hopper famously wore an 11.8-inch-long coil of copper wire. She used this to illustrate that for every such length of copper in a system, a nanosecond of delay is introduced in that system. Across a large AI cluster, Dan noted, miles of copper and fiber add up to real delay.
DigitalOcean’s response is to build the densest possible clusters, adopt liquid cooling, and keep local Solidigm storage on each server so data moves host to host in microseconds rather than milliseconds. That translates directly into shorter compute time, faster time to first token, and lower customer costs.
Ty reinforced the point from the storage side. “Storage used to be a passive layer where everybody just thought of it as where the data would sit. That doesn’t hold up anymore,” he said. “Storage isn’t where the data sits. It’s more about how fast intelligence moves.” He added that Solidigm’s high density quad-level call (QLC) NAND puts more data closer to compute at lower power, which reshapes the total cost of ownership math at deployment scale.
DigitalOcean’s message to technology decision makers is a useful reminder to a market fixated on GPU counts. Buying accelerators and fast networking is only part of the equation. If data cannot reach those GPUs quickly, or if teams overspend on oversized models for narrow jobs, the investment underperforms. Refreshing legacy storage, right-sizing a model to a task, treating proximity between data and compute as a design decision rather than an afterthought are practical places to start to get the most out of an AI infrastructure investment. As agentic workloads raise the pressure further, the providers building for velocity, not just capacity, look best positioned for what comes next.
For more information, listen to the full podcast and visit DigitalOcean.com.

The same chip architecture that Axelera AI built for the power, cost and latency constraints of edge vision AI can now also handle multiuser generative AI workloads at the edge, using the same Metis and Europa chip family and toolchain.
We sat down with Manuel Botija of Axelera AI, which builds AI chips for edge devices. Manuel said that expanding capability matters as more AI workloads move beyond the data center and edge devices are asked to run increasingly complex models. Europa, he said, is built to run vision-language models and agentic AI directly on edge hardware.
We also talked about what digital in-memory computing changes about chip performance; how customers worldwide are using Axelera AI’s technology; and what it takes to offer a predictable cost per token as inference workloads grow.
Here’s what we learned.
A: Most of the industry has tried to scale AI from the data center out to the edge, and that direction has never really worked. Data center architectures are built for abundant power, abundant budget and abundant time. None of that is available at the edge. So the chips and systems built for one environment were getting forced into a location they were never designed for, which meant serious tradeoffs had to be made.
We believed there was a different way and started by holding those “limits” as real limits. We built an architecture that solves for small power envelopes, limited budgets and time from the beginning, because at the edge, those constraints are real.
Our first architecture, the Metis AIPU, had to hit real performance within a tight power envelope and a real price point, and that discipline is exactly what let us scale to the Europa AIPU without compromising the efficiency we perfected. We have even taken that edge architecture with both Metis and Europa and built server-class products supporting customers globally with full-length, full-height PCIe cards built around multiple AIPUs.
A: Customers are putting Metis to work solving real problems in places most people never think to look.
A large global convenience store chain uses it for real-time shelf monitoring, catching stockouts as they happen. An agritech company in New Zealand built it into automated apple sorters, and a healthcare technology company in Europe uses it inside an automated pill sorting machine, both cases where accuracy and speed directly affect people’s safety.
The utilization keeps expanding. A drone company in Croatia deploys Metis in a drone pack built for search and rescue missions. In India, a company built a cargo monitoring system for the country’s busiest port, bringing computer vision to logistics at scale.
These are just a handful of the deployments in production, with new uses for vision and localized models emerging around the world. What stands out is how different these use cases are from one another, and how each one found its way to edge AI because it was the only architecture that could meet their constraints on power, cost and latency at once. It’s a reminder of how much innovation is happening at the edge, in every corner of the world.
A: Traditional processors spend most of their energy and time moving data back and forth between memory and compute, especially for the matrix-vector multiplication that makes up roughly 80% of AI inference workloads.
Digital in-memory computing performs that math directly inside the memory cell, so the data barely has to move at all. That single change removes the biggest bottleneck in conventional architectures and turns into real gains in both speed and power. The RISC-V controlled dataflow is what makes this programmable and precise rather than a fixed-function shortcut. It gives us four independently programmable cores that can run parallel model execution with deterministic, digital accuracy, unlike analog in-memory approaches that trade precision for efficiency.
That combination is why the Metis architecture delivers up to 15 TOPS per watt and three times better performance per watt than GPU-based solutions, all while keeping accuracy indistinguishable from a full floating-point model. The result is an architecture that scales cleanly. The same principles that made the Metis AIPU efficient at 214 TOPS carry through to the Europa AIPU at 629 TOPS, which means moving the math into memory isn’t just a one-time efficiency trick. It’s the foundation of our roadmap.
A: The Europa AIPU is built on the same digital in-memory computing and RISC-V dataflow foundation as Metis, but with meaningful additions: twice the AI cores, native video decode, and 16 vector cores dedicated to pre- and post-processing, all of which make Europa perfectly suited for larger and more complex models and analytics. Those additions are what make native transformer support and multiuser generative AI workloads possible at the edge.
What stays constant for customers is the development experience. The same Voyager toolchain carries across both chips, so a team that built a pipeline on Metis isn’t starting over on Europa. That continuity is what turns an architectural leap into a growth path instead of a rebuild.
Metis remains an efficient powerhouse for computer vision use cases at the edge, including embedded systems where power and space are tightly constrained. Europa extends that foundation into a highly flexible AI accelerator built for a broader range of inferencing needs, from vision-language models and agentic AI to dozens of concurrent high-resolution video streams running natively at the edge. For customers, that means the same tooling and philosophy they already trust, now available across an even wider set of use cases.
A: Digital in-memory computing removes the data movement bottleneck that makes costs scale unpredictably as workloads grow, so the economics hold steady whether an operator is running a handful of streams or scaling to dozens of concurrent workloads.
The other separator is architectural flexibility. Operators locked into separate chips for CNNs and separate chips for transformer-based models face a cost structure that multiplies every time they add a new model type. A unified architecture that handles both on the same silicon keeps cost per token stable as workloads evolve from vision to generative AI, which is the discipline that will define the platforms operators can build a business on two years from now. Europa delivers that flexibility to customers.
Lastly, when models are run on one’s own infrastructure, the cost per token is much more manageable than running models in the clouds through providers that change cost structures at their discretion. With the Axelera Voyager toolchain, customers can use open-source models tuned to their weights and deploy them in-house. This includes Voyager Wingman, an agentic platform that can build pipelines based on natural language prompts.

Solidigm and io.net discuss neo-cloud economics, GPU pricing, decentralized AI compute, and inference cost cutting versus hyperscaler cloud infrastructure.

We sat down with David Driggers, CEO and founder of Cirrascale Cloud Services, a neocloud that runs dedicated hardware for AI workloads across NVIDIA, AMD, Qualcomm and Tenstorrent accelerators. Cirrascale’s approach is relevant now as enterprises move from proof-of-concept AI deployments toward production and need infrastructure that matches each workload to the chip built for it.
We talked about how Cirrascale matches workloads to the right accelerator as models grow, how its billing model creates cost predictability, and what separates neoclouds built for low-latency inference from those built for training alone. Here's what we learned.
A: Cirrascale gets classified as a neocloud, but we were building this hardware before the category had a name. We designed the industry’s first 8-GPU server back in 2012 working directly with NVIDIA, and we were infrastructure partners to OpenAI when they were still an 8-10 person team. That hardware DNA is why we can run NVIDIA, AMD, Qualcomm and Tenstorrent side by side, alongside our work with Google on Google Distributed Cloud, and actually know what each one is good for. No single chip serves a 1B-parameter model and a 400B-parameter model economically, so our customers get the accelerator that fits their workload instead of bending their workload to fit whatever one vendor is selling.
A: We don’t charge for data ingress or egress, which removes one of the biggest hidden cost swings customers deal with elsewhere. We also give customers the choice between token-based pricing for intermittent workloads and GPU-hour billing for anything running 24/7, so the bill follows how they use the platform. And because we track token consumption at a granular level, customers get real insight into where they’re burning input tokens inefficiently. That data helps them optimize the model, not just the invoice. Inference is forever, as we like to put it, and those costs compound. That’s why we built a cost calculator: Customers can model their own workload before they commit to anything.
A: Most enterprises are still in the POC phase with models like Gemini, GPT and Claude, trying to prove utilization before they can get more approved budget. Nobody’s handing out incremental AI spend without proof it moves the needle. That’s part of why we run Private Gemini on Google Distributed Cloud. It gets a frontier model into production inside the customer’s own security boundary instead of leaving it stalled in evaluation. On the hardware side, we’re pushing into multimodal, larger-context inference. What’s changed in the last year is that hybrid deployment, on-prem for steady workloads plus neocloud for real-time workloads across regions, is becoming the norm. I expect that to be standard by the end of this year and accelerate hard through 2027.
A: We built the load balancing in the Cirrascale Inference Platform ourselves, specifically because a model’s ideal hardware changes as it grows. An 8B model that becomes a 70B model isn’t cost-optimal on its original chip anymore, and that kind of migration is really hard to do on-prem. The single biggest cost lever is capacity utilization. The platform puts each workload on the most efficient capacity that still clears its latency target, so customers aren’t paying premium rates for headroom they don’t need.
A: Training and inference are completely different animals. In training, you can run a thousand GPUs and barely notice a network hiccup. In inference, that same hiccup is a failed request in front of a customer. The neoclouds that survive are the ones built for real-time workloads and low latency from the ground up, not the ones bolting inference onto training infrastructure. Carrier-hotel-grade connectivity is a big part of that, which is why we partnered with Telehouse to deploy the Cirrascale Inference Platform inside their data centers. It also means being able to serve regulated industries, where FedRAMP High and CMMC 2.0 are the price of entry rather than a roadmap item. We also give customers real ownership flexibility. Some customers own their GPUs outright while we run the networking and storage around them. And we’re expanding our data center footprint this year to put that capacity and failover closer to where the inference actually happens.

Nicklas Frahm of Corti on domain-specific AI, built-in compliance, and the Kubernetes Engine powering sovereign healthcare infrastructure.

This summer, TechArena has been talking to the companies building the AI data centers of tomorrow.
We caught up with Steven Carlini, chief advocate for AI and data centers at Schneider Electric, which designs power, cooling and controls for AI facilities as one integrated system. As rack densities climb and utility power becomes the tightest constraint on new AI capacity, how quickly that infrastructure can be designed and deployed is becoming one of the sharpest bottlenecks in the industry.
We talked about the changing role of power in AI projects, what’s driving how fast new AI capacity can come online today, and where the gap is widest between what AI buildouts need and what today’s grid and supply chain can deliver.
A: AI infrastructure performs best when power, cooling, controls and software are designed as one integrated technology system instead of assembled piece by piece. Schneider Electric acts as an energy technology partner by bringing together electrical distribution, liquid cooling, racks, digital management and services into a unified architecture in the form of reference designs and prefabricated modules. These help customers reduce complexity, uncertainty and risk, and they help accelerate deployment. By designing these systems together from the outset, operators can improve efficiency, simplify operations and reduce time to power for new AI capacity.
A: Power has become a strategic enabler of AI deployment. As rack densities continue to climb, we are bumping up against the laws of physics and running out of physical space for the power inputs. Delivering reliable power to each rack is as challenging as delivering liquid cooling and as important as delivering power to the facility. Every decision around electrical distribution, cooling and controls affects how quickly infrastructure can be commissioned. Designing around the rack helps operators maximize available power, support higher densities and bring AI capacity online faster.
A: Today, speed is driven by time to power. Access to utility power or constructing onsite or adjacent prime power (energy parks) dedicated to the data center, electrical infrastructure, cooling and coordinated supply chains all influence how quickly AI clusters can be deployed. Organizations that take an integrated approach to infrastructure planning and use validated, end-to-end solutions can reduce project complexity, shorten commissioning timelines, and accelerate time to revenue.
A: Demand for AI infrastructure is growing much faster than traditional power infrastructure can expand. Operators need more capacity, higher density cooling and shorter deployment timelines than existing facilities were designed to support. Closing that gap requires maximizing every available megawatt through efficient infrastructure, modernizing the grid, streamlining processes for expansion and access, building new energy parks where needed, and deploying integrated and intelligent power distribution solutions that reduce time to power.
A: The next challenge is not only securing more power; it is reducing the time it takes to bring that power into service and making every available kilowatt work harder to generate tokens and intelligence. As AI workloads continue to scale, operators need infrastructure that can adapt dynamically across power, cooling and operations. The path forward is not simply adding more capacity but creating smarter infrastructure that helps organizations accelerate time to power while meeting the growing demands of AI. That requires moving from disconnected systems to intelligent, integrated infrastructure. With EcoStruxure, Schneider Electric connects energy, power and facility systems to help operators gain real-time visibility, improve efficiency and optimize performance.
Want to learn more about Schneider Electric's AI data center deployments? Check out this article.

Per aspera ad astra: through hardship, to the stars. It’s a fitting concept for this industry, where the work that gets us somewhere extraordinary can look ordinary. Cooling systems. Fabrics. Data layers. Control planes. These critical systems often don’t get the attention they deserve given they determine how far AI can actually go. When we launched Ad Astra: AI Infrastructure Competition as an official media partner of the 2026 AI Infra Summit, we wanted to aim a light at that work.
We asked for 200 words. Make the case for why your technology belongs among this year’s most significant advances in AI infrastructure. Our judging panel composed of TechArena Advisors evaluated every entry against three criteria:
Today I’m pleased to announce the results.
Airsys took the top spot with LiquidRack, a liquid cooling system built around a straightforward observation: AI is driving rack densities that existing cooling infrastructure was never designed to handle.
Rather than routing cooling through centralized facility infrastructure, LiquidRack integrates it directly into the rack. Airsys makes the case that this changes the deployment math for operators, who can add AI capacity incrementally, expand over time, and extract more compute from the power they already have rather than overhauling an entire facility to get there.
The company backs the approach with specifics. Airsys reports that its patented spray cooling technology transfers heat up to three times more efficiently than conventional approaches, freeing more power for compute. It cites a PUE below 1.02, no water consumption, and roughly 80% less dielectric fluid than immersion cooling. The company also notes that LiquidRack has been validated in customer deployments across multiple industries and geographies, which moves the conversation past the lab.
Our TechArena Advisor Laura St. John gave the following reasoning for choosing Airsys as the winner:
“Airsys’s LiquidRack paired a clear innovation with the strongest real-world evidence in the pool, backed by validated deployments across multiple industries, not just projections. In a field tackling one of AI infrastructure's biggest bottlenecks, cooling density, LiquidRack stood out as ready to deploy today.”
Three additional companies earned recognition from our Advisors.
VAST Data argued that AI has transformed computing while much of the underlying data infrastructure still reflects an earlier era of enterprise IT, with separate storage systems, databases, event streams, and vector stores endlessly copying the same information. The VAST AI OS collapses those layers into a single shared platform. The company says it now powers more than 3 million GPUs worldwide, including AI factories where one system serves over a billion CUDA cores.
Cornelis Networks made a data movement case that emphasizes the network’s critical role in the AI era. The company points to production AI clusters running at 30 to 50 percent model FLOPs utilization, with the gap lost to congestion and synchronization stalls. Its CN6000 SuperNIC carries Omni-Path, RoCEv2, and Ultra Ethernet on a single 800 Gbps adapter. Cornelis modeling indicates AllReduce collectives complete roughly 24 percent faster than standard Ethernet, which the company translates to about 13% less training time for a 250-billion-parameter model on a simulated 10,000 GPU cluster.
Rafay addressed the operational sprawl of running an AI data center, unifying Kubernetes, virtual machines, SLURM, bare metal, GPUs, and AI services under one control plane, with Token Factory extending that model into inference by turning GPU capacity into metered, standardized services. In production environments, Rafay reports improving GPU cluster efficiency by 25% to 40% and supporting three to five times more tenants without adding headcount.
The binding constraint on AI infrastructure has moved off the chip and into everything surrounding it, and the companies competing hardest right now are the ones working that perimeter. Airsys addressed thermal and power headroom, VAST Data addressed the data layer, Cornelis Networks addressed movement between GPUs, and Rafay addressed the operational model that holds the whole thing together. Congratulations to Airsys, VAST Data, Cornelis Networks, and Rafay. We will see you in Santa Clara this September.
Read more about our winners here.
If you haven’t yet registered for the AI Infra Summit, now is the time to act! As part of our partnership, we’re pleased to offer benefits to our TechArena followers who want to join the conversation in Santa Clara this fall: