We are building the
Affine Processing Unit.
A new class of fully programmable AI processors for hyperscale workloads, designed to deliver better token economics as AI moves from occasional chatbot queries to always-on agentic systems running in the background.
GPUs are hitting diminishing returns for AI.
By 2030, larger dies, more high-bandwidth memory, and more power will stop translating into better tokens per dollar. Meanwhile, model architectures keep evolving and data centers fight for utilization. Hyperscalers need hardware that is fully programmable like a GPU, can adapt to changing transformer designs, and has a higher performance ceiling for AI workloads.
The Affine Processing Unit keeps a familiar GPU-style SIMT programming model to ease adoption, but replaces the internal hardware architecture with an AI-native design. It is not a fixed-function accelerator betting on a fixed model shape. It is a fully programmable AI processor designed to capture the next decade of token growth as GPU economics start to break.
Higher compute & SRAM density.
Much more of the die goes to useful compute and on-chip memory (SRAM) than in GPUs.
A programming model you already know.
SIMT execution and software-defined kernels keep it programmable like a GPU, so existing intuition and workloads carry over.
Built for agentic AI.
Kernels are easier to write, and CPU-driven kernel launches drop drastically. Thousands of small background tasks should not pay repeated launch overhead.
Higher real-world utilization.
Tuned for real model workloads, not just peak-FLOPS benchmarks.
Every major compute era created a new processor category. AI is no different.
The best tokenomics in AI infrastructure.