The AI Infra race is entering a new phase

Published on: 08/11/2026

By Keshav Malhotra

SHARE

For years, the focus was on training: ๐’๐’‚๐’“๐’ˆ๐’†๐’“ ๐’Ž๐’๐’…๐’†๐’๐’”, ๐’ƒ๐’Š๐’ˆ๐’ˆ๐’†๐’“ ๐’„๐’๐’–๐’”๐’•๐’†๐’“๐’”, ๐’Ž๐’๐’“๐’† ๐‘ฎ๐‘ท๐‘ผ๐’”.

But as AI moves into real products, enterprise workflows, and agentic systems, the heavier load is shifting to inference.

The challenge is no longer just GPU access. It is GPU efficiency.

In this newsletter, we look at three approaches emerging across the stack:

๐Ÿ“€ Hardware built for agentic workloads.
๐Ÿ“€ Data pipelines that keep GPUs fed.
๐Ÿ“€ Scheduling layers that improve utilisation across multi-model workflows.

Download Playbook