Xscaping Limits Blog Series #5: The Optical Networks Connecting AI

Having established the need for optics inside AI data centers in blog #4 of this series, let’s now look at the different optical connections within a data center network. Modern clusters consist of thousands upon thousands of accelerators, connected across numerous racks, data halls, and — increasingly — entire campuses of interlinked data centers.
Modern AI data center fabrics can be divided into three broad categories based on the scale of connectivity — that is, the number of XPUs each fabric interconnects:
1. Scale-Up fabric (~100 kW): Interconnects hundreds of XPUs in a low-latency, cache-coherent manner over an all-to-all network. This is the preferred fabric for inference workloads.
2. Scale-Out fabric (~1 GW): Interconnects up to a million XPUs within a data center through a two- to three-tier Ethernet or InfiniBand network, and is typically deployed for training workloads.
3. Scale-Across fabric (2–5 GW): Interconnects multiple data centers over long-haul networks, typically to run a single training job across multiple buildings or to aggregate capacity across a campus.
Each of these fabrics has different reach, bandwidth, and latency requirements — and those differences translate directly into different optics requirements.

Scale-Up Fabric:
Current AI data centers deploy the scale-up fabric within a rack. Inside the rack, copper interconnects remain an effective solution for moving data between tightly coupled components: the short distances allow electrical links to deliver the bandwidth today’s systems require with acceptable latency. But the pressure to grow this domain is mounting: reasoning-heavy inference workloads — with their long context windows and multi-step token generation — demand ever-larger pools of XPUs and memory operating as a single coherent system, pushing the scale-up domain beyond the boundaries of a single rack.
Beyond the rack, however, electrical signaling becomes increasingly difficult to scale as the distance between compute resources grows. This is the central challenge in expanding the size of the scale-up domain: maintaining high bandwidth over longer reaches with electrical infrastructure demands substantially more power and additional signal conditioning.
Scale-Out Fabric:
Connecting racks together for AI training workloads requires networking switches to provide the aggregation layer — traffic directors that collect data from across the cluster and steer it over high-capacity links to its destination on the broader fabric. Today, the connections between these switches typically span up to 2–3 km and rely on 800G–1.6T single-mode optical transceivers carrying 100–200G per wavelength. Because longer links mean more (and more expensive) fiber, the economics increasingly favor packing more wavelengths onto each fiber.
Scale-Across Fabric:
With campus-scale deployments and, eventually, geographically distributed interconnected data centers, the amount of fiber required skyrockets. Every optical link needs fiber infrastructure, and every additional fiber brings more optical components and physical routing hardware, adding significant operational overhead. Inter-building fiber is also conduit-constrained and expensive to trench, so spectral efficiency directly displaces fiber count: a building pair carrying tens of petabits per second on parallel single-mode links would require an impractical number of fibers, whereas packing 100+ wavelengths onto each fiber pair collapses that count dramatically.

Instead of tacking on more fiber, a far more viable strategy is to get more out of the fiber you already have — specifically, by increasing the bandwidth each fiber can carry. Multi-wavelength photonic architectures based on dense wavelength-division multiplexing (DWDM) allow many independent data channels to travel through a single optical fiber. Each channel is assigned a different wavelength of light, so multiple lanes of communication run in parallel — multiplying bandwidth without growing the physical footprint of the optics powering the cluster.
DWDM is already widely used in telecommunications and large-scale optical networks, where operators expand the capacity of existing fiber rather than continuously deploying more cable. These coherent networks using C band tunable lasers for DWDM application. However, the hyperscalers moved away from it for scale out applications within datacenters due to laser cost and economies of scale. For AI data centers, the same DWDM principle becomes ever more important as clusters grow and communication demands scale with them. With every new generation of AI — and the corresponding jump in demand on data center infrastructure — DWDM architectures carrying many data channels per fiber become an increasingly obvious choice. DWDM is only as scalable as its light source: every channel needs a laser line, and provisioning one discrete laser per wavelength quickly becomes untenable in cost, power, and manufacturability. To meet AI’s volume demand, DWDM needs a scalable laser solution — one that can generate the full spectrum of channels efficiently, at data center scale.


