
Photonic TPU
This is a PCIe 6.0 x16 optical AI accelerator designed for matrix operations and neural network inference. The device is built upon a silicon photonic interposer with integrated optical waveguides, hosting compute chiplets, memory, laser sources, and electronic controllers. The core computing system consists of photonic matrix processors utilizing Mach-Zehnder interferometers and microring resonators; digital data is converted into optical signals, with amplitude and phase encoding the values for matrix operations. The laser module employs Wavelength Division Multiplexing (WDM), enabling the transmission of multiple independent data streams over a single waveguide at distinct wavelengths. Non-volatile optical memory (O-PCM)—where weights are represented by levels of the material's optical properties—stores neural network parameters, while HBM4 handles dynamic data and activations. An electronic controller equipped with high-speed DACs and ADCs converts digital server data into electrical control signals for the photonic compute units and returns the results in digital form. Data transmission between boards occurs via optical ports, allowing accelerators to be linked into a distributed computing cluster. Primary server connectivity is established via PCIe or CXL, and a standard software stack supports machine learning workloads. The P-TPU delivers performance exceeding 50 PFLOPS for FP16/INT8, achieves a memory bandwidth of 50 TB/s, and operates with a power consumption of approximately 50 W. Computations are performed directly within the optical matrices as light signals pass through interferometric structures, after which photodetectors convert the optical output back into an electrical signal. This architecture integrates photonic computing, HBM, non-volatile optical memory, and high-speed optical interconnects onto a single compute board.
Model Information
Anton Dubina
In Progress
Car
In Development
TBA











