THE BREAKTHROUGH
TECHNICAL DEEP-DIVE

How we
solved prefill.

Prefill hit its compute wall. Memory optimisations won't get past it. Photonics will.

01 WHAT AI ACTUALLY DOES

Inference has two steps.

The ratio of input to output tokens is 15:1, and growing. 93% of tokens are prefill. The only way to speed up prefill is to speed up the core itself.

STEP 1 — PREFILL

Processes the prompt.
Generates the KV cache.

93%
of tokens
88%
of GPU time
66%
of energy
Compute scales with (input length)².
  • OperationMATRIX × MATRIX
  • BottleneckCOMPUTE
STEP 2 — DECODE

Generates output tokens, one at a time.

7%
of tokens
Bounded by memory bandwidth, not compute.
  • OperationVECTOR × MATRIX
  • BottleneckMEMORY BW
02 THE WORKLOAD THE MARKET IS MOVING TOWARD

Coding is eating
the AI market.

50% of all tokens now go to programming. Share increased 4.5× in six months. AI revenue is growing 10×/yr. Coding workloads are almost entirely prefill — 99% of agentic coding tokens are input tokens. The workload the market is shifting toward is exactly the one OPUs dominate.

SOURCE · OPENROUTER
Q1 '23
11% Coding
89% Chat · Search · Other
Now
50% Coding
50% Chat · Search · Other
02·B WHERE GPU TIME ACTUALLY GOES
% OF GPU TIME ON PREFILL AS INPUT:OUTPUT RATIO RISES
47% Q1 2024 · ALL USES 57% Q4 2025 · ALL USES 88% AGENTIC CODING 100% 75% 50% 25% INPUT / OUTPUT RATIO →
88%AGENTIC CODING
INPUT / OUTPUT RATIO →

The industry is solving decode.
Photonic compute is the only path through prefill.

Decode has a roadmap: MHA → GQA → MLP → LRKV → QSVD. HBM to distributed SRAM. Memory bandwidth is the bottleneck, and the industry is making rapid progress there. Prefill is compute-bound. Not solvable with memory optimizations. The only way forward is to speed up the core itself — and transistor scaling has slowed. This is what the OPU was built for.

03 WHY OPTICAL COMPUTE FAILED FOR 60 YEARS

Three locked doors.

Each one has now opened.

There was no market for 4-bit GEMM.

NOW
AI inference is almost entirely matrix multiplication at low precision. For the first time there is a massive market for exactly the operation photonics does natively.

There were no small optical transistors.

NOW
Neurophos shrank the modulators 10,000×. 1 million in 5×5 mm² on standard CMOS. Previously required 1 m².

Moore's Law was still delivering.

NOW
Transistor scaling has slowed. GPU efficiency gains are incremental. The compounding advantage of electronics is over. Photonics is just beginning.
04 THE BREAKTHROUGH
For 60 years, optical compute promised everything. And failed. Every time. For one reason:
“There are no small optical transistors.”

Neurophos
shrank them
10,000×.

10,000×
Smaller than state-of-the-art optical modulators
1,000,000
Photonic units in 5×5 mm². Was 1 m².
CMOS
Standard process. No exotic materials.
300+
Patents. Computation demonstrated.
05 HOW CAN YOU BE SO FAST?

Speed = clock
× ops per clock.

56 GHz clock — 30× NVIDIA Blackwell. 8× operations per clock as arrays reach 3k × 3k through quadratic compute scaling. In metaphor: a 56,000-round clip fired at 56 GHz, each bullet several million ops — reload at a MHz.

NVIDIA BLACKWELL · ELECTRONIC
1.84 GHz
SPARSE ELECTRICAL PULSES
T100 OPU · PHOTONIC
56 GHz
OVERWHELMING PHOTONIC STREAM
06 SCALING ROADMAP · PERFORMANCE

420× faster than
NVIDIA Blackwell.

Photonic cores run at 56 GHz — 30× faster than Blackwell. 8× more operations per clock, scaling quadratically with array size. Energy efficiency compounds as we scale — optical cores consume power only for I/O.

56GHz
CLOCK SPEED
30× NVIDIA Blackwell (1.84 GHz).
8×
OPERATIONS PER CLOCK
Scales quadratically with array size.
200×
ENERGY EFFICIENCY
TOPS per watt vs Blackwell.
07 VS. THE STATUS QUO

The entire industry
fits in the bottom corner.

Speed (TOPS/mm²) on the vertical. Efficiency (TOPS/W) on the horizontal. Everyone else clusters near the origin. Neurophos OPUs plot exponentially off-chart.

NEUROPHOS OPU
STATUS QUO
07·B SCALING ROADMAP · SPECIFICATION COMPARISON
METRIC NEUROPHOS OPU · MANWE NVIDIA B200
PEAK SPEED4.2 ExaOPS~10 PetaOPS
ENERGY EFFICIENCY1,400 TOPS/W~9 TOPS/W
COMPUTE DENSITY~2,800 TOPS/mm²~5 TOPS/mm²
CLOCK SPEED56 GHz1.84 GHz
POWER DRAW3,000 W1,000 W
PHYSICSPhotonic tensor coresElectronic transistors
SCALING LAWExponential (optical)Slowing (Moore's Law)
THE BRIDGE

The physics works.
The chip is real.
Now meet the team.

Meet the Team Contact Us