PS — Cortex-A9 @ 666 MHz (single core)
PL — Your NPU (~7,500 LUTs / 44 DSPs / 31 BRAM)

llama2.c — golden model

Tokenizer, sampling, and the C reference every RTL block is validated against.

See Phase 1

Descriptor builder

Builds a chain of {opcode, src, dst, M, N, K, M0, shift} for a whole layer, then rings the doorbell.

See Phase 6

⚠️ Cache coherency

A9 writes weights → they sit in L1/L2. PL reads stale DDR.

Xil_DCacheFlushRange() before kick Xil_DCacheInvalidateRange() after

==Looks exactly like an RTL bug. Is not one.==

UART @ 115200 → your generated stories

Command sequencer

Fetches descriptors, runs a full layer with zero PS involvement. Interrupts on done.

A baby PM4 ring. ~1,000 LUTs

AXI-Lite slave

16 control/status registers. Doorbell, kick, busy-cycle counters.

~300 LUTs

See Phase 0

8×4 systolic array

32 PEs = 32 DSPs int8 × int8 → int32, weight-stationary, output-stationary accumulation.

==6.4 GOPS @ 100 MHz== ≈ 5× the A9

See Phase 3

Vector unit

RMSNorm · Softmax · SiLU · RoPE ~12 DSPs + 4 BRAM of LUTs

Amdahl's ambush lives here — 10% becomes 35% once matmul is fast.

See Phase 5

Requantize

int32 → ×M0 → >>n → saturate int8

⚠️ Round half-away-from-zero, and check negatives. Arithmetic >> rounds toward −∞.

See Phase 2

Double-buffered tile RAM

Ping-pong weight + activation buffers. Load tile N+1 while computing tile N.

Target >80% PE utilization

12 BRAM · ~800 LUTs

KV cache

128-token context = 41 KB = 9 BRAM

256 tokens leaves only 3 blocks free → too tight. This is why 128.

See Phase 7 §7.1

Custom AXI4 burst master

~700 LUTs — you write this yourself. Xilinx AXI DMA IP is 2–3K LUTs = 20% of your whole fabric.

⚠️ Bursts cannot cross 4 KB boundaries.

See Phase 4

DDR3 — the real bottleneck

512 MB, 16-bit @ 533 MHz Peak ~2.1 GB/s · realistic ~1.2 GB/s


$$\text{tok/s} \approx \frac{1.2\ \text{GB/s}}{\text{model size}}$$

260K @ 4-bit → fits in BRAM, DDR idle15M @ int8 → ~80 tok/s, bandwidth-bound

No amount of added DSPs would help.

Read Board Specs first

XC7Z007S — 14,400 LUTs · 60 DSPs · 225 KB BRAM

==Roughly ¼ of a Pynq-Z2.== Almost every FPGA-ML tutorial online assumes the bigger part. Check every resource claim.

doorbellmatmul opvector opint32 accint8weights + actsburst fillAXI-HP 64-bitK, Vctrl / statuslogits → tokenflush before kick