Board Specs — Know your silicon before you design for it

Part of Zero to NPU.

The Blackboard carries a Xilinx/AMD Zynq XC7Z007S-1CLG400C. This is the smallest Zynq-7000 part.

Do not trust Pynq-Z2 tutorials

Almost every FPGA-ML tutorial online assumes a XC7Z020 (53K LUTs, 220 DSPs, 630 KB BRAM). You have roughly one quarter of that. Designs that fit a Pynq-Z2 will not fit here. Check every resource claim against the table below.


The numbers that constrain you

ResourceAmountWhat it means for you
LUTs14,400Very tight. A stock Xilinx AXI DMA IP eats ~2–3K of these. You will write your own.
Flip-flops28,800Fine. Pipeline aggressively; FFs are the cheap resource.
DSP48E1 slices60 (datasheet says 66; plan for 60)1 int8 MAC each. Caps compute at ~32–64 MACs.
Block RAM50 × 36 Kb = 225 KBYour entire on-chip working set. A 260K-param int8 model barely doesn’t fit.
PS1× Cortex-A9 @ 666 MHzSingle core. No SMP. NEON is available and matters.
DDR3512 MB, 16-bit @ 533 MHzPeak ~2.1 GB/s, realistic ~1.2 GB/s. This is your real bottleneck.
PS↔PL4× AXI-HP (64-bit), 2× AXI-GP, 1× ACPHP for bulk weight streaming, GP for control, ACP for cache-coherent access.

Speed grade is -1 (slowest). Plan PL designs at 100 MHz, push to 150 MHz only after you’ve closed timing once.


What fits, and what doesn’t

Modelint8 sizeThroughput ceilingWhere weights live
stories260K260 KBthousands of tok/sOn-chip BRAM (at 4-bit) — no DDR traffic
stories15M15 MB~80 tok/sStreamed from DDR every token
TinyLlama 1.1B1.1 GB~1 tok/sNot happening on this board

This is why the capstone is two-tier — see phase-7-end-to-end.


Resource budget

The design does close. I worked this out against the real part:

BlockLUTsDSPsBRAM
8×4 systolic array (32 PEs)~1,50032
Weight + activation buffers~80012
Vector unit (norm/softmax/act/RoPE)~2,000124
Custom AXI4 burst master~7002
AXI-Lite register slave~300
Command sequencer~1,0002
Xilinx AXI interconnect~1,200
KV cache (128-token context)9
Total~7,500 / 14,40044 / 6031 / 50

~48% LUT utilization

That leaves real room for timing closure and debug logic. Good.

The remaining 19 BRAM blocks hold model weights for the Tier-1 capstone. 4-bit stories260K needs 29, so you’ll spill the embedding table to DDR and keep the 5 transformer layers on-chip. See §7.1 for the full arithmetic — that budget is what forces the 128-token context limit.


Why 8×4 and not 8×8

64 PEs needs 64 DSPs and you have 60 — and the vector unit needs a dozen.

32 MACs @ 100 MHz = 6.4 GOPS, roughly 5× what the A9 will do on int8. That’s a real win and it fits.