Getting Started¶
This page runs Stream end-to-end twice. Part 1 prices a small workload on a multi-core accelerator - no code generation, only the base install needed. Part 2 adds AIE code generation: it maps a SwiGLU block onto an AMD Ryzen AI NPU and emits the MLIR that AMD's toolchain compiles for the device.
Both assume you have installed Stream and are in the repository root, so the relative stream/inputs/... paths resolve.
Part 1 - A first run: 2-conv on a TPU-like accelerator¶
This needs only the base install (pip install -e .). We map a tiny two-layer convolution onto a multi-core accelerator and let Stream's MILP solver place every tensor and choose every transfer path.
The inputs¶
- Hardware -
tpu_like_quad_core.yamlis a system of four TPU-like compute cores plus a pooling engine, a SIMD unit, and an off-chip DRAM controller, wired together by an on-chip interconnect. - Workload -
2conv_1_8_32_32_16_32_3.onnxis two chainedConvlayers (a committed test fixture; only the tensor shapes matter for cost estimation, so the weights are cleared and the file stays tiny). - Mapping (optional) - a hand-written mapping YAML. Omit it and the mapping generator that claims the hardware proposes one: which cores each layer may run on, and how layers are tiled across cores.
Run it¶
from stream.api import evaluate_mapping
estimate = evaluate_mapping(
"stream/inputs/examples/hardware/tpu_like_quad_core.yaml",
"stream/inputs/testing/workload/2conv_1_8_32_32_16_32_3.onnx",
"outputs/first-run",
)
print("cycles:", estimate.cycles)
print("per group:", estimate.group_cycles)
Stream parses the hardware and workload, proposes a mapping, and runs the allocation of each fused group - generate tilings → estimate per-core cost → MILP allocation (the TransferAndTensorAllocator) → memory estimation. It finishes in a few seconds. cycles is the steady-state estimate summed over the fused groups, plus whatever reconfiguring the array between them costs on hardware that declares it. The solved estimate.context holds the scheduler, workload, accelerator and group_latencies.
evaluate_mapping takes the mapping YAML as its fourth argument, select_mapping picks the cheapest of several candidate mappings, and SolveOptions sets the solver backend (default "ortools_gscip"; also "ortools_highs" and "gurobi"), the columns and the constraint selection.
What you get¶
Everything lands under the output directory, one folder per fused group:
outputs/first-run/
└── group_0/ # one fused group of layers
├── mapping.yaml # the generated mapping that was used
├── tiled_workload.png # the workload after inter-core tiling
├── core_cost_lut.yaml # per-node, per-core cost estimates
└── tetra/ # the MILP allocation result
├── optimization_metrics.yaml # objective, solve time, gap, ...
├── slot_latency_breakdown.yaml # where the latency is spent
├── steady_state_trace.json # schedule trace (open in Perfetto)
└── steady_state_workload_final.png
The PNGs are the quickest way to see what happened: group_0/tiled_workload.png (how the layers were split across cores) and group_0/tetra/steady_state_workload_final.png (the resulting steady-state schedule). steady_state_trace.json opens in Perfetto for a timeline view. See Outputs for the full reference.
You can run the same call against any of the bundled example architectures or the swiglu workload - see the User Guide for the input formats.
Part 2 - AIE code generation: SwiGLU on the AMD Strix NPU¶
generate_code runs the same solve and then lowers it through the code generation backend that claims the hardware. Here we map a SwiGLU block onto the AMD Strix NPU and emit the MLIR that AMD's toolchain turns into an NPU binary.
Prerequisites¶
Code generation needs the AIE toolchain, which is not part of the base install (the wheels are platform-specific and git-hosted, so they cannot live in PyPI metadata). Install it once with the console script:
stream-setup-aie # add --dry-run to preview the steps first
This requires Linux x86_64 and CPython 3.12 or 3.13. See Installation for details.
The inputs¶
- Hardware -
stream/inputs/aie/hardware/whole_array_strix.yaml: the AIE array of the AMD Strix NPU. It has eight columns, each with a shim-DMA tile, a 512 KB memory tile, and four AIE compute tiles - a 4×8 grid of compute tiles. - Workload - a SwiGLU block: two projection
Gemms, aSiLUactivation, an elementwiseMul, and a down-projectionGemm, built for a problem size bymake_swiglu_workload. - Mapping - built from tile sizes by
make_swiglu_mapping.
Run it¶
from stream.api import SolveOptions, generate_code
from stream.inputs.aie.mapping.make_swiglu_mapping import make_swiglu_mapping
from stream.inputs.aie.workload.make_onnx_swiglu import make_swiglu_workload
workload = make_swiglu_workload(256, 512, 2048, "bf16", "bf16", last_gemm_down=True)
mapping = make_swiglu_mapping(256, 512, 2048, True, 32, 32, 64)
estimate = generate_code(
"stream/inputs/aie/hardware/whole_array_strix.yaml",
workload,
"outputs/swiglu",
mapping,
SolveOptions(nb_cols_to_use=8, stage_options={"npu": "npu2"}),
)
print(estimate.context.get("module"))
nb_cols_to_use=8 uses the full 4×8 compute-tile array, and npu targets the Strix (XDNA2) NPU. The MILP allocation over the whole array takes a minute or two; each fused group's design is written under outputs/swiglu/group_<index>/codegen/.
The generated MLIR¶
The output is an MLIR module in AMD's aie / aiex dialects - tile placement, compute cores, and the object-FIFO data movement for the whole SwiGLU block:
builtin.module {
aie.device(npu2) {
%0 = aie.tile(0, 0)
%1 = aie.tile(1, 0)
...
}
}
From MLIR to a running NPU binary¶
This .mlir is the hand-off point to AMD's AIE toolchain. The aie / aiex dialects it uses are exactly those of mlir-aie and its IRON programming framework. mlir-aie lowers and compiles the module - placing the cores, building the object-FIFOs, and generating the host control program - into an NPU binary (an xclbin plus an instruction sequence) that runs on AMD Ryzen AI NPUs (the npu2 target here is the XDNA2 NPU in AMD Strix).
In short: Stream decides what runs where and emits the MLIR; mlir-aie and IRON build that MLIR and deploy it on the device.
Where to go next¶
- User Guide - the workload, hardware, and mapping input formats in detail.
- Stages - what each pipeline stage does and how to extend the pipeline.
- Using Stream with an AI agent - the MCP server and IR models.