yasp-toolkit embedded
Build a TensorRT engine for your target GPU from a PyTorch model: turn the model into compile inputs, run the compile, verify the engine's numerics against torch, and check on tasks and the workers that ran them.
The router host defaults to https://compile.yasp.ai/embedded and can be overridden with $YASP_EMBEDDED_ROUTER_HOST or the advanced-auth --router-* flags shown in --help.
wrap
Inspect a HuggingFace repo and write the model.py that embedded prepare takes.
Usage:
$ yasp-toolkit embedded wrap [OPTIONS]
| Option | Default | Description |
|---|---|---|
--hf <str> |
required | HuggingFace repo id to wrap (a local directory of a saved model also works). |
--output <file> |
model.py |
Where to write the generated model.py. |
--seq-len <int> |
128 |
Text sequence length to inspect the model with. |
--image-size <int> |
— | Square image edge; default is the processor's own. |
--audio-seconds <float> |
1.0 |
Seconds of audio to inspect with. |
--trust-remote-code |
— | Run modelling code the repo ships. Off by default. |
--task-heads / --no-task-heads |
task-heads |
Prefer a task head (logits out) on a bare checkpoint. --no-task-heads exports the pretrained model itself, with no randomly initialised head. |
--force |
— | Overwrite an existing file. |
yasp-toolkit embedded wrap --hf <HF_REPO_ID>
wrap loads the repo, works out its inputs, and writes the two-symbol model.py (Model plus get_fwd_args) that prepare takes, so a HuggingFace model needs no hand-written wrapper. Pass --trust-remote-code only for a repo whose modelling code you trust.
prepare
Generate the ONNX + NPZ inputs that embedded compile expects from a torch model.
Also captures the torch-eager outputs into a reference npz so
embedded verify can compare compiled artifacts against ground truth
without re-loading the torch model. Unless --bench-iters 0, the same
npz additionally records per-iteration host-torch timings (on CUDA when
available) that verify renders as the reference row of its benchmark
table.
The model file must define a Model class (nn.Module subclass) and a
get_fwd_args() -> tuple returning example positional inputs.
Ends with the plugin table. A HAS_PLUGIN false row there is a blocker, not
a note: the export left a placeholder node no .so implements, so embedded
compile cannot build an engine until that kernel is generated with Gaia.
prepare prints the commands for it.
Usage:
$ yasp-toolkit embedded prepare [OPTIONS] {model_file}
| Argument | Default | Description |
|---|---|---|
model_file |
required | Torch model .py defining Model and get_fwd_args. |
| Option | Default | Description |
|---|---|---|
--onnx-out <path> |
required | Destination path for the ONNX model. |
--npz-out <path> |
required | Destination path for the .npz sample inputs. |
--reference-out <path> |
— | Destination path for the torch-eager reference outputs npz. Default: |
--bench-iters <int> |
100 |
Timed host-torch iterations stored as benchmark_timings_ms in the reference npz, so embedded verify can show host latency next to each worker's. 0 disables. |
--bench-warmup <int> |
10 |
Warmup iterations before the timed host-torch benchmark. |
--trt-version <str> |
— | TensorRT version of the compile target (e.g. '8.5'). Enables the target-specific ONNX passes; without it TRT-8 targets silently miss rewrites worth ~10%+ latency. |
--capture-safe / --no-capture-safe |
— | Export a CUDA-graph-capturable engine graph: swaps capture-unsafe patchers and externalizes the sparse spatial-shape derivation as a coors_max engine input (I/O contract gains one INT32[4] input; the sample npz includes it automatically). |
--plugin <str> |
— | Bring-your-own plugin selection (repeatable): ' |
--debug-taps <str> |
— | Expose tap() markers in the model as extra reference outputs so verify compares intermediate tensors too. Repeat once per tap name (e.g. --debug-taps alpha --debug-taps beta), or --debug-taps all for every tap. Omit to expose none. |
--debug-plugin-taps |
— | Enables debug plugin taps. |
--stage <str> |
— | Staged model files only: export just these stages (repeatable; default: every stage in chain order). A stage whose export is known to fail can then be left out. |
--chain-min-cosine <float> |
0.999 |
Staged model files only: the cosine every result of the chain's own torch run must reach against the model file's entry() for the prepare to pass (integer results must match exactly). |
--dynamic-dim <str> |
— | Let one input axis vary at runtime, so a single engine accepts a range of sizes instead of only the example's (repeatable): 'get_fwd_args() returns; unlisted inputs keep the fixed size of their example. min/max bound what the engine will accept. The middle value, when you give one, is the size it is tuned to run fastest at; omit it and that comes from your example input instead (with a list of calibration frames, the FIRST frame). Either way the example's size must fall inside min..max, and verify both scores accuracy and measures latency at THAT size -- not at the tuned one, so a pinned opt makes the reported milliseconds describe a size the engine was not optimized for. Repeat the flag for further axes or inputs. |
See Debug taps for more information on taps.
--plugin selects the regions to carve out as custom-op nodes; the export then also
writes plugins.json and one kernel-generation candidate per plugin under
plugin_candidates/. See Optimize your model
for the selector grammar and the artifacts.
propose
Print the carve to build, from measured engine time.
Default: the recommended carve (per-family static plugins where a kernel
repays its generation job, one dynamic remainder per op, exclusions listed
with reasons). --merge-shape-families: one dynamic candidate per op.
--plugin ...: attribute cost onto an existing carve instead. Plugins
already inside the profiled engine are recognized automatically. Time that
cannot be attributed with certainty is reported as one 'unattributed'
bucket so the percentages stay honest.
embedded optimize builds this carve end to end — carve, generate, build,
compile, verify — with no further input.
Usage:
$ yasp-toolkit embedded propose [OPTIONS]
| Option | Default | Description |
|---|---|---|
--profile <file> |
— | Per-layer timings: the embedded compile output .tgz (profiled on the target by the system), its profile_layers.csv, or a trtexec --exportProfile JSON. |
--onnx <file> |
— | The prepared ONNX the profile was measured on. Needed (with --fx-nodes) only for FLAT frontend exports; scoped ONNX layer names carry their own provenance. |
--fx-nodes <file> |
— | fx_nodes.json written by embedded prepare — enables provenance for flat exports and the from...to span suggestion. |
--plugins <file> |
— | plugins.json from embedded prepare (default: looked for beside the profile). Already-plugged regions then rank as their own rows with a re-optimization proposal. |
--plugin <str> |
— | Attribute the measured cost onto THIS carve (same selector grammar as prepare --plugin) and report exactly these plugins. Repeatable. What the carve does not cover is still shown: coverable leftovers as recommended-carve suggestions (table only; embedded optimize never builds them), the rest under NOT WORTH. |
--checkpoint <directory> |
— | A model checkpoint directory (HF layout, config.json): print which shipped attention kernels could serve its attention geometry, before any export or profile exists. Alone, that is the whole report; with --profile it prints first. |
--no-carve |
— | Propose from scratch even when the workspace's plugins.json records a carve (by default an engine not yet plugged is reported against the carve you are about to build, under its own plugin names). |
--merge-shape-families |
— | Emit ONE selector per (kind, kernel, stride) with the channel axis left dynamic, instead of the default one-selector-per-channel-family. Only for deployments that genuinely run several shapes through one engine: measured on a fixed-shape engine, the channel-general kernel ran at 82% of the memory roof against 95% for the specialised one — 0.94x, slower than the kernel it replaced. |
--top <int> |
5 |
How many hotspots to show. |
--depth <int> |
5 |
Aggregate modules at most this many path levels deep. |
--estimates / --no-estimates |
estimates |
Attribute the anonymous fused (Myelin) engine layers by engine-order bracketing. ON by default: without it 98% of a TRT>=10 engine reports as one 'unattributed glue' line. --no-estimates reports ONLY measured, certain attribution. |
--verbose |
— | Also print the torch-class time table, per-proposal detail blocks, and re-optimization hints. |
propose reads the per-layer profile the worker already captured on the target — pass
the compile output .tgz directly — and attributes every millisecond back to the torch
module it came from, using the fx_nodes.json that prepare wrote beside the ONNX. Keep
that file (and plugins.json, for an engine that already contains plugins) next to the
prepare output it came from, or the regions cannot be named.
Three output modes, least to most user knowledge:
- default — the recommended carve: heavy regions split per shape family, a dedicated
static plugin where a generated kernel repays its job, one dynamic remainder plugin per
op for the rest. Regions that should not be plugged print under
NOT WORTH PLUGGINGwith the reason. --merge-shape-families— one dynamic candidate per op, channel axis left free. Only for deployments that genuinely run several shapes through one engine.--plugin <selector>— attribute the measured cost onto a carve you already ship and report exactly those plugins, with whatever the carve does not cover still visible.
Every row is a pasteable prepare --plugin selector, and optimize builds
exactly those rows in one run. See Optimize your model
for the whole flow, and Profile the model for the profile
it consumes.
optimize
Optimize a baselined model: propose, carve, generate, build, compile, verify -- in one run.
Starts from the baseline embedded prepare + embedded compile output and
ends with an engine carrying a generated plugin for every region propose
recommends, verified against the torch reference and measured against the
baseline. Every stage and every hosted job is recorded in the run
directory, so an interrupted run resumes with the same command, and a job
still running server-side is reattached to rather than paid for twice.
Usage:
$ yasp-toolkit embedded optimize [OPTIONS] {model_file}
| Argument | Default | Description |
|---|---|---|
model_file |
required | Torch model .py defining Model and get_fwd_args: the one the baseline was prepared from. |
| Option | Default | Description |
|---|---|---|
-b, --baseline <directory> |
required | The baseline directory: the embedded prepare output (model.onnx, fx_nodes.json) with the profiled embedded compile bundle beside it. A previous optimize run dir works too, to optimize further from an engine that already carries plugins. |
-a, --arch-spec <str> |
required | Target arch spec, as for embedded compile (e.g. arm64-sm87-dos6010-cu114-trt86). Routes the compile and the plugin builds, and names the TensorRT version prepare targets. |
-g, --gpu-name <str> |
required | GPU of the Gaia eval worker kernels are scored on, exactly as eval workers lists it (e.g. NVIDIA Jetson AGX Orin). |
-o, --output <directory> |
optimized |
Run directory for everything this flow writes. Rerun with the same directory to resume. |
--profile <file> |
— | The baseline's profiled compile bundle. Default: the one .tgz in --baseline. |
-p, --precision <str> |
native |
Engine precision for the compile, as for embedded compile (same default); must match the baseline's. |
-c, --kernelgen-config <str> |
— | KEY=VALUE forwarded to every kernel-generation job, exactly as kernelgen compile --config (e.g. -c iterations=4 -c min_speedup=1.0). Repeatable. The flow adds only deployment_target=tensorrt, which the plugin build requires. |
--from <str> |
— | Redo the run from this stage on: propose, prepare, kernelgen, build, compile or verify. |
--restart |
— | Discard the run's state and start over. |
--stop-after <str> |
— | Stop once this stage is done (e.g. propose to read the carve before spending on jobs). |
--attempts <int> |
3 |
Tries per hosted job on a transport failure. |
--backoff <float> |
60.0 |
Seconds to wait between those tries. |
--gaia-routing <str> |
— | Pin the kernelgen evaluation to Gaia workers registered with this routing name. |
--kronos-routing <str> |
— | Pin the compile to Kronos workers registered with this routing name. |
--progress / --no-progress |
progress |
Render transfer progress bars. |
optimize runs the whole optimize flow from a baselined model: propose on the baseline's
profile, prepare again with one --plugin per proposed row, one kernelgen compile per
plugin, one build-plugin per kernel, a compile carrying every .so, and verify. Each
stage's outputs land in the run directory (-o):
| Path | Written by |
|---|---|
carve.json |
propose — the rows being built, with their selectors |
model.onnx, plugin_candidates/ |
the carved prepare |
model.npz, model.reference.npz |
copied from the baseline, so both engines are compiled and verified on the same inputs |
sources/<Name>.module_source.py, sources/<Name>.report.json |
kernel generation |
plugins/<Name>.so |
the plugin builds |
<run>.tgz |
the compile |
optimize.json |
the run's state: every stage and every hosted job's task id |
Resuming. Rerun the same command on the same -o directory and it continues from the
first incomplete stage. Transport failures are retried --attempts times with --backoff seconds
between tries; a job that genuinely failed (a kernel that did not pass correctness, a build
error) is not retried, and the run stops at the end of that stage naming every plugin
without a result, since the compile needs all of them. --from <stage> redoes a stage and
everything after it; --restart discards the run; --stop-after propose is the cheap way
to read the carve before any job is submitted.
Two things the flow derives. The TensorRT version prepare targets comes from the
arch spec (...-trt86 is 8.6), so it cannot drift from the target the compile routes to.
Kernel-generation jobs carry deployment_target=tensorrt, which the plugin build requires;
every other setting reaches them through --kernelgen-config KEY=VALUE, exactly as
kernelgen compile --config would pass it.
Going further. A finished run is a valid --baseline for the next: propose recognises
the plugins already in its engine, and the run carries their kernels and .so files forward
so the next engine keeps them.
inspect
Summarize the plugin candidates emitted by embedded prepare.
Usage:
$ yasp-toolkit embedded inspect [OPTIONS] {prepare_dir}
| Argument | Default | Description |
|---|---|---|
prepare_dir |
required | A prepare --onnx-out dir, its plugin_candidates dir, or candidates.json. |
| Option | Default | Description |
|---|---|---|
--run / --no-run |
run |
Load each .pt2 and check it reproduces the captured output (needs torch). |
--guidance |
— | Write a kernelgen compile --config guidance=@. |
--emit-kernelgen <file> |
— | Write a runnable batch-submit script at this path (one ytk kernelgen compile per candidate, run in turn). Consistent with propose --emit-prepare: no flag, no script. |
yasp-toolkit embedded inspect <PREPARE_OUT_DIR>
inspect summarizes the kernel-generation candidates a prepare --plugin run wrote under plugin_candidates/, and by default loads each .pt2 to check it reproduces the captured output. --guidance drafts a guidance file per candidate for kernelgen compile --config guidance=@<file>, and --emit-kernelgen <path> writes a script that submits one kernelgen compile per candidate.
compile
Build a TensorRT engine from a prepared model on your target hardware.
Phases: create the task, upload model.onnx + sample.npz, start, poll until the task reaches a terminal state, then optionally download the result.
The task id is printed as soon as it's obtained, so a failure mid-flow
(upload, wait, download) is recoverable via the granular commands
(embedded get <task-id>, download …).
Usage:
$ yasp-toolkit embedded compile [OPTIONS]
| Option | Default | Description |
|---|---|---|
-a, --arch-spec <str> |
required | Worker arch_spec tag the task should route to (e.g. amd64-sm120-u2404-cu128-trt00). |
--routing <str> |
— | Pin the task to workers registered with this routing name (e.g. a dev/test pool registered with --tag routing=<name> / worker.tags.routing). Omit to use the shared stable pool. |
-p, --precision <str> |
native |
Precision tokens joined with _, in any order (e.g. native, fp16, fp16_int8, fp32_tf32). native builds strongly typed: the ONNX decides every layer's precision. Each worker advertises the tokens its target can build as its precision_tokens tag (see embedded workers). Ignored with --model-dir: each stage declares its own. |
--preflight / --no-preflight |
preflight |
Before creating the task, check that a worker for --arch-spec (and --routing) advertises every token of --precision (with --model-dir, every token any stage declares). Skip it for a worker registered since, or a router that cannot list workers. |
--onnx, --onnx-model <file> |
— | Local ONNX model to upload as model.onnx. |
--npz, --samples-npz <file> |
— | Local NPZ sample input to upload as sample.npz. |
--model-dir <directory> |
— | A staged model as embedded prepare left it (the --onnx-out directory: chain.json, one export per stage, the chain's gate). ONE task builds every stage with the precision and plugins its declaration names and packages ONE artifact whose runner executes the chain. Exclusive with --onnx/--npz/--custom-plugin/--calibration-cache/--dataset; --precision is ignored (each stage declares its own). |
-o, --output <file> |
— | Where to write the resulting output.tgz (the demo log lives inside it as share/.../generated/demo_run.log — extract with transfer extract). Omit to skip the download step. |
--dataset <str> |
— | Name of a registered dataset (see dataset add/list) to attach to the task. |
--refresh-manifest |
— | Force a fresh dataset manifest even if a still-valid one is cached. |
--calibration-cache <file> |
— | Path to a pre-built INT8 calibration cache (TRT-XYZ-EntropyCalibration2 format, ~1 kB), produced by embedded calibrate with the same --precision. Skips live calibration on the worker. When omitted and the precision permits int8, the worker calibrates against the dataset attached via --dataset (regular manifest path). |
--custom-plugin <file> |
— | Path to an extra TensorRT plugin shared library to register for this compile (repeatable). |
--tag KEY=VALUE |
— | Attach a user-facing key=value tag to the task (repeatable, e.g. --tag env=prod --tag region=ca). Tags are for identification only; they do not affect routing. |
--progress / --no-progress |
progress |
Render transfer progress bars. |
The demo log lives inside output.tgz as share/.../generated/demo_run.log — pull it out with transfer extract.
calibrate
Produce an INT8 calibration cache from a dataset on your target hardware.
Phases: create the task, upload model.onnx + the dataset manifest, start,
poll until the task reaches a terminal state, then optionally download the
resulting calibration.cache.
Feed that cache to compile --calibration-cache, with the same --precision,
to skip live calibration on subsequent builds.
The task id is printed as soon as it's obtained, so a failure mid-flow (upload, wait, download) is recoverable via the granular commands.
Usage:
$ yasp-toolkit embedded calibrate [OPTIONS]
| Option | Default | Description |
|---|---|---|
-a, --arch-spec <str> |
required | Worker arch_spec tag the task should route to (e.g. amd64-sm120-u2404-cu128-trt00). |
--routing <str> |
— | Pin the task to workers registered with this routing name (e.g. a dev/test pool registered with --tag routing=<name> / worker.tags.routing). Omit to use the shared stable pool. |
-p, --precision <str> |
required | The precision the cache is for, which must permit int8 (e.g. fp32_int8, fp16_int8). Required, because a cache only fits the precision it was calibrated for: build the engine with compile --precision set to the same value. |
--preflight / --no-preflight |
preflight |
Before creating the task, check that a worker for --arch-spec (and --routing) advertises every token of --precision. Skip it for a worker registered since, or a router that cannot list workers. |
--onnx, --onnx-model <file> |
required | Local ONNX model to upload as model.onnx. |
--dataset <str> |
required | Name of a registered dataset (see dataset add/list) the worker calibrates against. |
--input-name <str> |
— | Name of the model input the dataset feeds; disambiguates models with multiple inputs. |
-o, --output <file> |
— | Where to write the resulting binary calibration.cache. Omit to skip the download step. |
--refresh-manifest |
— | Force a fresh dataset manifest even if a still-valid one is cached. |
--custom-plugin <file> |
— | Path to an extra TensorRT plugin shared library to register for this calibration (repeatable). |
--progress / --no-progress |
progress |
Render transfer progress bars. |
build-plugin
Cross-compile a TensorRT plugin from source for your target hardware.
Phases: create the build, upload the source file, start, poll until the build reaches a terminal state, then download the resulting shared object.
The build id is printed as soon as it's obtained, so a failure mid-flow
(upload, wait, download) is recoverable via download.
Usage:
$ yasp-toolkit embedded build-plugin [OPTIONS]
| Option | Default | Description |
|---|---|---|
-a, --arch-spec <str> |
required | Build spec to cross-compile for; routes to plugin builders advertising this tag (e.g. jp61-cu126-trt103-py311). |
-s, --source <file> |
required | Plugin source file to build. |
--plugin-name <str> |
— | Override plugin_name in the bundle's plugin_meta.yaml before building. Kernel generation names a plugin after the operation it implements, not after your selector, and TensorRT resolves plugins by (namespace, name, version) — so a generated bundle usually needs renaming to match the node prepare emitted. Use this instead of hand-editing generated source. |
--input-format <str> |
cuda |
Source language of --source. The build backend validates the values it supports. |
-o, --output <file> |
required | Where to write the compiled shared object (plugin.so). |
--progress / --no-progress |
progress |
Render transfer progress bars. |
The source is the module_source.py a Gaia kernel generation produced — CUDA header, CUDA source, and the plugin_meta.yaml declaring the TensorRT contract. Build one .so per plugin, then pass them all to compile with --custom-plugin.
Pass --plugin-name on every call, set to the <PluginName> the matching prepare --plugin
selector inserted into the ONNX: TensorRT resolves a plugin by (namespace, name, version), so a
.so whose name does not match its node cannot be resolved at build time. It rewrites
plugin_name: in a copy of the source, never in the generated file itself.
cancel
Usage:
$ yasp-toolkit embedded cancel [OPTIONS] COMMAND [ARGS]...
Options:
* --help: Show this message and exit.
Commands:
compilation: Cancel a running compilation task.calibration: Cancel a running calibration task.
yasp-toolkit embedded cancel compilation
Cancel a running compilation task.
Usage:
$ yasp-toolkit embedded cancel compilation [OPTIONS] {task_id}
| Argument | Default | Description |
|---|---|---|
task_id |
required | Task ID. |
Options:
* --help: Show this message and exit.
yasp-toolkit embedded cancel calibration
Cancel a running calibration task.
Usage:
$ yasp-toolkit embedded cancel calibration [OPTIONS] {task_id}
| Argument | Default | Description |
|---|---|---|
task_id |
required | Task ID. |
Options:
* --help: Show this message and exit.
Cancellation only applies while the task is still pending or started; the router answers with a 409 once it reached a terminal state (completed / failed / cancelled).
get
Fetch a task's (compile, calibration or plugin build) current state.
With --wait, poll until terminal.
Usage:
$ yasp-toolkit embedded get [OPTIONS] {task_id}
| Argument | Default | Description |
|---|---|---|
task_id |
required | Task ID (a compile, calibration or plugin-build id). |
| Option | Default | Description |
|---|---|---|
--wait |
— | Poll every 15s until the task is completed or failed. |
key
Print a scoped key a Kronos worker can register with.
Use the JWT wherever a Kronos worker asks for YASP_API_KEY:
export WORKER_KEY=$(yasp-toolkit embedded key)
kubectl create secret generic yasp-api-token --from-literal=token="$WORKER_KEY"
Usage:
$ yasp-toolkit embedded key [OPTIONS]
| Option | Default | Description |
|---|---|---|
--audience <str> |
embedded-api |
Token audience claim. |
--scope <str> |
— | Token scope (repeatable). Defaults to embedded-register only. |
workers
List embedded workers (optionally filtered by tag).
Usage:
$ yasp-toolkit embedded workers [OPTIONS]
| Option | Default | Description |
|---|---|---|
-t, --tag <str> |
— | KEY=VALUE filter on worker tags. Repeatable. |
--detailed |
— | Include worker ID and user columns. |
verify
Compare each compiled .tgz's outputs against the torch reference.
Reads the reference npz produced by embedded prepare, then extracts
share/.../sample_inputs.npz, share/.../generated/sample_outputs.npz and, when the
worker shipped one, share/yasp_binary/build.json from each tgz. Sanity-checks that the
worker saw the same bytes across every tgz (otherwise comparing outputs is meaningless),
then pairs every output tensor BY NAME with the reference's and scores it — MAPE / MSE /
max-diff / cosine (the same definitions the worker's build-time check uses), and top-K
agreement when the output shape looks classification-y. Every dtype the engine emits is
compared at its true shape and dtype (float16, int64 and bool included).
When a tgz's outputs npz carries the worker's benchmark_timings_ms
telemetry array, a per-tgz benchmark table (min/mean/p50/p95/p99/max ms)
is appended after the accuracy table. A reference npz produced by
embedded prepare with --bench-iters > 0 carries host-torch timings
under the same key; those render as the table's first row, so each
target's latency reads against the host baseline.
Usage:
$ yasp-toolkit embedded verify [OPTIONS] {reference} {tgzs}...
| Argument | Default | Description |
|---|---|---|
reference |
required | Reference outputs npz produced by embedded prepare -- or, for a staged model, the prepare directory itself (its chain_reference.npz: the chain's results by ref, plus every loop's step count). |
tgzs... |
required | One or more compiled output .tgz files to score against the reference. |
| Option | Default | Description |
|---|---|---|
--max-mape <float> |
— | Exit 1 if any output tensor's MAPE (100 * sum|ref-out| / sum|ref|) exceeds this percentage. |
--min-cosine <float> |
— | Exit 1 if any output tensor's cosine similarity to the reference falls below this. |
--fx-nodes <file> |
— | fx_nodes.json from embedded prepare: with it, the per-output cosines (debug taps included) are attributed to modules and the regions that INTRODUCE the drift are ranked. |
--remove-outliers |
— | Drop MAPE outliers via 1.5x IQR before averaging (mirrors the inference notebook). |
--topk <int> |
— | Report top-K classification agreement vs reference. Auto-on for 2D outputs with last-dim > 100. |
--json |
— | Emit machine-readable JSON instead of a Rich table. |
--verbose |
— | Output all missing outputs and profiling layers. |
verify compares every array in the reference against the corresponding array in each build's outputs — the model outputs plus any debug taps exposed at prepare time — skipping only the benchmark telemetry. Each becomes a row scored with MAPE / MSE / max-diff / cosine (and top-K when a shape looks classification-y).
After the accuracy table, verify also prints a benchmark table (min/mean/p50/p95/p99/max ms). Each .tgz contributes a row from the worker's benchmark_timings_ms telemetry, and — when the reference npz was produced by prepare with --bench-iters > 0 (the default) — a torch host row from the host-torch timings, so you can read the compiled-vs-eager speedup directly. This replaces pulling demo_run.log by hand for latency.