Loomcore

A C++ runtime that schedules ONNX models as a graph: a policy router, dynamic batching on two lanes, INT8 paths, binding deadlines, zero-downtime hot-swap.

Live runtime: checking

Timeline · One job

job-9 · Red fox · 39.2 ms

Recorded 1 Oct 2026

Recorded on 1 Oct 2026 on a local build of the same runtime (Intel Core i5-1145G7, Windows 11). Connecting to the live runtime.

FP32INT8router decisionskipreject / cancelDAG edge
40.2 ms

Ctrl + wheel zooms, drag pans, ← → step through events.Tap a bar or marker for details; + and − zoom, then drag to pan. Open in Perfetto: copy the trace, save it as trace.json, drop it on ui.perfetto.dev.

All 4 events in time order
mstrackwhatdetail
0.000Jobsjob-9completed after 39.0 ms
8.67CPU mobilenet FP3213.1 ms · job-9
21.9RouterRouter decisionCompositeRouter: LatencyBudgetPolicy: remaining budget 18.2 ms < threshold 30.0 ms; downgrading to INT8 for node 'bert_tiny'
35.0GPU_SIM bert_tiny INT84.00 ms · job-9

Result

What the job produced

Recorded 1 Oct 2026
Red fox
Red fox

mobilenet says

red fox

77.5% confidence

  1. red fox77.5%
  2. dhole14.4%
  3. kit fox4.3%
  4. red wolf3.0%
  5. grey fox0.2%
Per node
nodelaneprecisionms
mobilenetCPUFP3213.1
bert_tinyGPU_SIMINT84.00
jobbudget 40.0 ms39.2

Router decisions

  • bert_tinyLatencyBudgetPolicyremaining budget 18.2 ms < threshold 30.0 ms; downgrading to INT8 for node 'bert_tiny'

bert_tiny embedding

Embedded a photo of a red fox. The first 16 of 128 values:

  1. -1.000
  2. 0.027
  3. -0.999
  4. 0.978
  5. -1.000
  6. 0.765
  7. -0.996
  8. -0.540
  9. 0.055
  10. 0.022
  11. -0.582
  12. 0.060
  13. -0.013
  14. 1.000
  15. -0.878
  16. -0.574

Benchmark

FP32 vs INT8, measured

Recorded 1 Oct 2026

Intel Core i5-1145G7AVX2 presentAVX-512 presentAVX-512 VNNI presentAVX-VNNI absent

p50 latency per inference, milliseconds, single item, one thread per session
modelFP32 p50INT8 p50FP32 INT8INT8 speed-upon AMD Ryzen 9 6900HX
MobileNetV2 · static QDQ10.49.981.04x0.93x
bert_tiny · dynamic0.5360.4501.19x1.13x

This CPU has VNNI (int8 dot-product instructions), so ONNX Runtime's INT8 convolutions use them and INT8 MobileNetV2 is 1.04x as fast as FP32. The gain is modest: static QDQ quantization adds Quantize/Dequantize nodes, and at batch size 1 on one thread their cost eats into what VNNI saves.

Recorded machine: loomcore_bench, 20 warm-up and 100 measured runs per variant, 1 Oct 2026; ONNX Runtime 1.30, one intra-op thread per session (the scheduler supplies the parallelism). Latency only: the INT8 models are not checked for accuracy. Reference numbers from docs/BENCHMARKS.md.

Controls

Drive the runtime

Image

JPEG or PNG up to 2 MB, or one of six Commons photos.

Deadline
none

Binding: admission control rejects a job that cannot fit, and the reaper cancels one that runs out of time.

Checking the live runtime…

Graph

mobilenet → bert_tiny

Recorded 1 Oct 2026

Router chain

In order. A skip ends the chain; otherwise the first policy to set precision or lane wins. Changes apply to the next run, which hot-swaps the graph to the new router first.

  1. 1
    circuit-breaker

    CircuitBreakerPolicy(0.5, 4, 2000)

    Reads the node's recent success and failure outcomes; skips a node whose error rate reached 50% over at least 4 runs; lets one probe through after 2 s.

  2. 2
    bulkhead

    BulkheadPolicy(6)

    Reads how many requests for the node are in flight; sheds (skips) a request once 6 are already in flight for that node.

  3. 3
    confidence-gate

    ConfidenceGatePolicy(0.85)

    Reads mobilenet's softmax confidence for this image; skips bert_tiny once mobilenet is at least 85% sure; there is nothing left to describe.

  4. 4
    precision-planner

    PlannedPrecisionPolicy()

    Reads measured FP32 and INT8 p95 per node, and the job's budget; solves a 0/1 knapsack over the critical path for which nodes to run INT8 so the job fits.

  5. 5
    latency-budget

    LatencyBudgetPolicy(30)

    Reads the job's remaining time budget; drops a node to INT8 once less than 30 ms of the budget is left.

  6. 6
    load-aware

    LoadAwareBackendPolicy()

    Reads the CPU and GPU_SIM queue depths; sends the node to whichever lane has the shallower queue.

Scheduleradmission controldeadline cancellationprecision planningEDF lanes

Claims

Every claim, next to what checks it

From docs/CLAIMS.md. Each links to the test case or program that asserts it.

#claimchecked by
1The full test suite passes
2A job's promise is settled exactly once, even under a deadline reaper racing a node's natural completion (no std::terminate)
3Batching never mixes incompatible non-batch shapes for the same node
4Scheduler::shutdown() actually stops accepting new jobs (not a no-op)
5Admission control rejects a job the precision planner judges undeliverable, and dispatches nothing first
6The precision-downgrade knapsack finds the true minimum-quality-loss set, which a savings-only greedy would not
7The circuit breaker lets through exactly one probe after cooldown, never more, under concurrent load
8Runtime::reloadGraph hot-swaps the graph under continuous concurrent load with zero jobs lost
9The same hot-swap holds against the real mobilenet+bert_tiny pipeline, with per-swap timing reported
10The reference pipeline runs end to end (real image → classification → confidence-gated embedding) and reports real per-node p50/p95
11INT8 MobileNetV2 is measurably *slower* than FP32 on the reference (non-VNNI) machine — not fabricated, not silently skipped
12The ORT dependency is fully encapsulated: the benchmark links and runs with no ONNX Runtime include/link dependency of its own
13The C API is a real, separate consumable surface: a pure-C translation unit links and runs against it with no C++ Loomcore headers
14The installed package is consumable via plain find_package(Loomcore) — not just "it builds in-tree"
15CI builds and tests both platforms (Windows + Linux) on every push/PR
16A run's JSON-lines log converts to a real, structurally valid Perfetto/Chrome Trace Event Format timeline — lane tracks, per-execution duration bars, routing markers, DAG flow arrows