The dev loop
One number decides whether a framework is pleasant to work in: the time
between saving a Rust file and seeing the result in the browser. arc dev
exists to make that number small, and this page says what it currently is,
how it was measured, and where the time goes.
Numbers without a method are folklore, so the method is here in full and the machine is named. Reproduce it before believing it.
What one save costs
arc dev holds the TCP port itself and runs the application as a child
process, so a rebuild replaces only the child. The supervisor prints what each
part of the trip took:
cargo 7.85s (check 2.73s, codegen+link 5.06s) swap 0.21s spawn 1.44s boot 0.31s total 9.81s
Those stages are:
| Stage | What it is |
|---|---|
cargo | cargo build --features dev, spawn to exit. |
check | Start of the build to the last non-executable artifact. |
codegen+link | That boundary to the linked executable. |
swap | Stopping the old process and staging the new binary. |
spawn | Asking the operating system to start it. |
boot | The new process starting to its first accepted IPC connection. |
| Vite | Nothing. A .tsx, .vue or .css edit never reaches this loop. |
cargo dominates, so it is what the measurement below isolates.
How this was measured
arc new demo --stack react --db postgres, witharcaturepatched to the working tree so the framework under test is the one in the repository.cargo build --features devonce, cold, to filltarget/.- Change one line in
app/controllers/home_controller.rs– the string the welcome page renders – and timecargo build --features dev. Three times, each with a different string, so no run can be answered from a previous run’s cache. - A fourth run with
--timings, for the per-unit breakdown and the fresh/dirty split.
--features dev is what arc dev itself runs, so this is the same build the
loop performs, not an approximation of it.
The baseline
Measured 2026-08-21 on the machine described below.
| Measurement | Result |
|---|---|
Cold build, empty target/ | 52m 38s |
cargo build with nothing to do | 4.1s |
| One-line handler change | 33.2s / 34.3s / 37.6s |
demo.exe | 18.8 MB |
demo.pdb | 71.2 MB |
The --timings run breaks the rebuild into exactly two units of work out of
489 in the graph:
| Unit | Time | Of which |
|---|---|---|
demo lib | 50.6s | frontend 6.9s, codegen 43.7s |
demo bin | 39.4s | codegen of a nine-line main.rs, then the link |
Two dirty units, 487 fresh. That is the first thing the numbers settle: on a
Rust-only change nothing is rebuilt that need not be. Not arcature, not
arcature-macros, not the embedded scaffold templates, not a dependency. The
loop is not slow because it recompiles too much; it is slow because the two
units it does compile are expensive.
The second thing they settle is where inside those two units the time is. Type-checking the application crate – the part a developer thinks of as “compiling” – is 6.9 seconds of a 90-second trip. Everything else is code generation and linking, and the 71 MB of debug information is why: every frame of it has to be written by rustc, read by the linker, and merged into a program database on each save.
The machine
This is a small, busy machine, and the absolute numbers are worse than a developer laptop would show:
- Windows 11, x86_64-pc-windows-msvc, 4 logical CPUs.
rustc 1.98.0,cargo 1.98.0.- Microsoft Defender watching
target/. - Other Cargo builds running concurrently throughout. Cargo reported
Max concurrency: 1 (jobs=4 ncpu=4)for the timed run, and that run took 96.8s against 33-38s for the same work untimed – a two-to-three times spread from contention alone.
Treat the absolute figures as an upper bound and the shape – 5% frontend, 95% codegen and link, nothing spurious rebuilt – as the finding. The shape is what any change has to move.
Because the load varied, only measurements taken under --timings are
compared against each other below: those report per-unit compile time rather
than wall clock, and both the before and the after run reported the same
Max concurrency: 1 (jobs=4 ncpu=4). Plain wall-clock series taken minutes
apart on this machine differ by more than any change being measured, and are
not used as evidence for anything.
What was cut
The baseline points at one thing: debug information. Not the application’s
own – the scaffold has always built it with line-tables-only – but its
dependencies’.
The instinct is that a dependency compiles once and then sits in target/,
so its profile is a one-time cost. That is wrong for generic code. Every
Vec<MyThing>, every tokio combinator, every sea-orm query builder used
with the application’s own types is monomorphised into the application’s
crate, and its debug information is emitted by rustc and merged by the
linker there – on every save, for as long as the project exists. Nobody
steps through tokio while debugging a controller, so the scaffold now sets:
[profile.dev.package."*"]
opt-level = 2
debug = false
[profile.dev.build-override]
opt-level = 2
debug = false
Same machine, same application, same one-line change, both runs under
--timings:
| Before | After | |
|---|---|---|
demo lib | 50.6s (frontend 6.9s, codegen 43.7s) | 25.5s (frontend 4.9s, codegen 20.6s) |
demo bin | 39.4s | 19.8s |
| Both dirty units | 90.0s | 45.3s |
demo.pdb | 71.2 MB | 29.5 MB |
demo.exe | 18.8 MB | 18.8 MB |
Half, and the executable is byte-for-byte the same size, because none of this
was ever in it. Backtraces still carry file and line: the application’s own
crates were never touched. A developer who wants a step debugger through a
dependency can have it for one run with
CARGO_PROFILE_DEV_PACKAGE_tokio_DEBUG=2.
Three levers that look obvious are not taken, and the manifests say why:
opt-level = 0and a highcodegen-unitsare Cargo’s dev defaults. Writing them down changes nothing.split-debuginfois target-specific.rustc --print split-debuginforeportspackedas the only stable value on*-pc-windows-msvc, which is what MSVC already does by writing a.pdb. A fixed value in the manifest would be a no-op for some developers and a hard error for others.- A fast linker is already configured.
.cargo/config.tomlputs Windows on the toolchain’s ownrust-lld.exe, and leaves Linux and macOS on the system linker withmoldandwildas commented opt-ins – a config that fails on a machine without the tool is worse than a slow link.
What is left
The 2.5 second target is not met on this machine, and halving the cost was not enough to meet it. What remains, in order:
- Linking the executable. Even with a third of the debug information,
the
demobin unit is 19.8s for amain.rsof nine lines. Almost all of that isrust-lldpulling every rlib in the graph together. It is proportional to the size of the program, not to the size of the change, so it does not shrink as the diff shrinks. - Code generation for the application crate, 20.6s. This is
monomorphisation: the application instantiates a large amount of generic
machinery from
axum,tokioandsea-orm, and each instantiation is compiled into this crate.
Type-checking – 4.9s, and the only part proportional to what was actually edited – is already inside the budget. The loop is not slow because the compiler is slow at understanding the change; it is slow because the whole program is rebuilt around it.
The distance to the target, measured rather than scaled
The seconds above are per-unit compile time on a saturated four-core machine. They are the right numbers for comparing the before and after of this change, because both runs were taken the same way, and they are the wrong numbers for deciding how far 2.5s is. An earlier version of this page scaled them against issue #8’s quiet reading and put a post-change Cargo invocation near 3.8s, while saying plainly that the estimate was not settled and that somebody should re-measure on an idle machine.
Somebody has. Six one-line handler edits on an otherwise idle machine, warm
target, cargo build --features dev each time:
| Run | Wall clock |
|---|---|
| no-op build (nothing changed) | 1.2s |
| rebuild 1 | 5.8s |
| rebuild 2 | 4.3s |
| rebuild 3 | 4.5s |
| rebuild 4 | 4.3s |
| rebuild 5 | 4.4s |
| rebuild 6 | 5.5s |
warm cargo check (type-check only) | 1.6s |
So the Cargo half of the loop is about 4.4s, and the estimate was optimistic by roughly fifteen per cent – close enough to have been worth making, wrong enough to have been worth checking. Against a 2.5s target that is over by about 1.8x, not the fourfold the saturated per-unit figures suggest.
The split holds up and is the useful part. Type-checking a one-line change is 1.6s, comfortably inside the budget; everything above that is code generation and linking for the whole program, which does not shrink when the diff does. The loop is not slow because the compiler is slow at understanding the edit.
One trap for whoever measures next. A cargo check taken straight after a
cargo build reads about 90s on this project, and it is not the
type-check cost – check keeps its own fingerprints and artifacts, so the
first one after a build is cold. Run it twice and take the second; the 1.6s
above is a second run. A measurement script that interleaves build and
check will report the cold number every time and make type-checking look
like the bottleneck it is not.
The second thing is larger, and is missing from the list above because
--timings cannot see it. Issue #8’s own breakdown has spawn at 5.55s of
an 11.20s loop – bigger than the entire Cargo invocation – and identifies
it as Microsoft Defender scanning the 18.9 MB executable that was just
linked, reproducibly, at roughly 80x the cost of running a file it has
already seen. Nothing in this change touches it, and no profile setting can:
the scan happens after Cargo has exited. arc doctor already reports it with
the remediation. Anyone reading this page as the state of the dev loop should
read that stage as still the single largest one.
Getting to 2.5s therefore needs a structural change rather than another profile flag, and the candidates all have real costs:
- Fewer generics crossing the boundary.
-Zshare-genericsis nightly. Doing it by hand means erasing types at the framework’s public edges, which trades compile time against the type safety the framework exists to provide. - A different codegen backend.
rustc_codegen_craneliftis dramatically faster at-O0and is nightly-only, x86-64 Linux first. - Not relinking at all. Hot-patching the running process, as
subseconddoes, skips both remaining costs. It is a large piece of machinery and it does not survive every kind of change.
None of these is a patch-release change, so none of them is here. Issue #8 stays open with a measured number against it instead of a quoted one.