I was all about the work.
— Nicolas Cage, Interview Magazine, conversation with Marilyn Manson
An HTTP call is the easy case. It has one request and one answer. A speech detector is harder. Its model runs on a different engine on the server and in the browser, and the two engines will never be identical.
So we share the rules and leave the model to the host. The core turns each score into a decision: speech started, speech continues, speech ended. The host runs the model and feeds the core its scores.
The core makes the decision. The host carries out the duty.
This is part 3 of 3. Part 1 set out the method and the six steps, G1 to G6. Part 2 applied them to one carrier client and its HTTP adapters. This part applies them to voice activity detection, and then explains the build choices that both examples use.
VAD: share the rules, not the model¶
Voice activity detection (VAD) decides when a caller starts and stops speaking. A model scores each slice of audio, and that model runs differently on a server and in a browser. The rules that turn those scores into “speech started” and “speech ended” should not differ at all.
[G1] Requirement: use the same speech-state rules in native audio processing and in a browser worker, while each host supplies model inference.
[G2] Contract: PCM windows enter the engine. Each inference probability advances it to a VadFrame that contains the speech state and the audio output.
[G3] Core: the shared engine owns window preparation, smoothing and transitions. This snippet shows how an inference result advances the shared Rust state:
self.smoothed_prob = smooth(self.smoothed_prob, prob, self.smoothing_alpha);
let smoothed = Probability::new(self.smoothed_prob);
let (new_state, transition) = self.step_state(smoothed);
let (new_state, transition) = self.apply_speech_cap(new_state, transition);
self.sm = new_state;[G4] Host connection: native inference uses ONNX Runtime, and browser inference uses onnxruntime-web. Both feed probabilities back into the core.
The native pipeline calls Rust directly. Its inference boundary is:
let prob = {
let frame = self.vad.current_frame();
self.session.infer(&frame)?
};It then calls self.vad.consume_probability(prob.get()). The browser needs a JavaScript-facing wrapper around the same engine:
#[wasm_bindgen]
pub struct Vad {
inner: CoreVad,
}
#[wasm_bindgen]
impl Vad {
pub fn push_pcm(&mut self, pcm: &[i16]) -> usize {
self.inner.push_pcm(pcm)
}
}The wrapper also exposes construction, windows, configuration and probability consumption. The attributes mark the exported interface, and the build generates the JavaScript glue. Initialize that glue once, before you process audio:
import init, { Vad } from "./pkg/vox_vad_core.js";
await init();
const vad = new Vad(16000);
vad.push_pcm(pcm);pcm is an Int16Array from the audio input. For each prepared window, the browser worker runs ONNX inference and passes the returned probability to the same engine:
const r = await session.run({
input: inputTensor,
state: stateTensor,
sr: srTensor,
});
state.set(r.stateN!.data as Float32Array);
const prob = (r.output!.data as Float32Array)[0]!;
const frame = vad.consume_probability(prob);The surrounding worker keeps the context and tensors, converts each frame to speech events and releases the generated frame objects. The glue connects JavaScript and Rust. ONNX inference stays a separate integration.
[G5] Build VAD¶
[V1] One crate builds both library formats, with one feature for each host:
[lib]
crate-type = ["cdylib", "rlib"]
[features]
default = []
native = ["dep:ndarray", "dep:num-traits", "dep:ort", "dep:smallvec"]
wasm = ["dep:wasm-bindgen"]The native inference dependencies are target-gated. The wasm feature exports browser bindings for the shared engine.
The native build is a library, not a VAD server. Its inference loads the ONNX Runtime shared library at run time. The browser module uses a custom Cargo profile, wasm, which inherits release and sets lto = "fat" and codegen-units = 1. Run these commands from the repository root:
# [V2a] Native libraries: libvad.rlib and libvad.so
cargo build -p vad --lib --no-default-features --features native \
--target x86_64-unknown-linux-gnu --release
# [V2b] Browser module: target/wasm32-unknown-unknown/wasm/vad.wasm
cargo build --locked --target wasm32-unknown-unknown -p vad \
--no-default-features --features wasm --profile wasm
# [V3b] JavaScript bindings, named vox_vad_core
wasm-bindgen --target web --out-dir crates/vad/web/pkg --out-name vox_vad_core \
target/wasm32-unknown-unknown/wasm/vad.wasm
# [V4b] Optimize the Wasm module
wasm-opt -Oz --enable-reference-types --enable-bulk-memory \
--enable-nontrapping-float-to-int --enable-simd \
crates/vad/web/pkg/vox_vad_core_bg.wasm -o crates/vad/web/pkg/vox_vad_core_bg.wasmwasm-bindgen-cli must match the crate’s wasm-bindgen version. wasm-bindgen --target web picks a JavaScript module format; Cargo’s --target picks the platform. [V5b] Our build then bundles the worker with Vite and packs the onnxruntime-web files and the Silero model next to it. Hosting those files on Cloudflare still runs VAD in the browser. It does not create a Worker.
[G6] Verify¶
Today, the tests feed the engine scripted probabilities, either directly or through a scripted model session. One checks that a brief, confident burst counts as a barge-in on an 8 kHz telephony line. Two more check that splitting the audio into bursty chunks does not change the result.
The browser wrapper calls the same engine, but no test runs it yet. The next test to add feeds one probability sequence through both interfaces and compares the speech events. After that come end-to-end checks with the same PCM and model fixtures, with explicit inference tolerances and latency limits. A matching state machine alone does not prove matching inference or performance.
How one crate becomes several builds¶
Both examples in this series build one crate for more than one host. This section explains the choices behind those builds. Before any Wasm build, add the target once:
rustup target add wasm32-unknown-unknownWorkers also need worker-build, and Components need wasm-tools.
Five build choices¶
Every artifact in the series comes from the same five choices.
- Select the package with
-p. - Enable its Cargo features.
- Select the compilation target with
--target. - Build the configured executable or library outputs.
- Add the bindings or Component packaging that the host needs.
--targetchooses the platform. Features choose the code.crate-typechooses the output format.
Output formats: rlib and cdylib¶
The VAD crate sets both formats in its manifest:
| Format | Who uses it | Linux output | wasm32-unknown-unknown output |
|---|---|---|---|
rlib | Another Rust crate, through Cargo | libvad.rlib | libvad.rlib |
cdylib | A host outside Rust, through an exported interface | libvad.so | vad.wasm |
Listing both formats builds both for the selected target. It does not build Linux and Wasm together. cdylib does not export every Rust function, and it does not generate JavaScript or WIT bindings. The Rust Reference explains linkage and output formats.
Features choose the adapter¶
Cargo features select code and optional dependencies at compile time. They are not run-time switches. The platform crate gates each HTTP adapter on both the feature and the target:
#[cfg(all(feature = "cloudflare", target_arch = "wasm32"))]
pub mod cloudflare;
#[cfg(all(feature = "native", not(target_arch = "wasm32")))]
pub mod native;The host features of transport only forward to these, for example native = ["platform/native"]. Features are additive, so --all-features is not a portability check. Code for one target is not compiled when you build for another, so a clean native build says nothing about the Wasm adapters. Check each combination on its own:
cargo check -p transport --no-default-features \
--features telephony-control,native --target x86_64-unknown-linux-gnu
cargo check -p transport --no-default-features \
--features telephony-control,cloudflare --target wasm32-unknown-unknownRUSTFLAGS passes low-level options to rustc. It does not select a platform or turn on dependencies; use --target and --features for that. The Cargo Book covers features in full.
Bindings¶
We write the function. A binding generator writes the connection that lets another language or runtime call it. This is how each caller reaches the Rust code:
| Host | How the caller reaches Rust |
|---|---|
| Native | A direct Rust call |
| Browser | JavaScript calls a generated wasm-bindgen wrapper |
| Wasmtime | A generated host method invokes a WIT export |
| Cloudflare | The Worker’s JavaScript entry calls the Rust handler |
In Wasmtime, the Component is the guest. The host calls the functions that the guest exports, and the guest can call only the imports that the host supplies. An import is a request for a capability, not an implementation of it.
An adapter implements host behaviour. A binding lets a caller cross an interface boundary. A generated binding converts calls and values. It does not supply missing HTTP, storage, scheduling or inference. For C callers, cbindgen generates headers. For Kotlin or Swift, UniFFI generates bindings around a native Rust library.
One source, or one artifact?¶
- Source portability: one implementation, compiled separately for each host. The builds share code but ship different files.
- Artifact portability: one compiled file, run by several hosts. A Wasm Component can run in any host that provides its imports.
Both examples in this series give source portability. Artifact portability needs a common host contract, such as the same WIT imports.
Not shared yet: live sessions¶
Not everything fits the pattern yet. A live session with a speech model provider keeps a socket open for the whole call. The session code is already generic over its socket:
pub struct LiveSession<S: ClientConnection> {
reader: LiveReader<S::Reader>,
writer: S::Writer,
}Today, only the native socket connects to a provider. platform also has a browser WebSocket client, but only our widget uses it, for its own connection.
A Cloudflare session still needs its own socket handling, lifecycle and tests. A generic type parameter or a feature flag cannot create them. They only choose between adapters that already exist.
Where to draw the line¶
Go back to the bug from part 1. The response limit was a rule that had to behave the same on every host, and it lived in the adapters. Moving it into the one shared implementation fixed every host at once.
Both examples in this series follow that rule. Put everything that must behave the same into one implementation: carrier encoding and speech-state transitions. Give each host a small adapter for the part only it can do, such as HTTP or model inference. Then build each host on its own, and run the same fixtures through every adapter.
When you add a host, you should only add an adapter. If you find yourself changing the rules in the core, the line is probably in the wrong place.