Typed tool calls, right in this tab.
Needle 3 is Cactus Compute's 121M-parameter tool-calling model. This page runs it on needle-rs, a Rust engine compiled to WebAssembly. You type a command, you get back a schema-valid function call with a probability for every choice it made. Nothing leaves your browser: no server, no API key, and it keeps working offline after the first load.
// TRY_IT
// RACE
Cactus Compute ships its own browser engine for Needle 3, the one behind cactuscompute.com/needle. Race it against needle-rs, right here, in your browser: the same model, the same tools, the same commands, one after the other on the same thread.
// NEEDLE_VS_JEV
Jev, from TypeSafe AI, is a cloud model that answers typed questions (a Choice, a Score, a Noul) with probabilities instead of text. That's the same shape this page returns: typed values your code can branch on, plus how sure the model is. The difference is where it runs and what it costs.
| Jev (TypeSafe AI) | Needle 3 on needle-rs | |
|---|---|---|
| Returns | Typed answers with probabilities and confidence | Schema-valid tool calls, a confidence, and the probability of every enum option |
| Runs | TypeSafe's API | Your device: this tab, a native library, or the CLI |
| Latency | 70–500 ms end to end, per TypeSafe | About 0.25 s a command in this tab; about 25 ms native on an M3 Pro |
| Price | $0.042 per million input tokens, output free | Free, on hardware you already have |
| Weights | Closed, early access | Open (Apache-2.0, Cactus Compute); this engine is MIT |
| Privacy | Requests go to the API | Nothing leaves the page; works offline |
| Smarts | A frontier-scale model built for judgment calls | 121M parameters: good at turning commands into calls, weak at open-ended judgment |
They aren't really rivals so much as two tiers. Needle handles the easy, frequent calls on the device for free, and its confidence tells you when a request is past it. Send only those to Jev or a bigger model:
// Needle first; escalate only when it isn't sure. const r = JSON.parse(needle.decide(text, 256)); if (r.function_calls.length && r.confidence >= 0.6) { act(r.function_calls); // local, free, private } else { escalate(text); // Jev, an LLM, or a person }
Try "tell the innkeeper we need a room for the night" in the game example: Needle reaches for
attack instead of speak, and either its low confidence or the engine's own
grounding check flags it, which is exactly what the escalation path is for. And it's not a
general classifier. In our tests, sorting support tickets into teams whose names never appear in the text
got 2 of 4 right; that kind of judgment is Jev's pitch, not Needle's.
Jev figures come from TypeSafe's launch post
and docs; Needle figures are measured.
// HOW_IT_WORKS
- Same engine as native. needle-rs reproduces the native Needle library bit for bit on macOS: all 192 bundled cases and every multi-turn test conversation.
- In the browser, the math library differs from Apple's, so a few results drift: 188 of the same 192 cases make the same calls, and confidence matches on the median case.
- WebAssembly SIMD runs the int8 matmuls and attention. The engine needs exact fused multiply-add, which WebAssembly lacks, so the page uses the CPU's own fma and int8 dot through relaxed SIMD when a check proves the browser computes them exactly, and an exact software fma otherwise. Both give byte-identical answers.
- Faster than Cactus's own browser engine in Chrome, Firefox and Safari's WebKit: a median command takes 239–359 ms here against their 373–419 ms, and reading the tools is 15–38% faster.
- "Score every option" scores each enum value in full (teacher-forced, token by token) and normalizes over the options, instead of sampling one.
- The model (35 MB) comes straight from Cactus Compute's Hugging Face repo on the first visit and is cached after that.
// CREDITS
Thank you to Cactus Compute for Needle 3, its weights, and the reference package this port was checked against, all under Apache-2.0. needle-rs is an independent, MIT-licensed port by Scott Pierce.