$ needle-rs

Typed tool calls, right in this tab.

Needle 3 is Cactus Compute's 121M-parameter tool-calling model. This page runs it on needle-rs, a Rust engine compiled to WebAssembly. You type a command, you get back a schema-valid function call with a probability for every choice it made. Nothing leaves your browser: no server, no API key, and it keeps working offline after the first load.

$ starting the engine

// TRY_IT

tools

        
        
try one
$
history

    // RACE

    Cactus Compute ships its own browser engine for Needle 3, the one behind cactuscompute.com/needle. Race it against needle-rs, right here, in your browser: the same model, the same tools, the same commands, one after the other on the same thread.

    Loads Cactus's engine (0.7 MB) from their Hugging Face repo the first time.

    // NEEDLE_VS_JEV

    Jev, from TypeSafe AI, is a cloud model that answers typed questions (a Choice, a Score, a Noul) with probabilities instead of text. That's the same shape this page returns: typed values your code can branch on, plus how sure the model is. The difference is where it runs and what it costs.

    Jev (TypeSafe AI)Needle 3 on needle-rs
    ReturnsTyped answers with probabilities and confidenceSchema-valid tool calls, a confidence, and the probability of every enum option
    RunsTypeSafe's APIYour device: this tab, a native library, or the CLI
    Latency70–500 ms end to end, per TypeSafeAbout 0.25 s a command in this tab; about 25 ms native on an M3 Pro
    Price$0.042 per million input tokens, output freeFree, on hardware you already have
    WeightsClosed, early accessOpen (Apache-2.0, Cactus Compute); this engine is MIT
    PrivacyRequests go to the APINothing leaves the page; works offline
    SmartsA frontier-scale model built for judgment calls121M parameters: good at turning commands into calls, weak at open-ended judgment

    They aren't really rivals so much as two tiers. Needle handles the easy, frequent calls on the device for free, and its confidence tells you when a request is past it. Send only those to Jev or a bigger model:

    // Needle first; escalate only when it isn't sure.
    const r = JSON.parse(needle.decide(text, 256));
    if (r.function_calls.length && r.confidence >= 0.6) {
      act(r.function_calls);   // local, free, private
    } else {
      escalate(text);          // Jev, an LLM, or a person
    }

    Try "tell the innkeeper we need a room for the night" in the game example: Needle reaches for attack instead of speak, and either its low confidence or the engine's own grounding check flags it, which is exactly what the escalation path is for. And it's not a general classifier. In our tests, sorting support tickets into teams whose names never appear in the text got 2 of 4 right; that kind of judgment is Jev's pitch, not Needle's. Jev figures come from TypeSafe's launch post and docs; Needle figures are measured.

    // HOW_IT_WORKS

    // CREDITS

    Thank you to Cactus Compute for Needle 3, its weights, and the reference package this port was checked against, all under Apache-2.0. needle-rs is an independent, MIT-licensed port by Scott Pierce.