Rendered at 18:49:54 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
2 hours ago [-]
4 hours ago [-]
gingersnap 3 hours ago [-]
Is the local model similar to model2vec?
tgluck 3 hours ago [-]
Not really
they distill different things. model2vec distills a sentence transformer into static embeddings, so the output is a faster general-purpose encoder.
Jevstiller keeps the encoder frozen (bge-small by default) and distills Jev's decisions on one specific question into a small head on top of it
tgluck 7 hours ago [-]
Author here. This puts a proxy in front of repeated Jev classification calls. At first everything goes to Jev; from Jev's answers it trains a small head on frozen sentence embeddings, picks a confidence threshold with an exact finite-sample bound so that at most 2% of all requests get an answer Jev wouldn't have given, and then answers the confident share locally at ~15 ms on a CPU. A permanent 2% audit keeps checking; if agreement breaks, everything falls back to Jev and it retrains.
Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.
dotancohen 1 hours ago [-]
It would be great if we could correct Jev's incorrect answers, even on a separate endpoint. Let me tell it what Jev got wrong.
What type of head is that? What type of model is that head part of?
kodefreeze 2 hours ago [-]
Isn't this against their ToS? Useful for hobby stuff.
wedg_ 2 hours ago [-]
Woah cool idea. So it's almost a drop-in replacement for a typical Jev setup that just reduces your jev bill over time ?
they distill different things. model2vec distills a sentence transformer into static embeddings, so the output is a faster general-purpose encoder.
Jevstiller keeps the encoder frozen (bge-small by default) and distills Jev's decisions on one specific question into a small head on top of it
Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.
What type of head is that? What type of model is that head part of?