GLiNER2.5-Decide fine-tuned: Local System One model, Jev-level accuracy, ~7× faster
Hands-on guide to Jev and open-source GLiNER2.5-Decide System One models, their architecture, and how to fine-tune one for better accuracy and lower latency on your machine.
1. Introduction
Large language models are everywhere, and they are used for almost everything: from deciding whether text is a duplicate to solving million-dollar math problems.
The catch is that many tasks don't need a full LLM. Asking a chat model whether two questions are duplicates, or whether a ticket is urgent, is like bringing a bazooka to a gunfight: you pay for tokens you do not need, wait on token-by-token generation, and still have to scrape free text that can invent structure or wander off-schema. A short, typed answer would have been enough.
The shift came when TypeSafe shipped Jev: a “System One” model (referring to Kahneman’s "Thinking, Fast and Slow") that returns typed, calibrated answers instead of essays. With this, you get near-LLM performance with classical-ML latency: best of both worlds. But Jev runs behind an API, and while the model itself is quick, you still pay network latency on every call, and the weights are closed. For anything that has to stay in-house (privacy, cost, control), that is real friction. You want the same type of model, but something that runs locally.
So I fine-tuned an open decision model (GLiNER2.5-Decide) for duplicate detection and got Jev-level accuracy with ~7–8× faster serial inference than the API, fully on-device. The rest of this piece is how I got there.
2. Jev 101
If you are okay with a closed-source API, you can simply use Jev directly. Here’s the gist: Jev is TypeSafe’s System One model. You provide a state (all the facts needed for the decision) and a set of questions, and the API returns typed answers that are ready to use, not prose you have to parse. There are three question types: choice, score, and noul (a yes/no expressed as a calibrated probability). For more details, see TypeSafe’s introduction or the Jev model page.
Here is the smallest useful call, a single noul:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 | |
And that's it: you pass in a state, receive a probability, and can directly hook that into business logic without any complex parsing. Jev is now open to all (no more waitlist), and new signups get $5 in free credits; that's enough for about 119 million tokens at today’s rates (only input tokens are billed, not output).
3. Going local: GLiNER2.5-Decide
If you are not okay with a closed-source API because of privacy, cost at volume, or the need to beat a zero-shot model on your own data, you need something local. After Jev shipped, I looked at similar open-source releases and picked Fastino’s GLiNER2.5-Decide because its weights are on Hugging Face, it has a schema-driven Python SDK (classify_text / batch_classify_text), and inference supports Apple Silicon MPS.
What it is. Decide is not an LLM. It is a DeBERTa-v3-large encoder (overall ~435M parameters = ~340M in the transformer stack + ~95M embedding table). One model that can be used for classification, NER, structured records, and relations because it provides multiple heads (one for each task) on the same weights.
GLiNER2.5-Decide shares one DeBERTa-v3-large encoder across multiple task-specific heads. For this post, we will explore a model that can identify whether two questions are the same or not, i.e., duplicate detection. This task can be translated into a classification problem with two labels: duplicate and not_duplicate. Let's run the model on one example.
Suppose I have the following input:
{
"question1": "How do I reset my password?",
"question2": "What is the process to change my password?"
}
To run the model, I can use the classify_text module as follows:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 | |
Cool, it works! But how does it work under the hood? Let's see.
First, the input is formatted into a single string with the following schema:
( [P] label ( [L] duplicate [L] not_duplicate ) ) [SEP_TEXT] question 1 : … question 2 : …
Then, the string is tokenized into subwords using the DeBERTa tokenizer. The subwords are then passed to the DeBERTa encoder. The encoder returns one 1024-dimensional vector per subword. During preprocessing, the library records marker positions (subword indices for [P] and each [L]). At score time, it gathers only the [L] rows into a [n_labels, 1024] tensor. A shared MLP (1024 → 2048, ReLU, 2048 → 1) is applied to each label row independently (batched as one matrix multiplication), yielding one logit per label. For exclusive tasks like duplicate detection, softmax over those logits gives probabilities; argmax picks duplicate or not_duplicate.
A more detailed flow diagram is as follows:
To close it off, let's compare GLiNER2.5-Decide against classical BERT and Jev.
| BERT | Jev | GLiNER2.5-Decide | |
|---|---|---|---|
| Deployment | Local | Hosted API | Local (open weights) |
| Labels | Fixed head | NL criteria in request | Schema strings in-sequence |
| Output | Fixed logits | Typed + calibrated | Finite schema candidates |
| Generation | No | No (decisions) | No |
| Zero-shot new tasks | Weak | Strong (API) | Built for new schemas |
| Internals public | Yes | No | Yes |
4. Zero-shot run on GLiNER2.5-Decide and Jev
To test out the model, I used the GLUE QQP dataset. The dataset is a collection of pairs of Quora questions and a column that indicates if the questions are duplicates. The evaluation set is a balanced stratified sample of 1,500 pairs (750 duplicate / 750 not_duplicate), random_state=42 selected from the validation set. The inference for Jev is as mentioned above; for GLiNER2.5-Decide, I used the batch_classify_text module with a batch size of 8. The results are as follows:
Accuracy
| Model | Accuracy | F1 (duplicate) | Macro F1 |
|---|---|---|---|
| GLiNER2.5-Decide (zero-shot) | 69.7% | 0.724 | 0.695 |
Jev (jev-latest) |
79.3% | 0.759 | 0.788 |
Zero-shot on 1,500 balanced GLUE QQP pairs: Decide at 69.7% accuracy, trailing Jev's 79.3% by roughly a 10-point gap.
Latency
| Measurement | GLiNER | Jev | Ratio |
|---|---|---|---|
| Serial mean (headline) | ~68 ms | 520.1 ms | ~7.6× |
Amortized batch mean (batch_size=8) |
18.0 ms | 520.1 ms | ~29× |
| Throughput (Jev concurrency 12) | 55.6 ex/s | 22.6 ex/s | ~2.5× |
Local Decide inference vs the Jev API: ~7.6× lower serial latency and ~29× when amortized with batch_size=8 (Jev measured with concurrency 12).
As shown above, GLiNER2.5-Decide is about 7.6× faster than Jev in serial inference and about 29× faster in batch inference (note, Jev was run with a concurrency of 12). But it is still 10 points behind Jev in accuracy. This proves it's not really at the accuracy level of Jev.
I pointed this out in my post, and the Fastino team suggested fine-tuning the model, and a ➕ from their co-founder. So, I decided to take their advice.
Was the prompt unfair?
Looking at the prompt, you might suggest: Is it really fair? Jev is getting a little bit more context than GLiNER? I thought the same. I re-ran Decide with a Jev-aligned schema (same instruction and label descriptions), but the performance never improved, so it's safe to say prompt engineering is not the issue. Onwards to fine-tuning.
5. Fine-tuning Decide with LoRA
Zero-shot Decide was fast but roughly 10 points behind Jev on the same QQP slice. The Fastino team’s reply was straightforward: fine-tune it. I took the open checkpoint and adapted it with LoRA.
The training path is the one Fastino documents for GLiNER2: gliner2.training.ExtractorTrainer, with each example an InputExample (the same Question 1: / Question 2: string as the benchmark) and a Classification target on the label task. LoRA updates the encoder and task modules; only the adapter is saved. I sampled another 1500 rows from the validation set to use as the training set. The idea here was to prove the hypothesis that fine-tuning can close the gap between Decide and Jev, not to train a production-grade model. The training settings were as follows:
| Setting | Value |
|---|---|
| Base weights | fastino/GLiNER2.5-Decide |
| Adaptation | LoRA (r=8, α=16, dropout 0.05); checkpoint = adapter only |
| Epochs | 3 |
| Train batch size | 4 |
| Learning rates | encoder 1e-5, task heads 5e-4 |
| Schema at train & inference | {"label": ["duplicate", "not_duplicate"]} |
Early in training, eval accuracy hovered around ~58%, barely above chance on a balanced set. Over three epochs, train accuracy climbed toward ~91%, while eval (the disjoint 1,500) settled near ~80% by the end of training.
6. Results: now even better than Jev!
After training, I scored the LoRA adapter on the same 1,500-pair eval set used for the zero-shot comparison, and the results are much better now!
Same eval slice after LoRA: the adapter reaches 80.3% accuracy, matching Jev and lifting the zero-shot base by +10.5 percentage points.
| Model | Accuracy | F1 (dup) | Macro F1 | Correct / 1500 |
|---|---|---|---|---|
| Base (zero-shot) | 69.7% | 0.724 | 0.695 | 1,046 |
| LoRA adapter | 80.3% | 0.811 | 0.802 | 1,204 |
| Jev (earlier run) | 79.3% | 0.759 | 0.788 | 1,189 |
| Adapter minus base | +10.5 pp | +0.088 | +0.108 | +158 |
As shown above, the LoRA adapter led to an accuracy of 80.3%, an increase of 10.5 points from the base model, and an F1 score of 0.811, while still staying local and ~7–8× faster than Jev in serial inference.
Base zero-shot Decide leaned conservative on duplicates; the LoRA adapter improves both precision and recall on the duplicate class.
It also led to a decrease in false duplicates from 298 to 182, and an increase in correct predictions from 1,046 to 1,204. Precision and recall both improved as well.
Overall, the results are promising and show that fine-tuning can close the gap between Decide and Jev. Now we have a Jev class model that runs locally and is at least 7× faster!
7. When to use what
This experiment was one task (QQP duplicate detection) on one dataset (1,500 balanced validation pairs). It is not a universal leaderboard, but the tradeoffs are clear enough to choose a stack.
| Need | Suggestion |
|---|---|
| Best zero-shot + no ops | Jev API |
| Privacy, cost at volume, custom data | GLiNER2.5-Decide + LoRA |
| Inspectable encoder + schema API | GLiNER2.5-Decide |
| Complex reasoning tasks | Neither — LLMs are better for this |
I hope this article has helped you understand the trade-offs between Decide and Jev and how to choose the right model or fine-tune one for your use case.
Till next time, Cheers 👋