Introducing Lumma-Fev
Our decision model that scores your options in one pass.
The typical approach to structured decisions in production still asks a large language model to write a label, then parse whatever comes back. You pay for tokens you do not need, wait on decoding latency, and hope the format stays stable call to call. That pattern fits open-ended tasks. It is a poor fit when generation is not the product and all you need is to pick among options you already defined—for example route a ticket to the right team, rank urgency on a fixed scale, or pass a yes-or-no gate before automation runs.
Lumma-Fev does something narrower. Instead of asking the model to generate a decision, you ask it to score the options you supply. It reads your context once and returns probability distributions in a single forward pass—no autoregressive decoding, no JSON repair loop.
Lumma-Fev is a family of decision models: you pick a size, the interface stays the same. Every checkpoint is Apache 2.0 on our Lumma Decision Models collection— 0.1B (154M), 0.6B (649M), 4B, and 9B.
The interface
What you send and what comes back
You pass a state (ticket text, message, document snippet) and a dict of questions.
The model does not answer in prose. For each question key you provide, you get a structured result with a top choice and a probabilities map over the options you defined.
Example state:
I was charged twice for my March invoice. Please refund one of them. This is really urgent.
answers = model.decide(
state="My running shoes arrived in the wrong size. Can I exchange them for a size 10?",
questions={
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, wrong items, damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems",
},
}
},
)
Illustrative response:
{
"department": {
"type": "choice",
"choice": "returns",
"confidence": 0.9995,
"probabilities": {
"returns": 0.9997,
"shipping": 0.0003,
"billing": 0.0
}
}
}
Your application uses fields directly—e.g. answers["team"]["choice"] to route and answers["team"]["probabilities"]["billing"] to threshold automation.
Questions share the same encoded state but are scored independently; one answer does not change another.
Question types
Three question types
noulyes/no gate
choicepick one label from your list (up to 255 options)
scoreordered levels (low → high)
Same interface as TypeSafe Jev. Inference is prefill-only (no token generation), which keeps latency down versus LLM label parsing.
Comparison
Benchmarked on four tasks
We compared every Lumma-Fev checkpoint against TypeSafe Jev 1.13.0, Laya, and GLiNER-2.5-Decide on the same four accuracy benchmarks.
No matching rows.
Using the scores
Calibrate before you automate
Output scores are useful for ranking options, but they are not guaranteed to be calibrated probabilities out of the box. A score of 0.82 means “more confident than 0.60,” not “82% chance of being correct” until you fit thresholds on labeled data from your traffic. Many teams auto-route only above a high band, send middling scores to review, and escalate the rest—exact cutoffs depend on the cost of a wrong decision.
Try it
Try the interface
Demo on Hugging Face (4B & 9B)
For a live run on real weights, use our Spaces—they expose the same decide-style flow in the browser:
The family
The model family
- Lumma-Fev-0.1B (154M) — Built on Nandi-Mini-150M (pretrained from scratch by FrontiersMind), then fully fine-tuned for decisions. See the model card for a Super Mario demo on consumer hardware.
- Lumma-Fev-0.6B (649M) — Lumma-0.6B-Base with a frozen backbone and LoRA decision fine-tuning.
- Lumma-Fev-4B — Continual pre-training on the Qwen3.5-4B line, then decision fine-tuning.
- Lumma-Fev-9B — Continual pre-training on the Qwen3.5-9B line, then decision fine-tuning (~8B parameters on Hugging Face).
The API is the same at every size. We started at 154M to find a floor: useful routing decisions at tens of milliseconds when self-hosted.
Data
Training data
We trained Lumma-Fev on about 400k examples: roughly 200k synthetic samples and 200k merged from open datasets.
The mix includes English and ten Indic languages (Hindi, Malayalam, Marathi, Telugu, Tamil, Gujarati, Punjabi, Bengali, Kannada, and Odia).
Run it
Ways to run Lumma-Fev
1. Transformers
pip install -U "transformers>=5.17,<6"
from transformers import AutoModel
import torch
model = AutoModel.from_pretrained(
"FrontiersMind/Lumma-fev-9b",
trust_remote_code=True,
dtype=torch.bfloat16,
)
answers = model.decide(
state="I was charged twice for my March invoice. Please refund one of them.",
questions={
"billing": {"type": "noul", "instructions": "Is this about billing?"},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments and refunds",
"shipping": "Delivery problems",
"technical": "Bugs and outages",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today"],
},
},
)
print(answers["team"]["choice"], answers["team"]["probabilities"])
state can be text, a JSON object, or an array. Pin a revision with revision="…"; use a GPU with model.to("cuda"). Swap the repo id for 0.1B, 0.6B, or 4B.
2. HTTP API (lumma-fev-serve)
pip install "lumma-fev[serve]"
lumma-fev-serve --model FrontiersMind/Lumma-fev-0.6b --host 0.0.0.0 --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice. Please fix this ASAP.",
"questions": {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}}
}'
Same /v1/systemone contract as TypeSafe; optional LUMMA_FEV_API_KEY and --cors. Clients: curl, lumma_fev.Client, or typesafe-sdk.
Cost
What it cost to get here
Our total cloud GPU spend on Lumma-Fev training, synthetic data generation, and experiments came to about US$1,100.
FrontiersMind is a small AI research lab. We focus on efficient architectures and on improving cost and latency for useful model behavior.
Next
What we are building next
Next, we are training Lumma-BERT from scratch—our own encoder foundation. On top of that stack we plan to build the next generation of Fev models: the same Jev-like decision interface, with capacity and data we control end to end.
If you try Lumma-Fev on your own states and option lists, tell us where it breaks—that feedback feeds directly into Lumma-BERT and the Fev models that follow.