I Asked Jev 400 Questions at Once to See If It Would Flinch

· 9 min read

TL;DR — I tested whether Jev’s recommended “ask everything in one call” pattern actually holds up, over 250 calls and about a cent and a half.


TypeSafe’s documentation for Jev recommends a pattern they call speculative fan-out. Rather than asking one question at a time, you ask everything you might plausibly want to know in a single call, because the extra questions cost tokens but almost no extra time.

If you have not come across Jev, it is worth a moment of context. It is not a chatbot and it does not write sentences. You give it some state along with a set of questions whose possible answers you have defined in advance, and it returns those answers with probabilities attached. A very small example, which is the shape of everything that follows:

resp = client.system_one(
    model="jev-latest",
    state={"message": "You billed me twice for invoice A-2291."},
    questions={
        "team": Choice(
            instructions="Which team should own this",
            criteria={
                "billing":     "Payment, invoicing or refund issues",
                "engineering": "Bugs, outages or broken functionality",
            },
        ),
        "refund": Noul(instructions="The customer is asking for a refund"),
    },
)

resp.answers["team"].choice          # 'billing'
resp.answers["team"].probabilities   # {'billing': 0.94, 'engineering': 0.06}
resp.answers["refund"].noul          # 0.98

You get back 'billing' because 'billing' was one of the keys you handed it. There is no sentence to read and no JSON to pull out of a code fence. Since Jev never produces text one token at a time, it is not working through your questions in sequence either, and that is where the fan-out advice comes from.

The advice seems sound, but I could not find anyone who had tested it. The docs assert it and other write-ups repeat it, yet nobody shows what happens to the answers when a call goes from six questions to four hundred. Two things seemed worth knowing before building on the pattern: whether adding questions changes the answers already in the call, and whether Jev returns the same answer when the input does not change at all. It took 250 API calls and about one and a half cents to find out.

The Setup

The whole test depends on holding one signal still while everything around it changes.

The state is a support ticket with a bit of everything in it, so that the questions have something to disagree about:

TICKET = {
    "subject": "Charged twice, and now the export is broken",
    "plan": "Team, 14 seats, renewed 3 days ago",
    "messages": [
        {"from": "customer", "text": "You billed my card twice on the 3rd for invoice "
         "A-2291. I want the duplicate refunded today, not in 5-7 business days."},
        {"from": "customer", "text": "Also since the renewal our nightly CSV export "
         "returns a 500. This is blocking our finance close."},
        {"from": "customer", "text": "It has been two days. If this is not fixed by "
         "Friday we are moving to a competitor. I have been a customer four years."},
    ],
}

Then six probe questions about it, which between them use all three of the question types Jev offers:

PROBES = {
    "department": Choice(
        instructions="Which team should own this ticket",
        criteria={"billing":     "Payment, invoicing or refund issues",
                  "engineering": "Bugs, outages or broken functionality",
                  "success":     "Retention, relationship or account management"}),
    "churn":  Noul(instructions="This customer is likely to cancel in the next 30 days"),
    "anger":  Score(instructions="How angry the customer sounds",
                    criteria=["Calm and factual", "Irritated but civil",
                              "Openly threatening to leave"]),
    # plus priority (Choice), refund (Noul) and bug (Noul)
}

Choice picks one key from the criteria you supply. Noul is Jev’s term for a yes/no question and gives back a single probability between 0 and 1. Score places something on the short scale you describe, so anger comes back as a number near 2 on that three-step scale rather than as the word “irritated”.

For this ticket the answers are about what you would expect: refund returns 0.98, bug returns 0.80, churn returns 0.71, and anger returns 1.99 out of 3.

Those six never change. What changes is how much unrelated material surrounds them. The filler is generated from a list of topics — Noul(instructions="The customer mentions SSO"), the same for GDPR, tax, webhooks and so on — so every one of them is a perfectly reasonable question and all of them are beside the point.

questions = dict(PROBES)            # the six questions I care about

questions.update(filler(394, rng))  # 394 questions I do not care about

resp = client.system_one(
    model="jev-latest",
    state=TICKET,
    questions=questions,
)

There is no prompt anywhere in that. The state is the ticket as a dictionary, and each question carries its own instructions, so there is no system message to keep in sync and no output format to describe.

From there, three tests, each one the control for the next. The noise floor is the same request sent thirty times, so any variation can be pinned on Jev’s own jitter rather than on something I changed. The count sweep holds the probes fixed while padding the call to 6, 10, 25, 50, 100, 200, 255 and 400 questions, five repeats each. The order sweep keeps the total at fifty and shuffles where the probes sit, fifteen times.

All of it ran against two documents, the ticket and a two-star product review, so the result does not rest on one piece of text.

Adding 394 Questions Changed Nothing

Three probability answers held flat from 6 to 400 questions per call

Questions in the callChurnReal defectAsking for a refund
Asked alone0.7080.7930.980
60.7100.7980.980
500.7100.7920.980
2000.7080.7900.980
4000.7100.7940.980

The refund probability came back as exactly 0.980 in every condition, with a standard deviation across the whole experiment of 0.0000. The other two moved by roughly two thousandths, which is smaller than the variation I saw sending the identical call twice in a row.

Order did not matter either. Across fifteen shuffles at fifty questions the probes landed anywhere from position 1 to position 49, and I could not detect any effect from where they sat. The documentation is right on this point: the answers do not care what else is in the batch, and they do not care where in the batch they appear.

What It Costs

Milliseconds per question dropping from 60 to 1.8 as question count rises

QuestionsLatencyInput tokensCost per callCost per question
6360 ms673$0.0000283$0.0000047
50455 ms1,185$0.0000498$0.0000010
200418 ms3,291$0.0001382$0.0000007
400706 ms6,205$0.0002606$0.0000007

Those costs look like typos until you remember Jev charges $0.042 per million input tokens and does not meter output at all, so you are only really paying to send the document and the question text.

Going from six questions to four hundred is sixty-six times the work for less than double the wall clock. Per-question latency falls from 60ms to about 1.8ms, and per-question cost falls roughly sevenfold before flattening out, at which point you are simply paying for tokens. Four hundred typed judgements about one document, returned in about seven hundred milliseconds, for two hundredths of a cent.

One detail is worth noting rather than glossing over. The latency curve is not flat between 200 and 400 questions, where it moves from 418ms to 706ms. Nearly free is a fair description of the cost and roughly fair for the time, but the work does start to show up once you get into the hundreds. I also never hit a question limit, incidentally. Choice is documented as supporting up to 255 options, but a call carrying 400 separate questions went through without complaint.

It Is Nearly Deterministic, But Not Quite

This was the more interesting half for me, mostly because I could not find it written down anywhere.

Jev cannot produce an answer outside the options you define. Across every Choice answer in this experiment it never handed back anything outside the criteria I supplied, and that is structural rather than a matter of the model behaving well. It is worth being precise about what that buys you, though. It rules out output that does not fit your schema. It does not mean the category it picked is the right one, and it does not mean you will get the same category next time.

I sent the same request thirty times over. Not a similar request: the same ticket dictionary, the same six questions, the same keys in the same order, no timestamps or ids anywhere in the payload. The department classification came back as engineering twenty-nine times and success once.

Looking at the probabilities shows why. A typical call:

choice        'engineering'
confidence     0.12
probabilities {'billing': 0.27, 'engineering': 0.41, 'success': 0.32}

And the call that disagreed:

choice        'success'
confidence     0.05
probabilities {'billing': 0.28, 'engineering': 0.36, 'success': 0.36}

Two options at 0.36 each. The ticket genuinely is both an engineering problem and a retention problem, the model is reporting that honestly, and something has to break the tie. That was enough to look at it properly, so I took twelve Choice questions across the two documents and sent each identical request forty times.

Flip rate against reported confidence, 22 of 24 questions perfectly stable

Twenty-two of the twenty-four were perfectly stable: forty calls, forty identical answers. Only two moved at all. The department question on the support ticket flipped 7.5% of the time at a reported confidence of 0.107, and the complexity question on the product review flipped 2.5% of the time at 0.312.

The department question carried the lowest confidence in the dataset, which is the pattern you would hope for. It also says something about what that confidence number represents, and it is not what I had assumed. It is clearly not the probability that the answer is right: in the typical call above the confidence was 0.12 while the chosen option held 0.41. The gap to the runner-up was 0.41 minus 0.32, or 0.09, which is much closer to the reported figure, and on the tied call the gap was zero and the confidence 0.05.

I would put that no more strongly than this: across these results, confidence appears to reflect how separated the leading options are rather than how likely the chosen one is to be correct. Two examples are consistent with a gap, they do not establish the formula. Either way, a low value is telling you the top options were close together, not that the model expects to be wrong.

I would rather not make this tidier than the data allows, though. Questions with confidence scores of 0.198, 0.267 and 0.287 never flipped once in forty calls, so confidence is not a clean predictor of instability across the range. The narrower claim is the one the data supports: every unstable answer I found sat near the bottom of the confidence range, but plenty of low-confidence answers were perfectly stable.

There is also a real limit on what forty repeats can tell you. A question that flips once in two hundred calls would look completely stable here. So where I report zero flips, read that as a rate below roughly seven percent rather than as a guarantee.

What I Would Take From This

The fan-out advice is worth following. If you are already making a Jev call to classify a ticket, there is little reason not to ask the other twenty things you might need: the extra questions had no measurable effect on the existing answers and the marginal latency is tiny. That changes the shape of the design work, because you stop rationing questions.

Repeatability is the thing I would be careful about. If you are caching results, diffing outputs, or writing tests that assert on one exact answer, low-confidence Choice results can move even when the input is identical, so it is worth storing the confidence alongside the answer. Treating anything below roughly 0.2 as a flag seems reasonable, not because the answer is likely wrong, which this test does not show, but because it may not be stable, and an unstable answer eventually means a system that disagrees with itself.

Finally, keep an eye on latency as counts grow. Four hundred questions took nearly double the wall clock of six, which is still good but is not nothing, and if you are considering a thousand it is worth measuring on your own documents first.


The whole experiment cost about $0.015 and took an afternoon. The code is on GitHub and reruns with a TypeSafe API key.