Just Because an LLM Can Do It Doesn’t Mean It Should

Just Because an LLM Can Do It Doesn’t Mean It Should

#ai #ai agents #architecture #genai #llm
Gilad Neiger
October 11, 2026

A new class of models is starting to emerge, and I think it highlights something increasingly important about the architecture of AI systems.

For much of the generative AI wave, one question has dominated technical conversations:

Which LLM should we use?

GPT or Claude? Large or small? Fast or reasoning? Commercial or open-weight?

That made sense.

For the first few years, much of the industry was focused on discovering what these models were capable of.

But as more AI systems move from experiments into production, the questions are changing.

The question is becoming less:

Can an LLM do this?

And more:

What kind of intelligence does this operation actually require?

At Develeap, our work with engineering organizations increasingly puts us close to what happens after the AI demo works: architecture, reliability, latency, cost, security, deployment, observability, and the operational reality of running AI systems in production.

Across these engagements, we’re seeing the conversation shift. The challenge is no longer only whether AI can perform a task, but how that capability should be engineered, operated and owned as part of a production system.

And one pattern is becoming increasingly visible.

As LLM capabilities have expanded, we’ve seen engineering teams use them across a much broader set of responsibilities – from generation and reasoning to classification, routing, scoring, verification and policy-related semantic checks.

That is understandable.

LLMs are extraordinarily capable general-purpose models, and in many cases they perform these tasks very well.

But capability alone does not make something the right architectural choice.

Generative models are increasingly being used for tasks that aren’t fundamentally generative

Imagine a customer sends this message:

“I was charged twice this month and need someone to refund the duplicate payment.”

Our system may want to understand:

  • Which department owns this?
  • How urgent is it?
  • Is the customer requesting a refund?

Once semantic understanding is required, a very natural implementation today is to send the message to an LLM and request structured output:

Structured LLM output: department billing, urgency high, refund_requested true

And to be clear: modern LLM APIs are increasingly good at this.

Structured Outputs can constrain model responses to a predefined schema, including enums, booleans, and required properties.

So I don’t think the old argument that “LLMs may return invalid JSON” is particularly interesting anymore.

The more important architectural question is different:

What capability are we actually asking the model to provide?

In this example, we are not asking it to write.

We are not asking it to investigate.

We are not asking it to design.

We are not asking it to solve an open-ended problem.

We already know the possible answers.

What we need is semantic understanding in order to make a bounded decision.

That is a different kind of operation from generation.

Yet in many architectures, both may still appear as exactly the same component:

LLM

A new category is starting to emerge

This is what makes models like Jev interesting.

Jev is a model developed by TypeSafe AI.

TypeSafe describes it as a “System One model” – their term for a model designed around fast, structured judgments rather than open-ended text generation.

Its interface is based on three main primitives:

  • Choice – choose between predefined alternatives.
  • Score – evaluate something against an ordered rubric.
  • Noul – estimate the probability that a statement is true.

So our earlier example could conceptually look more like:

Jev primitives: STATE, CHOICE between billing / technical / sales, SCORE urgency 0-2, NOUL "The customer is requesting a refund" → 0.96

The interesting part is not the syntax.

It is the contract.

The output is not text that happens to contain a decision.

The decision itself is the output.

And I think the emergence of models designed specifically around this kind of interaction is architecturally significant.

Because it challenges an assumption that has become easy to make:

semantic intelligence does not necessarily imply generative intelligence.

The real question isn’t LLM or Jev

There is an easy overcorrection here.

We could look at Jev and conclude:

We should replace LLM-based classification with decision models.

I don’t think that is the right conclusion either.

Classification is a problem.

Routing is a problem.

Scoring is a problem.

Verification is a problem.

None of them inherently belong to a single type of model.

For example, if the rule is:

sender_domain == "develeap.com"

then the right implementation is probably code.

If an organization has millions of labelled examples and a stable taxonomy, a specialized classifier may be cheaper, faster and more predictable than either a decision model or a general-purpose LLM.

If the problem is primarily semantic similarity, embeddings or a reranker may be the right tool.

If the requirement is a flexible semantic judgment over unstructured information with a bounded answer space, a decision-oriented model becomes particularly interesting.

And if the task requires multi-step reasoning about an unfamiliar situation, a reasoning model may still be the right choice.

So the engineering mistake is not:

using an LLM instead of Jev.

The mistake is:

choosing a model before understanding the kind of intelligence the operation requires.

From model selection to intelligence architecture

In many of the AI systems we see being built today, the architecture still tends to converge on a similar pattern:

Need intelligence? -> LLM

This is often completely reasonable in the early stages of a product.

One capable model reduces complexity.

It allows teams to move quickly.

It avoids premature optimization.

But as systems move into production and workloads grow, that abstraction starts to deserve more scrutiny.

Because “intelligence” is not one operation.

A production system may need to:

  • calculate,
  • enforce rules,
  • predict,
  • retrieve,
  • rank,
  • classify,
  • judge,
  • verify,
  • reason,
  • plan,
  • generate.

These are different requirements.

And increasingly, we have different primitives available for them.

I think this is pushing us toward a more explicit concept of intelligence architecture.

Something closer to:

Intelligence architecture: a system branching into Rules / Code, Specialized ML, Retrieval & Ranking, Judgment Models, Reasoning Models and Generation Models

Not every system needs all of these.

And importantly, Judgment is not a layer that every request should pass through.

It is another architectural primitive that can be introduced where a system specifically needs semantic judgment.

The engineering question becomes:

What is the most appropriate primitive that achieves the quality, cost and operational characteristics this specific responsibility requires?

Sometimes the answer will be a frontier LLM.

Sometimes it will be specialized ML.

Sometimes retrieval infrastructure.

Sometimes a decision model.

And sometimes it should still just be code.

Judgment as a first-class capability

This is the part I find most interesting about models like Jev.

Semantic judgment is not new.

LLMs have been performing it successfully for years.

What is changing is that judgment can now be treated as a first-class capability rather than something implicitly bundled into a general-purpose generative model.

Consider an agentic system.

It may need to determine:

  • Which agent should handle this request?
  • How risky does this action appear?
  • Does this output satisfy a requirement?
  • Is this document relevant enough to include?
  • Should this request be escalated?
  • Does this tool call look consistent with the user’s intent?

These are not necessarily generation tasks.

They are judgments.

Historically, it has often been convenient to ask the same LLM already participating in the workflow to make those decisions as well.

In practice, it is common for the same model to handle several of these responsibilities – from understanding the request and selecting tools to evaluating the result and producing the final response.

There is nothing inherently wrong with that. In many systems, the simplicity is worth it.

But as systems become more complex, separating these responsibilities can become operationally meaningful.

A dedicated judgment capability creates an explicit architectural boundary.

And explicit boundaries are useful because they can be evaluated, measured, optimized, and replaced independently.

Production changes the economics

This is not only an architectural discussion.

It also affects the economics of operating AI systems.

A decision-oriented model may be significantly cheaper than using a large general-purpose model for the same bounded semantic judgment.

But even that comparison is incomplete.

A deterministic rule may be cheaper still.

A specialized model running locally might outperform both economically.

A retrieval model may solve the actual problem without needing either.

So the meaningful question is not:

Which model is cheaper?

It is:

What is the lowest-cost architecture that achieves the quality and operational characteristics we need?

And I think “cost” itself needs a broader definition.

The cost of intelligence

When we work with production AI systems, model pricing is only one part of the equation.

There are at least four costs worth considering:

The cost of intelligence: compute cost, latency cost, engineering cost and failure cost

Compute cost

How expensive is each inference?

Differences that look insignificant during a pilot become material when a system performs millions of operations.

This becomes even more relevant in agentic systems, where one user interaction may trigger multiple internal model calls.

Latency cost

A production agent may perform several intelligent operations before returning an answer.

For example:

  1. understand intent,
  2. assess risk,
  3. select context,
  4. choose a tool,
  5. decide which model to use,
  6. verify the output,
  7. generate the final response.

If every one of those becomes a sequential call to a general-purpose generative model, latency compounds.

This does not mean every step should be extracted into a specialized model.

But it does mean the architecture becomes worth examining.

Engineering cost

Structured Outputs have removed much of the old friction around forcing LLMs into machine-readable schemas.

So “the LLM may return bad JSON” is no longer, in my view, a strong architectural argument.

The more relevant distinction is responsibility.

With a generative model, we are using a highly general model and constraining its output to behave like a bounded function.

With a decision-oriented model, bounded judgment is the native abstraction.

Architecturally, these communicate different intentions:

This component generates.

versus:

This component judges.

That distinction may look semantic at first.

In production systems, explicit responsibilities tend to matter.

They affect ownership, testing, monitoring and future optimization.

Failure cost

And then there is the cost that becomes increasingly important as systems become autonomous:

What happens when the model is wrong?

Every probabilistic model can make mistakes.

A decision model can make the wrong judgment.

An LLM can make the wrong judgment.

A specialized classifier can make the wrong prediction.

So production design needs to consider:

  • confidence,
  • thresholds,
  • fallbacks,
  • escalation,
  • evaluation datasets,
  • monitoring,
  • and the blast radius of an incorrect decision.

The important question becomes less:

How intelligent is the model?

and more:

How safely does the system behave when that intelligence is uncertain or wrong?

This is the kind of question that becomes much more visible once AI moves into real operational environments.

Judgment becomes something we need to engineer

Once judgment becomes an explicit architectural responsibility, it also becomes something we can engineer independently.

That matters.

A semantic decision such as:

“Does this action appear consistent with the user’s intent?”

is no longer just an instruction buried inside a larger prompt.

It becomes a component with its own:

  • evaluation set,
  • confidence thresholds,
  • failure modes,
  • latency and cost profile,
  • monitoring,
  • versioning,
  • and fallback behavior.

That changes the development lifecycle.

Instead of evaluating an AI system only by whether the end-to-end experience “seems to work,” teams can begin evaluating individual intelligent responsibilities independently.

Once judgment becomes an explicit component, it can be evaluated like one.

Teams can build evaluation sets around specific decisions, calibrate thresholds against production data, compare model versions before rollout, and monitor where those decisions begin to drift.

That makes the boundary useful not only architecturally, but operationally.

Decomposing intelligence does not only create more architectural options. It creates more things we can measure.

And measurability matters when AI moves into production.

The more intelligence remains hidden inside one large model call, the harder it is to understand which responsibility failed, why it failed, and what should be improved.

Making judgment explicit gives engineering teams another boundary they can observe, evaluate and evolve independently.

Production changes the questions

This is ultimately the part that matters most to me.

When AI is still in the prototype stage, the dominant question is naturally:

Can we make this work?

Once it reaches production, engineering organizations start asking different questions:

  • Can we run it reliably?
  • Can we measure it?
  • Can we afford it at scale?
  • Can we keep latency under control?
  • Can we understand how it fails?
  • Can we upgrade it safely?
  • Can the engineering organization actually own it?
  • What happens when a model is uncertain?
  • Which decisions should never have been probabilistic in the first place?

These are the questions we increasingly encounter as organizations move beyond AI pilots.

And they require a different level of architectural thinking.

The challenge is no longer only connecting a capable model to a product.

It is deciding how intelligence itself should be distributed across the system.

This is bigger than Jev

Jev is new.

“System One model” is TypeSafe’s terminology, not yet an established industry category.

We don’t know whether that terminology will become standard.

And we don’t know whether Jev specifically will become an important long-term part of the AI stack.

Like any new model, it needs to be evaluated against real workloads, real data and real alternatives.

But I think the direction it represents is worth paying attention to.

Because the appearance of decision-oriented models suggests that the AI stack is beginning to specialize.

And specialization is usually what happens when a technology matures.

The conversation becomes less:

Which model is smartest?

And even less:

Which model should we use everywhere?

The better question becomes:

Why are we assuming the same kind of intelligence should own all of these responsibilities?

From model selection to intelligence architecture

Software engineering has always been about choosing the right abstractions.

We don’t use the same database for every workload simply because one database could technically support them all.

We don’t run every workload on the most powerful machine available because more powerful means more capable.

We don’t make every architectural responsibility belong to the same component because that component can technically perform it.

We specialize.

We compose.

We make trade-offs.

AI engineering is beginning to look the same.

The architecture becomes less:

From APPLICATION → BIG SMART MODEL → TOOLS, to a SYSTEM branching into deterministic logic, specialized ML, retrieval & ranking, judgment capability and reasoning / generation

Mature AI engineering is not about specialization for its own sake.

It is about designing systems with clearer responsibilities, measurable boundaries, and explicit trade-offs around cost, uncertainty and failure.

For the last few years, the industry has invested enormous energy in making individual models more capable.

The next challenge is different:

How do we turn that capability into systems we can actually understand, operate and trust in production?

That is what I find significant about models like Jev.

Not because they give us another model to choose from.

Because they are another sign that AI engineering is moving from capability-first thinking toward deliberate system design.

The next phase of AI engineering may be less about making models smarter – and more about making architectures more intentional.