28 Sep AI Decision Workflows in Omniscope: Jev, LLMs and Automation
The AI Decision block in Omniscope is powered by TypeSafe AI’s Jev.
It takes the information in each row of your data, evaluates questions against a set of possible answers, and returns the results as fields you can use in the rest of the workflow.
This gives us another way to build applications around data. A support message can be classified before it reaches a queue. A product description can be checked for useful detail. A project proposal can be assessed against defined criteria, with the results available for someone to inspect and compare.
And here in this articel we want to show why the combination of LLMs + Jev + Omniscope is something to be aware of: an LLM can extract information from a document or generate a response; Jev can evaluate a question with defined outcomes; Omniscope can combine those results with the rest of the data and apply the rules around them. You can inspect each step separately instead of leaving the whole process inside one prompt.
What the AI Decision block does
The starting point is a question you need to ask about your records. You provide the relevant fields and define the possible answers, including descriptions that explain what each answer means. Be accurate as the quality of those definitions really matters: “Is this a good fit?” needs a clear account of what you are looking for.
Jev supports three kinds of answer:
1) A choice selects from a defined set, such as Billing, Delivery or Returns.
2) A yes/no question evaluates a condition.
3) A score evaluates the record against an ordered scale, with the possibility of values between levels.
The AI Decision block exposes these results, along with probability information and confidence for Choice and Score questions. The block documentation describes the configuration.
Those answers become columns beside your original data.
A classification can feed a filter, a score can enter a calculation, and an uncertain answer can become a record for someone to review.
You can build the next step around the result without first extracting it from a paragraph of generated text.


Combining Jev, LLMs and ordinary data logic
Consider a workflow that starts with a long support conversation. An LLM could extract a summary and the customer request. Jev could then classify the topic, assess urgency and identify whether escalation appears necessary. Omniscope could combine those assessments with account status, contractual response times and routing rules. An LLM could draft a reply once the relevant information has been assembled.
That is one possible arrangement. If the original message already contains everything needed for classification, you may not need an extraction step at all. If a simple rule can resolve a case, apply it before asking a model. Each model call should have a reason to be there.
LLMs can also classify and score records. Using Jev is a choice about the fit between a decision model’s structured outputs and the assessment you need to perform, rather than a claim that only one model can do the job. Accuracy and suitability still need testing against your own examples.
The arithmetic belongs in the workflow. Adding scores, comparing dates, joining tables or applying a threshold does not require a language model. Keeping those operations separate lets you inspect the calculation, change a rule and understand which part of the result changed. It also makes it easier to locate a failure: was the extracted information wrong, was the assessment wrong, or did the routing rule do something you hadn’t intended?
Classification, triage, data quality, and more.
Customer support is one application of this pattern. Jev could identify the topic of an incoming message and whether it contains a refund request, while Omniscope applies the rules for assigning the resulting record. A report can show the volume of each category and the cases waiting for review. Response generation can remain a separate step with its own checks.
Lead qualification uses a similar structure. The model could assess the apparent use case or fit against defined criteria, then Omniscope could combine that assessment with territory, company size, existing account status and previous activity. The resulting priority depends on both the model’s judgement and the business information already available. You can change the priority calculation without asking the model to reassess every lead.
Data quality has another useful boundary. Ordinary validation can identify missing values, duplicates and invalid dates. A semantic check can assess whether a description contains enough information, whether a category appears consistent with the text, or whether two parts of a record contradict each other. Those assessments can feed a review queue alongside the conventional checks.
Document processing can combine extraction with the same approach. Once requirements have been extracted into records, they could be assessed as supported, partially supported or unsupported against the supplied evidence. Omniscope can group the results, expose gaps and produce a report for review. The assessment is only as useful as the evidence and definitions behind it.
I could carry on, but I’ll finish by saying Jev can also evaluate generated content. A question about whether an answer addresses a requirement can supply another signal for accepting it, checking it further or sending it to a person. However… a second model’s approval still needs validation; agreement between models is not proof of correctness.
One example: an interactive Decision Scorecard
The Omniscope Decision Scorecard is a small example of this approach. You upload a CSV, explain what you are trying to decide, and enter questions and possible answers. Jev assesses each record, and Omniscope presents the results in an interactive ranking.
In this experience, each question’s score is converted to a common 0-100 scale and combined using the importance you assign to it.
The report presents a ranked list of records, with coloured bars showing each question’s contribution to the overall score.
Suppose reliability and ease of setup initially count equally. Moving the reliability slider gives that assessment more influence, and the ranking recalculates immediately.
A supplier with a strong delivery record and a more difficult setup may move above one that is easy to implement but less reliable. The movement indicators show how far each record has moved from its original rank.
There is no new call to Jev when you move a slider. The assessments stay the same; the calculation changes on the fly. That makes it possible to explore different priorities without mixing a change in your preferences with a fresh set of model answers.
Then, open a record and you can inspect the selected answer for each question, its score, its contribution to the total and the confidence or probability where available. The original data is there too. You can search for a particular record, leave a question out by setting its importance to zero, or reset the original weights.

Deciding what can continue automatically
A supplier at the top of the scorecard has scored highest under the questions, answer definitions and weights you supplied. That is the extent of the claim.
A missing fact, a poorly defined criterion or an incorrect assessment can all change the result, and moving sliders won’t repair a wrong answer.
Probability and confidence help you inspect uncertainty, but they are different outputs and neither guarantees correctness.
In a workflow, you could use them to decide which records need additional checks or human review. Any threshold for continuing automatically needs testing against examples where you know the expected outcome and understand the cost of getting it wrong.
Keeping the calculation separate obviously makes this easier to examine. You can trace a total back to its contributing scores, change the weights and check the arithmetic independently of the model. You still have to decide whether the questions are useful and whether the answers are credible.
Start with a repeated assessment you already understand, define its possible answers and check the output against familiar records. Then build the next step around it: a calculation, a review queue, a routing rule, another model call or an interactive application.
Omniscope lets that assessment participate in the rest of the data workflow, where you can inspect both the answer and what you have chosen to do with it.
That’s all folks!

No Comments