An app and a skill make Jev genuinely useful

An app and a skill make Jev genuinely useful
The workflow needed for Jev Classify

TL;DR You can use Jev easily with the Jev Classify web app. You can transform somewhat hazy questions or rubrics into a Jev-friendly question set using the jev-question-translator skill.

Suppose you want to compare how all fifty states regulate a particular activity. You have gathered the statutes. You may even know which features you care about: who is covered, what conduct is prohibited, which exceptions apply, who may enforce the law, and what remedies are available. Knowing what you want to learn, however, does not get you the answers. Someone still has to read all fifty statutes and record, jurisdiction by jurisdiction, how each one handles each feature. Until recently, that someone was a person working by hand.

Or suppose your materials are a thousand judicial opinions on shareholder oppression. You want to know which theories the plaintiffs asserted and which succeeded, which defenses the defendants raised and which succeeded, and how each case came out. Again, the information is in the opinions. Extracting it means reading every one of them against a lengthy coding form. No one but an extremely motivated professor or their youthful research assistants would engage in such a quest.

Or suppose you are writing a brief. You would like useful feedback as you revis. But your greedy. You don't want a single lengthy critique from a colleague delivered a week after you gave it to them (followed by the gentlest of nags). And you don't even want an AI assessment delivered a ten minutes after you ask for it. Rather, you want your draft treated as a sporting event with a running commentary that updates as the draft changes. Would that not be cool?

These three projects look different, but they share a structure. In each there is a body of text to examine, a set of judgments to make about that text, and a reason to make those same judgments over and over: across fifty statutes, across a thousand opinions, or across successive versions of one draft. Whenever a project has that shape, three things decide whether it is practical: how fast each judgment can be made, how much each judgment costs, and whether the judgments are made consistently and intelligently from one document to the next.

Why frontier models have not solved this

You could, of course, hand much of this work to a frontier language model lab. Anthropic, OpenAI, x.AI, and some of their competitors can read a statute or an opinion and return an answer in a structured format. Many smaller models that run on a local machine can do so as well. So what is the difficulty?

The difficulty is cost and, with it, hesitation. Every document a frontier model reads consumes time and a substantial number of tokens, and tokens for high-quality models are not cheap. A million tokens fed into GPT6 Astra and 100,000 tokens out cost $15 as of right now. That expense compounds because a set of questions is rarely right the first time. You run the whole collection, look at the results, and discover that three of your questions were phrased badly. Fixing them means running the whole collection again—and paying the whole cost again. Faced with that prospect, many researchers on limited budgets never begin. A project that would be valuable if it could be iterated cheaply is abandoned because it can only be run expensively. Oh, and it is also slow. If I stopped every ten seconds to see if Claude felt my latest bon mot improved the exposition of this blog entry, I would never get to publishing it.

The marginal cost of local large language models is close to zero. But the sacrifice is speed and accuracy.

What changed on September 15, 2026?

On September 15, 2026, TypeSafe AI opened access to a model called Jev. The immense publicity it has received since that time is deserved. With Jev, the kind of work described above becomes fast and cheap. Just as important, correcting mistakes becomes fast and cheap. If a question turns out to be poorly framed—whether the fault is yours or the model’s—you revise the question and run the collection again, rather than living with the error because a rerun is unaffordable.

Jev is a classification model. It does not write prose, and it does not converse. It will not destroy the world. Instead, it takes a document and a list of structured questions and answers every question about that document, returning for each answer a probability rather than a paragraph. Run it across a whole collection of documents and you have, in effect, filled in a grid: every document down one side, every question across the top, an answer in every cell. (A mathematician would call this an “outer product” of the documents and the questions. A computer scientist might analogize the operation to a map reduce operator.) That grid is exactly what the fifty-state survey, the opinion study, and the brief-checking editor all need.

I described Jev in an earlier post, “What Jev Might Mean to the Legal World,” and listed legal tasks it seemed suited for. What that post did not solve is the practical problem. Jev was built for software developers. A lawyer who is not a programmer cannot easily use it at all, and even a programmer would want some structure that permits asking the same questions repeatedly and consistently rather than improvising each time.

Two new tools: Jev Classify and jev-question-translator

Today I am announcing two software products intended to close that gap and make Jev easy to use.

The first is a web application called Jev Classify. Right now, you can find it at https://jev-classify.netlify.app/ The second is a companion “skill” for AI assistants called jev-question-translator. It's been approved by lawve.ai and you can access it here: https://lawve.ai/@seth-chandler/skill/jev-question-translator. Together they let lawyers use Jev without learning to program and without learning the particular form in which Jev expects its questions. You never see a programming interface, and you never write a line of code or a line of the data format (called JSON) in which Jev’s questions must be expressed.

Instead, the process looks like this. You upload your documents to the web app, or paste in text. You state what you want to know, in ordinary English—even if your questions are still vague—or you supply a rough rubric or checklist. Then one of two tools converts those hazy notions into the precise, structured questions Jev needs. You can let the web app itself do the conversion. Or, for more careful work, you can use the jev-question-translator skill inside a large language model such as Claude, ChatGPT or Grok, have it produce a JSON file of questions, and load that file into the web app. Here, for example, is a sample prompt to the jev-question-translator skill.


/jev-question-translator Write questions that determine whether a liability insurer in Texas breached a duty to settle. Use the Midpage connector to help you find relevant cases and case law.

The output is the JSON file shown below that could be uploaded to the Jev Classify app. Here's a screen capture of some of that file.

Whether the structured Jev questions result from the app itself doing the work or an external soure such as the jev-question-translator skill doing the translation (with an upload of the results to Jev Classify), you end up with your documents and your questions in the form Jev requires. From there, hit "Run", and Jev does the rest.

I have used the skill within Claude to build two question sets: one for classifying felony-murder cases and one for federal justiciability. The justiciability set contains 62 questions. All I gave the skill was a single sentence:

“Let’s test it out. use /jev-question-translator to generate a battery of questions that determines factors relevant to whether the plaintiffs have standing in a federal case and whether there are any other barriers to justiciability.”

The web app can do much the same thing on its own, though perhaps not as well. My advice: if you want the most careful translation possible, use the skill to create the JSON question file and import it into the web app. If you want something quick, let the web app prepare the questions itself. For many purposes that will be fine.

Why translating questions is harder than it looks

You might wonder what work the frontier model is doing in this process. The answer is a lot; it actually takes a fair amount of thinking to move from our high-level human ideas about what to look for in a statute, case or brief, to something a basic language model can understand. This is so because a lawyer’s short question usually conceals a great deal. “Does this plaintiff have standing?” sounds like one question. To answer it, though, one must know the governing doctrine, attend to who the particular plaintiff is and what relief that plaintiff is requesting, and make several intermediate judgments along the way—about injury, about causation, and about whether the requested relief would actually help. “Does this brief use precedent effectively?” presupposes some notion of what effective use looks like and which precedents were the right ones to use. Even a seemingly crisp question such as “Does this statute create a private right of action?” requires decisions about express authorization, implied rights, cross-referenced provisions, and what the research project intends to count as a yes.

One could do this unpacking by hand and then format the results as JSON. It is tedious, however, and few lawyers would enjoy it. A better division of labor is to give the task—or at least a first draft of it—to a large language model, which is good at exactly this kind of decomposition. The language model takes the lawyer’s question and breaks it into what I will call atomic questions. An atomic question asks for a single judgment, one that can be made from the supplied text alone, without first having to resolve some other question. A checklist that a human could keep in mind may become dozens of atomic questions for Jev.

Application one: the fifty-state survey

Return now to the three opening problems and consider what Jev makes possible in each.

The fifty-state comparison is an instance of what public-health researchers call legal epidemiology: systematically describing legal rules so that one can study how laws vary across places and over time, and how those variations relate to outcomes in the world. The bottleneck has always been coding—reducing each statute to a common set of features. Jev can do that coding.

Consider how a single feature expands. A feature called “enforcement” might become several questions: Does the supplied text authorize enforcement by the state attorney general? By some other public agency? By a private party? Those possibilities can coexist—a statute may authorize all three—so separate yes/no questions fit better than a single multiple-choice question that forces one answer. A different feature might call for a choice among expressly defined categories, with “not addressed in the supplied text” available as one of the answers, so that silence is recorded as silence rather than forced into a category.

The payoff is a dataset with the same columns for every jurisdiction. A researcher can review the classifications Jev was unsure about, compare Jev’s coding against human coding on a sample, and—when statutes are amended—reuse the same question set rather than starting over. The resulting dataset can then be linked to health or social outcomes, which is where legal epidemiology earns its name.

Application two: a thousand opinions

The same method applies to the second opening problem, a large collection of judicial opinions. Suppose you are studying how courts treat a particular defense. For each opinion you could record whether the court mentions the defense at all, which of several specified tests it invokes, which factors it expressly discusses, and whether that defense actually contributes to the court's decision. With well-drafted questions, you can distinguish an argument a party made from a holding the court adopted, and an issue mentioned in passing from one the court actually resolved.

The answers become columns in a spreadsheet, one row per opinion, rather than facts scattered across hundreds of case summaries. Once the spreadsheet exists, the questions you can ask of it multiply. You might filter for cases in which the court applied one test but never discussed a particular factor, and then read that subset closely. You might look for patterns in the data itself—whether the success of one defense correlates with the presence of another, for example. The close reading of the text still happens, either by you or by an intelligent agent; Jev’s contribution is to tell you where to look.

Application three: a brief that is checked as you write

The third opening problem, feedback on a brief, points toward a different kind of use: continuous evaluation during writing. I should say plainly that this application does not yet exist; it is a possibility that Jev’s speed opens up. Maybe one of my readers should build it.

Imagine an editor that checks your draft every ten seconds, or after every substantial revision. It could ask whether the opening identifies the ruling you are requesting, whether a particular assertion is supported by the passage you cite, whether a given paragraph that you select actually answers the counterargument it purports to answer, and whether the conclusion follows from the reasons offered. It could rate the persuasiveness of each section on a scale from 0 to 10 and update the ratings as you type. Writing a brief could thus take on some of the character of an interactive game.

Naturally, such an editor would face the same translation problem discussed above. “How persuasive is this on a scale of 1 to 10?” cannot be put to Jev as it stands. (I mean it could, but I don't think the results would be good). Persuasiveness depends on the audience, the purpose, the evidence, and the applicable legal standards. The preparation step would identify several distinct dimensions of persuasiveness and describe recognizable levels of each. A short program would then take Jev’s answers on those dimensions and combine them into a composite response for the writer.

What would make such an editor practical is precisely speed and cost. An evaluation that is affordable once may be prohibitive when repeated hundreds of times in a single writing session. Jev makes that frequency affordable. This blog entry is about 6,000 tokens. I could check it 1080 times over about three hours and the total bill would about 26 cents. I can afford that! A dashboard vibe coded with an AI could further process Jev's responses and give me great feedback. (Like that last sentence is kind of lame: "great feedback.")

Here's a mockup of my imaginary Live Legal Editor app.

A fair question is why any of this intermediary machinery is needed. The answer is that TypeSafe built Jev for developers. Jev is offered as an API—an “application programming interface,” meaning a service that other programs call—and its answers come back as structured data meant to be read by software, not by people. To use Jev directly you would need to obtain an API key (a credential that identifies your account), write every question in the JSON format Jev requires, convert each PDF or Word file to plain text, send a separate request for every document, and then assemble the returned probabilities into a table. None of this is particularly hard for a programmer. For most lawyers, however, it is the point at which the project stops. (TypeSafe’s explanation)

A related question is why one cannot simply do all of this inside ChatGPT or Claude. You can ask either assistant to classify documents, and both are genuinely useful in developing questions. But telling an assistant to “use Jev” does not connect it to TypeSafe. An actual call to Jev requires credentials, correctly structured questions, and software that sends the request and receives the answers. An assistant equipped with the right tools can make that call—but the tools have to exist first.

The workflow, step by step

Here is the full workflow in Jev Classify, using a statutory comparison as the example.

1.     Add the documents. On the Prepare tab, drag files into the document area or browse your computer to select them. The web app accepts PDF, DOCX, plain text, Markdown, and HTML. You can also paste text and give the pasted document a name. Give each statute a useful name identifying its jurisdiction and version. Preview the extracted text. Scanned PDFs need to go through text recognition before the app can use them.

2.     Choose how the questions will be prepared. If you already have focused questions, enter them one at a time under Write in English, or use Paste a list. Keep each question’s answer choices together with the question and check the grouping preview to confirm the app has divided them correctly; you can type %%% between questions to make the breaks explicit. If you want the most careful translation, prepare the questions with the skill instead and load them in step 3.

3.     Load externally prepared questions, if you have them. Ask your assistant to save its questions as JSON. The jev-question-translator skill does this job for you. You put in hazy English language questions and it produces a JSON "rubric" that Jev can use.

Select Load questions JSON and choose that file from your computer. The same route lets you reuse questions saved from an earlier session. You never write JSON yourself.

4.     Open Connections. You need a TypeSafe API key to run Jev. Once you have an account with TypeSafe, you can get one here: https://console.typesafe.ai/keys If you are asking the web app to translate English questions, you also need an API key from OpenAI, Anthropic, or OpenRouter, and you must select a language model to do the translating. A frontier-grade model is worth using when the translation requires careful interpretation. If you loaded ready-made JSON questions, no translation is needed, and the TypeSafe connection is the only one required.

  1. Ask the language model to convert the questions to Jev form.

    

  1. Inspect and save the prepared questions. Check their wording, their definitions, and their permitted answers. Jev supports three kinds of answers: a yes/no probability (a "noul"), a choice among specified alternatives, and a score against described levels. These are different kinds of answers and should be read differently; in particular, a yes/no probability of 0.5 means Jev is uncertain, not that the document half complies. You can edit the prepared questions and download their JSON immediately, before running Jev at all.

7.     Run and use the results. Select Run with Jev. Each document is submitted with the prepared question set. The Results tab displays a table you can transpose and inspect, including the probability breakdowns for each answer. Export CSV to work in a spreadsheet, or JSON for further work with an assistant or another program.

A few practical details apply to every run. Approving individual questions is optional, and an Approve all button handles the entire set at once. Valid, included questions run without individual approval; incomplete questions are skipped. You can clear documents, clear questions, or start a new workspace. Save your exports before refreshing or leaving the page, because at present the working documents, questions, and results exist only in the open browser session.

What we have measured so far: speed and cost

Two things we can report with some precision are speed and cost.

In an exercise involving student writing, we ran 183 prepared questions across fourteen papers, producing 2,562 individual assessments. The recorded Jev request times totaled about 7.2 seconds. Almost all of the human effort went into preparing the questions, not running them.

As to cost: TypeSafe currently charges $0.042 per million input tokens, and output tokens are free. At that price, the reported usage for the student-writing run comes to roughly 1.3 cents for Jev’s work. All but the most impecunious legal writing professors might regard that as a bargain, even if they used it only to double-check grading they had already done, tediously, by hand. And the question set is a one-time investment: you pay to prepare it once and can reuse it as often as you like.

Turning Jev’s answers into a grade or into feedback took further judgment. We used the results to explore a weighted composite score and to draft a 500-word feedback document for a student. The detailed questions functioned as diagnostics. Simply averaging all of them would have been a mistake: it would have given extra weight to whichever rubric category happened to generate the most questions. The original purpose of the inquiry must govern how the smaller answers are combined.

Turning the output back into English

The justiciability example shows how Jev’s output can be converted back into ordinary language.

We ran the 62 justiciability questions against a hypothetical fact pattern. A Texas shrimpers’ trade association, one of its members, and a retired shrimper sue the Secretaries of State and Commerce over Mexican boats that fish illegally in U.S. waters. They ask the court to order the Secretary of State to request consultations with Mexico under an executive agreement, and to order the Secretary of Commerce to restrict Mexican shrimp imports under a statute providing that the Secretary “may” do so. The government responds that the harm is controlled by private Mexican fishermen and by the Mexican government, neither of them a party to the case, and that the President has chosen not to press the issue. A later diplomatic note and a proposed import rule add mootness and ripeness wrinkles.

Jev answered all 62 questions with a recorded request time of 193 milliseconds. Its classifications pointed to independent third parties as the source of the harm and to requested relief whose effectiveness depended on those parties changing their behavior.

A large language model given those classifications can reconstruct them into a judicial opinion, legal memo, or whatever other format is desired. Indeed, the attached file contains just such a judicial opinion reconstructed from the original hypothetical and the Jev Classify output. Here are two excerpts.

Excerpt 1:

Texas Gulf Shrimpers Association v. Secretary of State

Sep 27, 2026 · @Mathlawguy

UNITED STATES DISTRICT COURT, SOUTHERN DISTRICT OF TEXAS, BROWNSVILLE DIVISION

TEXAS GULF SHRIMPERS ASSOCIATION, ROSA DELGADO, and TOM REYES, Plaintiffs, v. SECRETARY OF STATE and SECRETARY OF COMMERCE, Defendants.

MEMORANDUM OPINION AND ORDER

The plaintiffs ask this court to order the Secretary of State to press Mexico for consultations and to order the Secretary of Commerce to restrict imports of Mexican shrimp. The defendants move to dismiss for lack of subject-matter jurisdiction under Federal Rule of Civil Procedure 12(b)(1). Because no plaintiff has shown an injury that the requested relief would likely redress, the motion is GRANTED and the complaint is DISMISSED without prejudice.

Background

The Texas Gulf Shrimpers Association is a nonprofit trade association of about 140 commercial shrimp-boat owners, based in Port Isabel. Its bylaws state its purpose as "protecting the livelihood of Texas shrimpers and the health of Gulf fisheries." Plaintiff Rosa Delgado, a member, owns and captains trawlers out of Brownsville. Plaintiff Tom Reyes is a retired shrimper who is not a member.

Excerpt 2:

II. The Association

The Association's only allegation of harm to itself is that it spent $80,000 on a public-awareness campaign and a catch study. That expenditure is not an injury caused by the defendants. "[A]n organization that has not suffered a concrete injury caused by a defendant's action cannot spend its way into standing simply by expending money to gather information and advocate against the defendant's action." Alliance for Hippocratic Medicine, 602 U.S. 367. The complaint does not allege that the defendants' conduct impaired any service the Association provides. It alleges nothing comparable to the conduct in Havens Realty Corp. v. Coleman, 455 U.S. 363 (1982), which, as the Supreme Court has since explained, "directly affected and interfered with" the plaintiff organization's "core business activities." Alliance for Hippocratic Medicine, 602 U.S. 367. The Association has no standing in its own right.

An association may also sue on behalf of its members when its members would otherwise have standing to sue in their own right, the interests it seeks to protect are germane to its purpose, and neither the claim asserted nor the relief requested requires the participation of individual members. Hunt v. Washington State Apple Advertising Commission, 432 U.S. 333 (1977). The second and third conditions are satisfied. The suit concerns the livelihood of Texas shrimpers, which is the purpose stated in the Association's bylaws, and the plaintiffs seek only declaratory and injunctive relief, which requires no member-by-member proof. The complaint also identifies a specific injured member, as Summers v. Earth Island Institute, 555 U.S. 488 (2009), requires. That member is Delgado. The Association's standing therefore depends on hers.

Limitations

Neither the app nor skill is perfect. Several limitations deserve to be stated plainly.

Jev sees only the text you upload. If your questions depend on definitions or on cross-referenced provisions, include them in the document. Jev chooses among the answers you supply; it cannot extract a name, quote a passage, or produce free text. For those tasks, give a language model the original documents along with Jev’s export.

The skill writes questions. It does not run Jev, set decision thresholds, or convert answers into conclusions. The grading and feedback in the student exercise were separate work performed after Jev ran, and the brief-checking editor described above does not yet exist. The timings reported here stop when Jev begins responding, and the cost figure covers Jev alone, not the language model used in preparation.

Test every question set against examples you already understand. In our justiciability testing, some errors were Jev’s misreadings. Others came from questions that mixed together different plaintiffs or different forms of relief, so that no single answer could be right. And some conditional questions produced answers that had to be disregarded because their premise turned out to be false on the facts at hand.

Finally, keep identifying information—student names, for instance—out of Jev. My earlier post explains my concerns about TypeSafe’s privacy terms.

We've only just begun

We are only beginning to discover what becomes possible when we combine the growing capabilities of now-traditional large language models with other forms of AI, such as Jev’s universal classifier. A language model can help design the inquiry, Jev can make thousands of focused judgments, and another model can turn those judgments into something a person can understand and use. That division of labor opens possibilities that are easy to miss when we expect one chatbot to do everything. And this is still Jev’s first generation. I am already seeing open-source variants such as Laya being built, and, if no less an authority than Reddit is to be believed, considerably more is coming. The tools will become faster, cheaper, and more capable; the interesting question is what we will think to do with them. As with large language models generally, the main limit increasingly looks like the user’s imagination: recognizing a useful task, asking what would make it possible, and trying the combination.

Notes

Three points about atomic questions

First, "atomic" describes how much work a question leaves undone, not how small it is. A question is atomic when it asks for a single judgment that someone who knows the field could make almost at once from the document and whatever the question supplies, such as a definition, a rule, or a list of permitted answers.

A question can be too large. "Was the defendant negligent?" hides several judgments: duty, standard of care, breach, causation, and harm. It also requires knowing law that Jev does not know. The skill therefore asks one question per element and writes the rule into each.

A question can also be cut too small. Asking whether a passage supports a proposition in a brief is one judgment, even though it compares two texts. If you split it into questions about whether each text uses the word "reliance," both can come back yes for a passage that says, "The court declined to hold that reliance is required." The keyword questions miss the relationship between the two texts.

The skill stops splitting when a question asks one thing and meets several further conditions, including:

  1. It can be answered from the document plus what the question supplies.
  2. It needs no intermediate step and no other question's answer.
  3. Its answers can be listed in advance.

Second, each atomic question must stand on its own. When Jev receives a batch of questions about a document, it answers each of them independently. The answer to question 7 is not available to Jev when it answers question 8. So if one judgment depends on another—if, say, a question about redressability matters only when the plaintiff has alleged an injury—that dependency must be handled either in how the questions are written or afterward, when the results are combined. Arithmetic, counting, and date comparisons should likewise be done by ordinary computation after Jev has answered, not asked of Jev itself. Jev does not yet come with skills or connectors to legal databases (Please, please add it in somehow in Version 2). And a table of probabilities does not turn itself into an overall score or a legal conclusion; someone has to write down an explicit rule for how the pieces combine.

Third, the sheer number of questions need not worry you. Jev evaluates all the questions in a request in parallel, so adding questions barely changes how long the response takes. That makes it practical to ask separately about distinctions we might otherwise be tempted to squeeze into a single ambiguous question. I have run sets of more than a hundred questions, and Jev handled them without complaint.