AI and Legal Education News Brief: September 7, 2026
OpenAI releases Astra
OpenAI launched GPT-6 Astra on September 3, beginning a rollout to paying ChatGPT users and developers. It combines stronger reasoning with the ability to operate software and complete assignments involving multiple steps. For lawyers, Astra's most important capability may be better "agentic browsing." Basically, the language models take control of your browser and navigates websites – including those not designed to be AI-friendly – working as you would. We've had this for a while, including OpenAI's pioneering but now defunct ChatGPT Atlas browser, but all reports indicate that Astra is considerably more intelligent in browsing than predecessor technology. As discussed in a recent blog post, terms of service rather than technical issues may now be the leading impediment to broader and better use of AI in the legal profession and elsewhere.
Astra has had strong benchmark performance. OpenAI reports 97.6% on FrontierMath Tier 4 (v2), compared with GPT-5.6 Sol’s 83%; 96% on GPQA Diamond’s science questions; and 64.6% on Terminal-Bench Science, up from 22.4%. Astra does not win everything: Fable 5.1 leads it on Humanity’s Last Exam with tools, 65% to 57.2%. These are reported evaluation results, with settings that can differ from ordinary ChatGPT use. Still, those are numbers worth sitting with. OpenAI’s benchmark table supplies the comparisons.
As to performance on legal benchmarks, we don't have a lot of evidence yet on Astra. Harvey reports better handling of unsupported assumptions and gaps in documents. Legora’s agent reviewed 41 financial documents in minutes and found all four planted errors, including a £500,000 discrepancy. Legora reports nearly 40% improvement on that workflow, alongside a much smaller 3% average improvement across its broader legal benchmark. Both numbers belong in the story. Harvey’s assessment; Legora’s results. And for a good discussion of preliminary results from Astra, I commend Richard Tromans' discussion on his Artificial Lawyer website.
The importance of Astra goes beyond better benchmarks, however. It may well portend an advance in language model architecture. Astra probably uses a looped transformer, although OpenAI has not publicly confirmed that architecture. The idea originated back in March of 2024. Instead of performing "inference" (producing new tokens or words) by sending information across a neural network once, or increasing capabilities by just adding yet more layers to the transformer architecture underlying modern AI, loop transformers instead send the information through the same computational layers more than once. Reusing layers adds computational depth without requiring a separate set of stored weights for every pass. Doing this may save memory, but it costs computing time; the extra work is not free. Sebastian Raschka explains the architecture and the uncertainty surrounding Astra.
Even if Astra itself does not use looped transformers, this new architecture is now more than theoretical. There is a working example in Nanbeige4.2-3B, from the Chinese research team at BOSS Zhipin. Its technical report describes passing information through the same transformer stack twice. Nanbeige is evidence that we may be able to continue to make progress in AI without quite as large a need for ever more chips.
There's another bigger picture point here too, one that the release of Astra and Fable 5.1 brings into focus. My reading is that we are approaching intelligence that competes favorably with all but perhaps the top 0.01% of humanity in many areas. But ask yourself: if Astra performs at essentially superhuman levels on demanding STEM problems, why are we confident humans remain better at law? Is law harder? Does it uniquely require reasoning that nonhumans cannot perform? It is kind of weird that the best AIs score 9.9% on the ARC-AGI benchmark designed to test learning on the fly and 59.1% on "Humanity's Last Exam," but still only score 20% on the Harvey benchmark of legal tasks! Or are we lying to ourselves because the alternative threatens how we teach, bill, and understand our professional worth? Repeating the “judgment” mantra or repeating the need for a "human in the loop" does not mean that, on balance, humans are systematically better at "judgment" – and might someone try to define that term? – than is modern AI. Nor does it mean that inserting a human into a workflow is always the right answer. An interesting recent paper on this latter point is here.
I anticipate posting shortly about working with Astra, using two papers I have had under development as examples. Sneak preview: unbelievably good at math; not so good at audience-oriented writing on complex issues. It's almost too smart. And, beware, it chews up tokens.
Agentic token use surpasses chat use
In 2025, agentic AI was the future, and no one was quite sure what it meant. As of fall 2026, agentic AI is most definitely here, and its shape is becoming ever more clear. Proof: in June 2026, agentic use of AI surpassed direct human use of AI at least on ChatGPT. Here's the graph you need to study.

To produce this graphic, OpenAI analyzed aggregated, de-identified 10 million messages from its enterprise customers. By June 2026, Codex produced 64% of combined Codex and ChatGPT output tokens. Frontier firms (those in the top 10% of tokens per active user) generated 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January. Their employees also used plugins and skills more often: 21% versus 9% for plugins, and 19% versus 3% for skills.
The implications of this study for legal are, I think, pretty clear. Legal organizations should stop treating AI as just a drafting box. Agents can gather matter information, inspect files, update work product, run multi-step workflows, and return results for review. The professional skill shifts toward defining the task, supplying trusted context, setting authority boundaries, and checking the finished work. And, I suspect, today's "Agentic AI" is only the beginning. On the horizon, for example, is "multi-player AI" in which teams of humans (or, I suppose, other agents) create a single context from which the AI generates various forms of responses.
Legal education must follow. Teaching students to polish a prompt for a chatbot is now AI kindergarten. Professional students need to learn how to design and supervise agentic workflows: connect sources, choose tools and skills, state stopping rules, protect confidential material, build verification tests, and decide which actions require human judgment. Prompt engineering remains useful. It is no longer the course. My own course, Large Language Models for Lawyers, makes an effort at this more modern approach to AI use.
Non-lawyers, particularly the young, using AI to meet legal needs
JUSTICE and the Administrative Fairness Lab have published evidence that people already use AI for legal help. I'm not sure this is much of a surprise for readers of this blog, but their survey of 1,428 respondents in the UK reporting a legal dispute within two years showed that 233 said they had used AI for help, information, or advice. I suspect the current figures are higher now in the United States both because AI may be used more widely here and because the data was collected back in the ancient days of December 2025 through January 2026.
The study found that use was higher among 18-to-24-year-olds—26%, versus 10% among those aged 55–64. People used AI across consumer, housing, employment, family, medical-negligence, debt, insurance, and court-procedure matters. They sought explanations, drafting, negotiation scripts, second opinions on solicitors, and emotional reassurance. Only 6% of chatbot users relied on AI alone. They consulted 3.7 sources on average, versus 1.9 among others, although they were less likely to consult a solicitor.
The survey results provide additional support for the idea that major AI providers such as OpenAI and Anthropic should start treating law like they treat computation. Don't rely on the model alone; instead route the query. ChatGPT and peers often send arithmetic and code to a Python/code-interpreter tool rather than trusting weights alone. By default, legal questions still hit just the base model and do not use, for example, readily available connectors to data sources such as CourtListener. I would love to see one of the frontier labs buy out Descrybe, Midpage, Dingduff, or other services and make grounded AI the default. (I suspect some of those services would not object too strongly were such an acquisition made at the right price). Until then, we are going to unfortunately see AI spew hallucinations in addition to valuable legal advice when lay people consult it without attempting to ground the responses in a connector to primary legal authority.
A new skill to evaluate legal scholarship
I have made available a new skill called thesis-assessor that, particularly when used in conjunction with two other techniques, provides useful feedback on scholarly drafts or publications. The skill, based in part on Professor Eugene Volokh’s work on academic legal writing, has been reviewed and published by Lawve.ai and can run without a paid research connector.
The skill starts by turning a topic into an assessable claim. It identifies the proposed thesis, contribution, payoff, supporting method, and scope. It then asks whether the project fits a faculty article, student note, seminar paper, book, or early-stage comparison among topics. That matters because a semester-long seminar paper should not face the same demands as an article seeking to change a field.
The thesis-assessor skill next searches for the closest scholarship and distinguishes four possibilities: full preemption, partial preemption, adjacent work, and no preemption found in the sources searched. The final category is deliberately modest. A search can fail without proving novelty. The skill then rates novelty, nonobviousness, utility, and soundness; runs tests suited to doctrinal, normative, empirical, historical, comparative, or theoretical claims; and keeps only objections that survive the author’s best fair reply. Its report gives a green, yellow, or red light, but the useful part is the repair: a revised thesis, a sharper contribution, and a ranked research plan.
Two other skills can be used in conjunction with thesis-assessor. Professor David Stuckler’s free ResearchFastTrack package, currently for Claude, tests whether a proposed topic contains a research project worth undertaking. Its companion connector can search roughly 250 million paper records, test duplication, measure whether a claimed gap is filling rapidly, map a debate, profile journals, find neighboring papers, and inspect a paper’s citation path. Sustainable Opposing Counsel Review, which I adapted from Larissa Meredith-Flister’s Opposing Counsel Review, attacks a developed argument, drafts the author’s best replies, and deletes attacks that collapse under those replies. What remains is an adversarial critique one could defend.
The skills should be useful for faculty evaluating their own scholarship, for law students evaluating their own drafts, or for anyone seeking to critique legal scholarship. In my own writing seminars, I am requiring that students run their papers through the trio of skills. My hope is that it opens students' eyes to what is possible and induces them to improve their research and writing.
A new skill to help lawyers read technical papers
I have released reader-first technical-edit, an AI skill that helps lawyers and legal scholars work through technical articles by making their exposition clearer while checking that equations, numbers, quotations, and citations survive the rewrite. The eager can download a .zip file containing the skill here. I anticipate that it will also be available on Lawve.ai shortly.
The skill works with mathematical papers, empirical studies, and systems research. For legal audiences, that includes law-and-economics models, epidemiological studies, and evaluations of legal-AI tools. It produces an edited draft and an audit memo explaining what changed, what remains questionable, and which verification checks were completed.
The skill is important because (1) lawyers frequently have very practical reasons to confront technical material and (2) technical writing often rewards speed rather than clarity of exposition. I have found that models like Fable and Astra tend to produce technical expositions that are highly compressed because the models assume the reader is at the same level of expertise that the AI has developed over some sequence of interactions. For example, a toxic-tort lawyer needs to understand the study behind an expert’s causation opinion. A patent lawyer working on machine learning needs to follow how the model works. A criminal lawyer confronting facial recognition needs to ask whether the validation conditions resemble the conditions that produced the identification.
The skill addresses common issues with technical papers through nine editing passes:
- Consistent names: Give each study, sample, variable, and other object one recognizable name.
- Clear antecedents: Replace ambiguous uses of “it,” “this,” and similar references.
- Definitions before use: Explain terms, symbols, and acronyms when they first appear.
- Section signposts: State what each section does and why the reader needs it.
- Unpacked reasoning: Supply missing explanatory steps and check whether verbal claims match reported statistics.
- Jargon review: Explain specialized language for the intended audience.
- Identifiable actors: Make clear who selected, measured, calculated, or concluded.
- Drafting cleanup: Replace unexplained references to earlier runs or internal project history.
- Reader check: Read through the revision for comprehension, then repeat the preservation checks.
Formal papers also receive notation checks. Empirical papers get a map of studies, samples, selection rules, and limitations. “122 out of what?” should have an answer where the reader needs it.
I tested the skill on my own work, a friend’s work, and several papers a litigator might find valuable. I found that it performed well—the Seth benchmark, admittedly an informal testing program. These trials encourage me to share it; they do not establish an error rate. Full mechanical verification requires suitable tools, and author review remains necessary.