Can AI Make Justice More Just? | American Enterprise Institute

Every new class of law students faces a harder learning curve than the last because the legal system often produces inconsistent outcomes. When Congress and state legislatures enact new laws, lawyers have to play catch-up. These complexities create a system that is expensive and often inaccessible to everyday Americans. Artificial intelligence has the potential to address these problems by accelerating the development of legal expertise and making legal outcomes more predictable.

Shane is joined by Benjamin Alarie, a legal scholar on tax law and artificial intelligence. Benjamin is a leading voice in this debate and the coauthor, with Samuel Becher, of a new book, Superjustice: Law in the Age of Artificial Intelligence. He originated the concept of the “legal singularity,” the idea that advances in artificial intelligence could eventually make the outcomes of legal disputes almost perfectly predictable and perhaps even automate much of the practice of law. Ben is also the cofounder and CEO of Blue J, which uses the latest large language models and a vast database of authoritative tax content to deliver high-quality, source-driven answers to challenging tax questions.

Below is a lightly edited and abridged transcript of our discussion. You can listen to this and other episodes of Explain to Shane on AEI.org and subscribe via your preferred listening platform. If you enjoyed this episode, leave us a review and tell your friends and colleagues to tune in.

Shane Tews: Law often turns on interpretation, difficult distinctions, and corner cases. Could artificial intelligence make legal outcomes more consistent and predictable?

Benjamin Alarie: One of the things that I found as a law student that drove me really crazy was exploring the edge cases of the law and the difficult distinctions that students are asked to make, courts are asked to make, in talking about the leading cases.

Frequently, I would be totally puzzled by some dimension of a case that we were talking about in class. I’d approach the professor after the class and say, “Well, what if the facts were a little bit different?” And invariably, it seemed the professor would say, “Well, it all depends.”

Of course it all depends. But aren’t we supposed to be living under the rule of law? Aren’t we supposed to know what the law requires in different circumstances? It’s a thin answer masquerading as a deep answer. It shouldn’t depend because the law is noisy, because you have this variation from judge to judge, or you have different readings of the purpose of precedents. It should depend because the facts depend.

The fact that law is so dependent on interpretation of past case law and statutory interpretation makes it really, really difficult to work through. The Internal Revenue Code itself spans several volumes. And then never mind all of the case law and the interpretations of that voluminous statutory law.

I think AI can be a massive help in increasing the consistency of the analysis of all of those materials, moving us from a system where it really depends on the quality of the lawyer that you have, the quality of the representation that you have, and the identity of the judge you happen to have. Even where judges are engaging in their very best good-faith effort to get to the best answer they can produce, there’s just way too much variance in how the law comes out.

And 99.99 percent of life is not lived on those edge cases. It’s lived in day-to-day circumstances where a lot of people just are not able to access the legal system, to access the legal outcomes that they ought to be entitled to.

So it’s really interesting to think about what happens when you decrease by one or two or three orders of magnitude the cost of getting extremely crisp and clear insight into what the law requires. And that’s what my book Superjustice engages in.

One concern with legal AI is hallucination. With Blue J, your AI tax research platform, you’ve built what I’d call a “walled garden” around authoritative source material. How does that work?

I totally agree about the advantages of using authoritative documents to ground the work that you’re inviting an AI to do.

At Blue J, this is really, really important to the work of tax professionals. Tax professionals will not tolerate a hallucinated source, a made-up case, a made-up provision of the Internal Revenue Code. That is just not going to pass muster at all.

What we’ve done is create an enormous corpus of information. It includes all the legislation, federal tax law, state-level legislation, and a really significant trove of international tax materials from jurisdictions around the world. We’ve also licensed materials from leading nonprofit tax publishers.

When somebody comes into Blue J, it’s a lot like using one of these frontier models like ChatGPT or Claude or Gemini, but it’s specifically for tax professionals to get tax research answers. Because we built this proprietary walled garden of content, we’re able to find the relevant materials in this big corpus and draw upon these materials, analyze those materials, and produce an answer that’s responsive to the tax professional’s question.

A large language model will often simply make something up. A lot of what we label as hallucinations are gaps in the materials those models have been trained on. If there’s a gap, the large language models fill in gaps in the knowledge with things that seem plausible to them. We know that tax professionals just won’t accept that.

So the way our system is engineered, it’s designed to say, “Sorry, we don’t have an answer. No answer appears in the database for that question.”

Sometimes you also get situations where there are conflicting sources. Documents A, B, and C have one answer, and documents D, E, and F may offer a different view. What we’ve done with Blue J is for it to acknowledge the fact that there are conflicting sources.

We don’t want an AI system to just pick a side. We want an AI system to scrutinize A, B, and C, scrutinize D, E, and F, and then maybe say it seems like the stronger view, all things considered, is the ABC view or the DEF view, but to properly and fairly treat all of the materials as potentially relevant.

And once you get an answer, one of the things that is very important for tax people to do is to check the provenance of the sources. Blue J allows you to click on any one of the sources used to inform the answer and read the source text. You want to not just trust the output, but also be able to audit where it is coming from and whether these are valid sources of authority.

AI can now compress legal research that once took a junior associate days into minutes. What happens to the training that entry-level lawyers used to get from doing that work themselves?

This is the million-, billion-, trillion-dollar question. What do we actually do with the next generation of legal professionals, of tax professionals who are coming through?

I think one way of thinking about it is these tools can be extraordinary in accelerating the learning of newcomers to the field.

A nice example is major legislative change. The moment new legislation is passed, everybody is a newcomer to that legislation. Everybody is trying to figure out, “What does this mean for my clients? What does it mean for the Internal Revenue Code? And what are the new things that we should be doing as a result of this new legislation?”

After the passage of the One Big Beautiful Act, we saw an unprecedented spike in the use of Blue J. What we’ve seen is junior people, mid-level people, and senior people all leveraging AI in order to get up to speed on the new legislation. The way to get up to speed very quickly is to leverage AI in order to generate specific summaries, answer specific questions, and strategize around specific kinds of clients with specific kinds of issues.

In a lot of ways, junior people are able to leverage it even more quickly than senior folks. Part of me laments that these kinds of tools weren’t available when I was a law student, when I was first encountering how to read tax legislation and thinking about the policy implications of different choices and how to read the case law alongside the legislation. This would have been so powerful as a way to get up to speed and to challenge myself in my understanding.

It’s much like we see now younger and younger chess prodigies emerging. I don’t think it’s because we’re cleverer than we ever were before. I think it’s because they’re able to leverage extremely fast feedback and play properly calibrated artificial opponents using chess software to increase their playing capability. They can get more reps in and find the right kinds of opponents who are going to challenge them in just the right way to improve their chess understanding.

I like to think that the same thing can apply to individuals who are trying to learn the craft, whether it’s tax law or another super-technical area of law. You can leverage this kind of technology to accelerate your learning and development.

Could these tools also make the law itself clearer over time by helping lawyers and judges identify where the real gray areas are?

I think we will see more clarity over time. It’s about accelerating the convergence of law over time. I think it really simply does accelerate legal development.

If you go back 200 years and think about the development in the common law, it’s been grinding along. Whether you want to think about contract law or tort law or property law, it’s decision by decision by decision by decision.

The wheels of justice have been turning quite slowly over time. You get new refinements, you get some legal innovation, then you have appellate courts grapple with the implications and try to define the boundaries. And if you have a circuit split, then how do we resolve that? We have to wait for it to go to the Supreme Court and wait for the Supreme Court to decide how to resolve these splits.

But with AI, the whole system has more visibility into its own inner workings in almost real time. And so you can accelerate the learning. You don’t have to wait for the reporters to be updated on the shelves of the library, and then for somebody to have an issue and to go research it, read it, cite it in a submission, that’s all happening but it’s happening much more quickly than it ever has before.

To the extent that we can accelerate the understanding of the narrowness of the gray areas over time, more and more cases will settle. You will have a better understanding of just where the gray areas are.

The decisions that the courts are actually being asked to render are going to be increasingly ones where there’s a genuine normative issue or there’s a genuine factual uncertainty in play that the court is being asked to resolve. Because if it’s otherwise, then those cases are going to settle because nobody wants to invest in a losing case.

We will end up with a better understanding of just where the uncertainties are, and this is going to lead to more consistency in judgments.

I think one thing that seems very clear to me is that judges are increasingly going to be leveraging AI tools to work through what the existing law requires.

If judges are leveraging AI, even privately, even if they’re writing their own judgments, they still, as judges, are not abdicating their responsibility. They’re using these tools to help accelerate the production process. They still need to stand behind the reasons that they’re giving.

This is not so different from how judges in appellate courts are using law clerks today. Judges ask their law clerks to do the research, and it’s well known that judges often ask their law clerks to take a stab at writing a first draft of reasons.

I think it’s the same thing if a judge is using AI to help perform the same kind of legal research, take a crack at producing these reasons, but then the judge scrutinizes whether those words capture the judge’s thinking about the particular case and decides to own them and says, “Yes, I’m accountable for this judgment.”

If judges are leveraging AI, I think we’re going to see more consistency, more predictability in the work product coming out of courts, and that’s going to make the entire legal system more predictable and more capable of being relied upon.

As AI makes legal outcomes more predictable and plays a larger role in legal decisions, what still requires human judgment, and how should transparency and accountability work?

I think there’s something very valuable in having your day in court.

For many litigants who feel like potentially they’ve been wronged, they want to have a public forum in order to be able to air their grievances and say, “This is what happened to me. This is wrong. I want it to be on the public record that I was wronged.”

Probably they’re also seeking a remedy. But for a lot of litigants, they want this to be recognized and to be acknowledged and to at least be heard, even if they don’t get the remedy that they’re seeking.

There’s real value to that in many circumstances for a lot of litigants. If my bot and the other bot negotiate through and the other bot concedes that there was something wrong, I don’t know if it satisfies the human desire to be validated in a public forum in quite the same way. I think it doesn’t.

On transparency, people ask, “What are we going to do if the legal system is really dependent on these AI algorithms? How do we make the system transparent?” I think that’s a bit of a misunderstanding about what transparency should mean in a context of Superjustice.

What transparency should mean is: Do we have contestability? Do we have the ability to have an account explaining why this was the outcome? What’s actually driving the outcome? What would have had to have been different for the outcome of a particular situation legally to have been different?

It would be very nice for the systems to be able to explain, “These are the factual findings based on the evidence here. This is the narrative account of what happened. These are the facts as found by the system. This is the appropriate legal outcome. If the facts were otherwise, then the result would have been different because of these other reasons.”

The ability to appeal it to a higher court has got to be there too. It has to be both legally contestable, but also explainable and capable of being tested.

If we have systems that are capable of explaining themselves in that way, then I think we have all the transparency that we need in the legal system. It’s even more accountable than our current system of judicial decision-making. The parties cannot ask questions of the judgment and find out what would have had to have been different for the outcome to be different. But these AI systems are going to be able to carry on a conversation about precisely that kind of thing.

At Blue J, as a tax research platform, we lean extremely heavily on the professional responsibilities of our users. Blue J is meant for tax lawyers and tax accountants, people who themselves have a professional obligation to double-check the analysis of Blue J. They also have a duty to stand behind their work, and they are accountable for the guidance that they provide their clients.

So we do both things for tax people. We provide a system that is clearly faster and better at getting to tax research answers than the traditional paper-based or search-based methods. But then we also provide them with the background sources for them to satisfy themselves that that’s the right answer.

A very powerful move that I see a lot of the top users of Blue J make is to ask a question, get an answer, and then say something like, “Okay, now produce the very strongest analysis of why that would not be the case,” or, “Produce the strongest argument that the IRS may mount in resistance to this argument or this position that you’ve taken.”

Using the AI as a bit of an adversary, or building in some of this back and forth that the adversarial legal system invites already, is a tool for really understanding the boundaries of a position or of a case. I find that it really sharpens my own thinking.

If I had to pick a favorite prompt, it would be the adversarial one: “Tell me how this could be wrong.”

It drives me crazy to have a sycophantic AI sucking up to me. I really, really don’t like that. I want a tool that is going to challenge me, is going to be critical. But I respect it a lot more for having followed those instructions. That’s exactly what I want.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *