On a recent episode of Stanford’s The Future of Everything, host Russell Altman sat down with Dan Ho, a professor of law, political science, and computer science at Stanford University, to talk about how AI is being deployed to tackle real problems in the legal system. Not theoretical ones. Not futuristic ones. Problems that have been festering for decades because no one had the tools — or the time — to solve them.

The conversation opens with a striking example. In 2021, California passed legislation requiring all 58 counties to comb through their property deed records and identify racially restrictive covenants — clauses that explicitly barred people of African, Japanese, Chinese, or “Mongolian” descent from owning or occupying property. Those covenants have been legally unenforceable since the mid-20th century, but they still sit in the records, embedded in the deeds that get pulled up every time a house changes hands.

The problem? Scale. Santa Clara County alone has 84 million pages of deed records going back to the 1800s. Before Ho’s team got involved, a team of two had manually reviewed roughly 90,000 pages to surface 400 covenants. Los Angeles contracted a vendor for $8 million and a seven-year timeline to do keyword searches across its own records.

Ho’s team built a fine-tuned AI model that processed five million deed records and identified the covenants in days. Along the way, they found that multimodal AI models outperformed — and cost less than — conventional optical character recognition tools for digitizing the old documents. The system is now being expanded to other California counties.

The second major project Ho describes is STARA, a statutory research assistant that can ingest entire legal codes and scan them systematically using large language models. His team partnered with the San Francisco City Attorney’s office to put STARA to work on the city’s municipal code — roughly 16 million words, comparable in size to the entire U.S. Code.

What they found was what Ho calls “regulatory sludge”: reporting requirements that may have served a purpose when they were enacted but have long since become dead weight. STARA identified approximately 528 reporting requirements. Among them: a mandate that the Director of Public Works regularly file reports on “fixed pedestal zones” for newspaper racks — concrete pedestals built to sell physical copies of the San Francisco Chronicle. An 80-year-old quarterly reporting requirement for a redevelopment agency. Reports on programs that no longer exist.

The Federal Reserve, each year, has to file a report on the Presidential Dollar Coin Program. And each year the Federal Reserve has filed this but also noted that the Presidential Dollar Coin Program ceased to exist in 2011. Could you please relieve us of this reporting obligation?

The City Attorney’s office ultimately proposed a 351-page resolution to delete or modify over a third of the identified obligations.

For all its promise, Ho is clear-eyed about AI’s limitations in legal practice. His research shows that general-purpose large language models hallucinate at alarming rates when faced with legal queries — between 60 and 80 per cent of the time on a benchmark of roughly 800,000 queries. Even retrieval-augmented generation systems, which pull in relevant source documents before generating answers, still hallucinate between one-fifth and one-third of the time.

And the failures aren’t random. They cluster in exactly the kinds of cases where people need the most help.

These general-purpose chatbots have a propensity to hallucinate at alarmingly high rates. And what we show is they hallucinate more frequently in exactly the types of cases that are likely to have underrepresented litigants.

Ho also flags a behavior he calls “model sycophancy” — the tendency of LLMs to affirm a user’s mistaken premise rather than push back on it. For people who don’t yet know the right question to ask, that’s a serious risk.

His recommendation: forget the dream of a general-purpose legal chatbot, at least for now. The real path forward is specialized, narrow-domain tools tailored to specific legal problems — systems that are easier to evaluate and far less likely to produce dangerous errors.

The open-versus-closed debate in AI gets airtime in the conversation, too. Ho points to a RAND study that found no statistically significant difference in bio-attack plans produced by teams with and without LLM access — and notes that the information LLMs provided was indistinguishable from what’s freely available on the internet. His takeaway: policy decisions about open and closed models need to be grounded in evidence of marginal risk, not assumptions.

What comes through most clearly in the episode is that AI’s most valuable legal applications right now aren’t about replacing lawyers. They’re about doing the work that nobody could do at scale — excavating 84 million pages of property records, scanning 16 million words of municipal code — and surfacing problems that have been hiding in plain sight for decades. That’s not a future promise. That’s happening now.

Shares: