From Key Numbers to Keywords: What the Search Box Cost Us

Part three of our four-part series on the impact of AI on the future of legal research.

In the first post of this series, we traced how the Abbott brothers and West Publishing organized American case law into digests, imposing a topical structure the profession relied on for a century. In the second post, we recounted how Simon Greenleaf, Frank Shepard, and King and Leonard built the citator, the tool that tells a lawyer whether a case remains good law. Both stories ended at the same place, the arrival of the computer.

This post covers what happened next. Search-based research made lawyers faster than any innovation before it. The profession's own scholars have also spent three decades documenting what it did to legal reasoning. Both halves of that story deserve telling, because together they explain why the next era of legal research looks the way it does.

Punch Cards in Pittsburgh

The search era began far from any legal publisher. At the University of Pittsburgh's Health Law Center, John Horty built a punch-card retrieval system for state health and hospital statutes, and at the American Bar Association's 1960 Annual Meeting he and a colleague demonstrated it, retrieving statutes by matching keywords. Full-text retrieval was computationally possible, and the profession noticed.

The Ohio State Bar Association noticed most of all. In 1964 it began exploring automated case law research, and after three years of study it contracted with Data Corporation to build the system. By 1969 the Ohio Bar Automated Research service was running on remote consoles in law firms, libraries, and government offices. The system malfunctioned often, cost more than projected, and met deep skepticism. It also worked. Mead Corporation acquired Data Corporation, and in 1973 Mead Data Central relaunched the service, soon a national one, under a new name, Lexis. Its signature idea was full-text searching, unconstrained by any human index, a controversial notion at the time. West Publishing answered in 1975 with a competing service, Westlaw.

The vendors then made a shrewd generational bet. They priced law school access at a small fraction of commercial rates, and students learned to research on screens while their senior partners never did. Robert Berring, whose work on "cognitive authority" we drew on in the first post, observed that this gap had a cost. The senior lawyers who traditionally served as the gatekeepers of research method were largely outside the transition, so a generation of young lawyers navigated it mostly on their own. The profession changed tools without ever quite deciding to.

The Liberation Was Real

It is worth pausing on what full-text search gave the profession, because the gains were real and they are permanent. For a century, finding a case meant thinking in the editors' categories. Full-text search dissolved that requirement. Any word in any opinion became an entry point. Research that consumed days in a library took minutes at a terminal, and a small firm with a subscription could reach the same corpus as a national firm with a grand library. Writing in the Law Library Journal, Carol Bast and Ransford Pyle asked whether the profession was living through a paradigm shift, and by any ordinary meaning of the phrase it was. No lawyer should want to go back.

Keywords Displaced Concepts

The cost arrived quietly, in the way lawyers began to think. Print-era research forced doctrine to come first. A lawyer could not look up a broken carriage wheel; she had to reason her way to negligence before the digest volume would yield anything. The search box inverted the sequence, and Barbara Bintliff put the inversion plainly in Context and Legal Research, published in the Law Library Journal in 2007. "Legal research no longer requires beginning with knowledge of the law," she wrote, "because the emphasis of electronic research is on facts and keywords, not legal concepts. Research now is truly a mechanical process of entering factual words into a database or search engine and retrieving results." F. Allan Hanson's title from the same journal tells the same story on its own: From Key Numbers to Keywords: How Automation Has Transformed the Law.

Something communal faded as well. When every lawyer, professor, and judge found authority through the same taxonomy, opposing counsel argued from a common map of the law. In Bintliff's words, West's structure "provided a shared context for legal research and analysis and, by extension, for the law itself." Electronic research replaced that shared context with individualized result sets. Each researcher, she observed, "creates an individualized context from the materials," and where researchers never reach the common core of rules, "there is no shared context and thus no communication." Two lawyers arguing the same motion may now arrive with cases that happen to share vocabulary but rest on different doctrinal footings, talking past each other in front of a judge holding a third set entirely. And the habit of hunting for a case on "all fours" with the client's facts trained researchers to prize factual coincidence over the rule and the reasoning, which is what precedent actually carries.

A Game of Go Fish

Keyword search has a mechanical weakness that every litigator knows by feel. Words are ambiguous, and the researcher must guess which ones a judge used decades ago. A search for "record" returns criminal records, vinyl records, safety records, and courts' own records. A controlling precedent that states the governing rule in different vocabulary is simply invisible. The whole exercise plays out like a game of Go Fish, and it produces an unforgiving trade. A narrow search misses relevant authority. A broad search returns hundreds of cases nobody has time to read. Because reading time is the scarcest resource in practice, lawyers choose narrow, accept the misses, and wade through the document dumps that come back anyway.

The Interface and the Algorithm Did the Steering

Two later Law Library Journal studies showed that the researcher was not entirely choosing this behavior. Julie Jones, in Not Just Key Numbers and Keywords Anymore, applied information-foraging theory to the databases' own interfaces. The simple search box sits in the most prominent position on the screen, while the indexes, tables of contents, and structured finding aids, the digest's conceptual scaffolding and the secondary sources that bridge into it, hide behind menus and extra clicks. Novice researchers get channeled toward quick keyword queries because the interface makes depth expensive and shallowness free.

Susan Nevelow Mart went further in The Algorithm as a Human Artifact, running the same searches across six research platforms and documenting how differently each one answered. Every platform ranks results by hidden, proprietary rules about proximity, stemming, and relevance. Lawyers who believe they are seeing an objective slice of the law are seeing a vendor's curated interpretation of it. The editorial authority of the digest era had returned, in other words, but this time invisibly, inside an algorithm no researcher could inspect.

Frederick Schauer and Virginia Wise added a final observation in the Journal of Legal Studies. As technology cut the cost of reaching nonlegal information even faster than the cost of reaching legal authority, courts cited more of it, absolutely and as a share of all citations, a drift the authors called the delegalization of law and a foreshadowing of what they described as "the decreased dominance of the traditional canon of legal information."

Faster Research, Thinner Reasoning

The second era of legal research kept its central promise. It made research fast, and it opened the full text of the law to anyone with a subscription. But the profession paid in kind. Reasoning drifted from doctrine toward word-matching. Depth stayed rationed by human reading time, so recall was sacrificed to keep the reading pile manageable. And the steering, once done openly by editors whose categories everyone could see, moved into interfaces and ranking algorithms that no one could.

The question the profession never got to ask was a simple one. What would legal research look like if the researcher could read everything a broad search returns, reason in legal concepts rather than keywords, and show exactly where every conclusion came from? The final post in this series answers that question.

← Back to Blog