Chapter Eleven

Evaluating AI Legal Research Tools

You evaluate AI legal research tools by moving from what a vendor says to what you can read for yourself. The evaluation starts with questions, moves to the finished work product a demonstration can show, and ends in your own use of the system on research questions whose answers you already know.

Two questions come before everything else. Ask whether the system reaches a database of primary law at all, and whether that law is curated and carries citator treatment data. A meeting can then show real evidence, because a vendor can open a completed memorandum with the research plan its system built, a revision it made to its own draft, the treatment record behind a cited authority, and the passage supporting each proposition.

The true measure comes from experiencing the system. An agentic research task runs for hours, which no meeting accommodates, so the judgment that decides the purchase is your own reading of the memoranda it delivers on your questions, in a trial and a pilot.

Three things have to be in place before an AI tool can do legal research at all. The first is a database of primary law the system can actually reach. The second is curation of that law, building the connections an agent follows to locate it. A citator is the third, supplying the treatment data that establishes whether an authority is still good law.

Take the database question literally. Ask which body of primary law the system reaches, by jurisdiction and by court level, and ask to see it. A language model's general knowledge is not a source of law. A system with no database of primary law behind it can produce fluent text about the law. It cannot conduct legal research, because it has no authority to read.

The consequence a buyer feels later is that nothing can be audited. A citation check asks whether a cited case exists in a database of primary law, which is a question that has no answer where no such database exists. The same is true of a quotation compared word for word against its source, since there is no source to compare it to. Every statement in the work product then rests on the model's own assertion that it is so, which is the condition Chapter 8's validation audit exists to end.

Curation is the second requirement because raw opinion text carries the law without its structure. A retrieval-augmented system can search raw text, pull passages from it, and quote them back with their sources, so pointing to a source does not itself require curation. What curation adds is the law's own structure, recorded as connections an agent can use: from a rule to the cases stating it, from a quotation to the opinion it came from, from a case to the cases treating it. Those connections are what let the agent locate the law bearing on an issue and then interrogate it, following one authority to the next the way a lawyer works a library.

Curation also strengthens validation. A curated passage was identified, extracted, and checked before any question was asked, so the audit compares the draft against that prevalidated content rather than against an extraction made at question time. The same editorial work builds the citator, with its direct history and treatment histories. Chapter 5 sets out what curating law for an agent involves. Legal research conducted without curated primary law is reckless. The recklessness sits in the practice rather than in the technology.

The citator is the third because good law is not a fact contained in the opinion being read. Whether a case still stands is a fact about every later decision that has treated it, accumulated across millions of opinions, so no amount of reading the case itself supplies it. Chapter 6 works through why that bookkeeping has to come from maintained citation data. Ask whether the system has a citator, whether the treatment data is the vendor's own or licensed, and how current it is.

These three requirements narrow the field, which is why they belong at the front of an evaluation rather than the end. A product can be well designed, pleasant to use, and genuinely useful for other legal work while failing the first question. The three are also easy to establish in a few minutes. They are expensive to discover late.

Public-domain collections of judicial opinions are a real advance in access to the law. They are not a substitute for curated primary law in an agentic system, because they carry the opinions without the editorial structure an agent retrieves against and without citator treatment data. Ask which collection a system reaches, then ask what that collection does not contain.

The connection method is what changed. A standard interface now lets an AI system reach an outside collection of opinions directly, the Model Context Protocol, or MCP, being the term a vendor is likely to use for it. A tool that had no law behind it last year can carry a connection to a large body of opinions this year. Vendors say so, accurately.

The opinions arrive through that connection. The editorial work that makes a body of law usable by an agent does not, which means no topical classification, no tagged rules of law or determinative facts, no quotations captured as objects with their sources recorded, and no treatment data. Each absence removes something specific. Without classification and tagged rules, retrieval by legal concept has nothing to match against. Without captured quotations and citations recorded as data, there are no connections to follow from one case to another. Without treatment data, there is no good-law check at all.

Pagination is the failing that does not show at first look. Many opinions in public collections carry no page numbers, because the version the collection holds never received the pagination of the official reporter. Court rules that require pinpoint citations ask for the page where the proposition sits. A corpus without page numbers cannot supply that citation, however faithful its text, so a memorandum built on it arrives citing whole cases in forums whose rules demand the page. Ask whether the collection carries the official reporter pagination for the courts your practice appears before.

Coverage and currency are separate questions. Ask which courts and which years the collection covers, how soon after issuance an opinion appears, and whether later corrections, withdrawals, and amended opinions propagate. Answers vary by collection and by court. A gap in a trial court's coverage is invisible in a memorandum that never mentions the case it left out.

The profession's reliance for commercial-grade legal research has rested on collections that are editorially maintained, so a new connection to an unmaintained collection does not change what that collection holds. The point is about contents rather than about anyone's good faith. It is checkable by asking the two questions above.

The fair conclusion is that these collections are excellent for what they are. They are a good way to read an opinion, a good way to confirm that a case exists, and a genuine public good. They are not the foundation for a research system whose work product goes out under a lawyer's signature.

An evaluation of AI legal research tools comes down to seven questions, covering agency, the source and preparation of the law, retrieval depth, good-law verification, validation, oversight, and confidentiality. Ask what the system does after its own first answer. Ask whether it reads what it retrieves or summarizes a ranking. Ask what is checked before the memorandum reaches the lawyer.

Line of inquiryThe question to askWhat the answer has to establish
AgencyWhat happens after the system's first answer?That the system acts again on its own draft, and that it can show you what it did
The lawWhere does the law come from, and how is it prepared?Which body of law the system works from, and how a passage inside it is located precisely enough to cite
RetrievalWho reads the results, the system or the lawyer?How much of what a search returns is actually read, and how the system reaches authority no query named
Good lawHow is an authority's treatment checked, and when?That treatment is checked before the authority appears, and that the check reaches the point of law relied on
ValidationWhat is checked before the lawyer sees the memorandum?Which checks are mechanical, which rest on a model's judgment, and what the system does when one fails
OversightCan the process be supervised?That the plan, the sources, and the checks are visible to the reviewing lawyer, and that a record of them survives
ConfidentialityWhat can the agent reach, and where does client data go?The scope of the security certification, the training-data term, the permission model, and the deletion terms, in writing

Every one of the seven ends in something the vendor can put on a screen, which is the test of a good evaluation question. None of them prescribes an architecture. Each can be answered with finished work product, once the threshold requirements above are satisfied.

The agency question is answered by a revision rather than by the word agentic. A system that plans an inquiry, notices that its own draft overstates a proposition, and narrows the statement to what its sources support has shown the reflection loop that separates an agent from a single prompt with a model behind it. Chapter 1 works through that distinction. Chapter 8 shows the same loop running inside the validation audit.

These questions also do double duty as the lawyer's own preparation. The ethics guidance Chapter 9 reviews asks for a reasonable understanding of the specific tool's capabilities and limitations, for the specific task at hand. An evaluation conducted to these questions produces that understanding. The notes from it are the record that it was acquired.

An AI legal research vendor should provide its validation methodology rather than its assurances. Ask how the system's accuracy is measured, on what question sets, how a misgrounded proposition is detected, and what happens when the system's own checks fail.

Published research has found grounding claims in this market overstated, so the burden of proof sits with the vendor. That published record is Chapter 7's subject. Its use in an evaluation is narrow. It entitles you to ask for methodology rather than assurances. The evidence that counts is the vendor's own measurement, on stated question sets, with stated failure modes.

Ask the failure question directly. A vendor that can describe where its system errs has told you something about everything else it says. A vendor with no failure modes to describe has also answered the question. The candid answer names the boundary. Validation of the kind Chapter 8 describes examines the propositions a memorandum makes, never the ones it omits, so a memorandum whose every statement passes the audit can still miss a controlling case. A vendor that volunteers that limit is describing an engineered system. A vendor claiming to have eliminated hallucination is describing an aspiration.

Then ask what happens on a failure. The answer worth hearing is a loop. A proposition flagged as unsupported returns to the drafting agent, which narrows the statement, restores the dropped qualifier, or removes the claim with its citation, after which the audit runs again on the corrected draft. A warning label presented to the lawyer is a different answer, one that moves the work back onto the reviewer.

The profession already has a working model for this discipline. In eDiscovery, statistical validation of a review and a documented record of the process have been presented to courts and opposing parties for years, a practice Servient's guide to agentic eDiscovery covers. What gets disclosed there is the validation protocol rather than the vendor's internal engineering, which is the same disclosure a research buyer should expect. Ask whether the validation audit report persists with the matter as documentation of the diligence applied, because a check whose result is not recorded cannot be shown to anyone later.

Ask an AI legal research vendor four things about client data and security, and take every answer in writing. Ask for the scope of its independent security certification rather than the fact of one. Ask where the contract states that client data is not used to train models. Ask how the agent's reach is bound to what the directing lawyer may already see. Ask the retention and deletion terms.

Scope is where a certification answers a different question than the one asked. A recognized independent certification covers a defined system boundary over a defined period, so the useful questions are which systems sat inside that boundary, what the report period was, and whether the AI processing itself was in scope. A certification covering the hosting environment while excluding the model layer is a real certification of something other than what a research buyer is asking about.

On training, the difference that matters is between a policy page and a contract term. A policy can be revised without notice to anyone. A contract term binds. Ask for the clause, and ask whether it reaches the vendor's own model providers as well as the vendor. Chapter 9 sets out why marketing assurances do not discharge the confidentiality duty on their own.

The permission model is testable in the demonstration. An agent that can reach every matter in the system creates a route for confidential information to move between matters that the firm's own access controls were built to prevent. The safeguard is binding the agent to the permissions of the lawyer who directed it, so its reach matches that lawyer's reach exactly.

Retention and deletion are the questions firms remember late. Ask what happens to the matter's research and to the validation audit reports when the engagement ends, and when the vendor relationship ends. The firm may need the audit record after both, as its own documentation of the diligence applied to work product it has already relied on, so deletion terms that sweep it away solve one problem and create another.

Ask which third parties see the data. A research system that calls a model provider outside the vendor's own infrastructure has extended the boundary, so the firm's diligence extends with it. The answer belongs in the same writing as the rest.

Five answers should stop an evaluation of an AI legal research tool. They are a claim that hallucination has been eliminated, a purchase process that never puts the system in your hands, a colored flag offered in place of the treating decisions, a validation claim with no stated failure mode, and a rip-and-replace ultimatum in either direction.

The answer you hearWhat it tells youWhat to ask next
Hallucination has been eliminatedThe claim is not engineering, since no system built on a language model carries itWhere does the system still err, and what happens when it does
Everything you need is in our demonstrationThe examples are rehearsed on the vendor's own question setHow do my lawyers run their own questions through the system before we decide
The system flags cases that are no longer good lawThe flag may be the whole of it, with no treating decision behind itOpen the treating decision, and show which point of law it addresses
Our accuracy is 99 percentA number with no question set, no method, and no failure mode is not a measurementOn what questions, measured how, and what did the failures look like
You will not need your current platformsA rip-and-replace framing puts risk on the firm that the technology does not requireWhat does a first year look like running both

None of the five is a reason to walk out of the room. Each is a reason to keep asking, because each one is an answer that sounds like a capability and carries no evidence. The pattern they share is a claim stated at a level of generality that cannot be checked, which is the level a careful evaluation is designed to leave behind.

One further answer belongs on the list, though it arrives as a silence. A vendor that cannot say who reads the authorities its system retrieves has not answered the retrieval question. The unstated answer is usually that a ranked list goes to the lawyer. Chapter 4 explains why a better search still leaves the reading where it always sat.

An AI legal research demonstration should show work product rather than slides. Ask to see the research plan the system built, a revision it made to its own draft, an authority it reached without a query, the treatment record behind a cited case, and the passage supporting a proposition the system stated in its own words. Decide what a passing answer looks like on each before the meeting.

A meeting can show all five, because all five live in completed work product. What no meeting shows is the system running your question end to end, since an agentic research task reads and validates for hours. The demonstration is where you inspect the evidence of runs already made. The verdict waits for your own use of the system, on your own questions, which the next answer takes up.

Each item comes with a thin version that can pass for the real one, so the checklist has to state what counts.

The askWhat countsThe thin version
The research planThe legal issues identified in the facts, with the lines of inquiry each one opens, produced before the searching startedA screen of search terms
A revisionA proposition the system narrowed, with the earlier statement and the corrected one both visibleA claim that the system checks its work
An authority reached without a queryA case the system arrived at by following a connection from another case, named as suchA second search with different terms
The treatment recordThe treating decisions themselves, reachable from the citation, with the point of law each one addressesA colored flag beside the citation
The supporting passageThe source language displayed beside the proposition it supports, reached in secondsA citation at the end of a paragraph

A case overruled on one point may remain good law on the point being cited, which is why the treatment record has to reach the point of law rather than the case. Chapter 6 works through that point-by-point tracking.

The authority reached without a query is the hardest of the five to stage. Retrieval that moves from a case to the decisions citing, quoting, following, or distinguishing it reaches law no search term would have produced, which is the capability Chapter 5 describes. Ask which authorities in the memorandum arrived that way, and ask what the system read after it found them.

A demonstration that closes on a list of cases to read has shown you a search feature with a newer model behind it, whatever category the vendor claims. The agentic claim is answered by a legal research memorandum, every legal proposition carrying a pinpoint citation to the passage that supports it, delivered already checked.

Deciding the pass criteria in advance is what keeps a polished demonstration from substituting for the evidence you came to see. Write down the questions you will later run yourself, the authority whose negative treatment should surface, and the proposition you will ask to trace. Bring the lawyers who will use the system. Bring the firm's law librarian if it has one, because the questions a librarian asks about coverage and currency are the ones a demonstration is least prepared for.

You test an AI legal research vendor's grounding claims by running your own questions through the system, which means asking for trial access rather than relying on a meeting. Choose a question whose answer you know well, plus one that turns on an authority carrying negative treatment. Then judge whether the answer is right, whether the treatment surfaced, and whether a proposition the analysis depends on traces to its support.

The reason the test runs on your own access is time. An agentic research task plans the inquiry, reads the authorities it uncovers, and validates its own draft, work that takes hours rather than minutes. Nothing about that fits between two agenda items, so no meeting can put your question through the system while you watch. Ask instead how your lawyers get trial access, run the questions there, and read the memoranda the way you would read an associate's.

A question the vendor chose has been run before. Its answer is known to work, which is why it was chosen. What your own question adds is your ability to grade the result. You know which case controls, which authority the answer should have reached, and how a competent memorandum on that question reads.

Prepare a third question where the practice allows it. Choose one whose controlling authority states the rule in words no researcher would type into a search box. Chapter 3 documents that vocabulary mismatch, which the retrieval techniques of Chapter 4 only partly answer. An answer that reaches the authority anyway has demonstrated retrieval by legal concept, plus the graph connections that lead from one case to another without a query. An answer that misses it has demonstrated the limits of the search box, whatever the interface looks like.

The trace test is the one part that needs no trial, because it works on any finished memorandum, including the one a demonstration opens. Pick a proposition the analysis depends on, one stated in the system's own words rather than by quotation, and ask where it comes from. The sound answer places the source passage beside the proposition, on screen, in the time it takes to read a sentence. A confident narrative with a list of citations at the end fails this test even when every citation is real, because nothing connects any statement to the language behind it. Chapter 8 explains why the pinpoint citation is what makes the trace possible.

Ask one further question about how the answer was composed. A system that assembles its answer only from the passages a retrieval step returned is bounded by that step. If those passages omit the controlling authority, the answer is wrong by construction, however capable the model. A system that reads each authority it uncovers in full is bounded by what it uncovered instead, a wider boundary and a visible one. Chapter 7 explains why that difference decides how often a system misstates what a case says. It shows up on any question whose controlling authority you can name yourself.

A vendor confident in its grounding will offer the access. Reluctance to put the system in a buyer's hands before the decision is itself an answer.

You test an AI legal research tool's jurisdictional coverage with authorities you already know. Name the courts your practice turns on, then use your trial access to run questions whose controlling authority sits in each of them. A system cannot research law it does not hold, so coverage is tested before anything else is credited.

Ask for the coverage list in writing before any trial begins. It should state which courts are included, from which years, and how statutes, regulations, and administrative decisions are handled. Then ask the currency questions separately. Ask how long after a decision issues it becomes available to the agent. Ask how often citator treatment is refreshed. Treatment data governs whether an authority is still good law, so a treatment record that lags by months undercuts every verification claim built on it.

Test the list with three or four questions whose controlling authority you can name yourself, one for each court level that matters to the practice. Read each memorandum for whether the authority appears at all, then for whether it appears as the controlling authority or as one entry among several. A system that reaches a close cousin of the controlling case, but not the case itself, has found the right doctrinal neighborhood without the authority in it. That result points at coverage rather than at retrieval.

What the system does when coverage runs thin is the more revealing test, so put a question to it that you know the law does not clearly answer. Read the result for whether the memorandum says so. Reported thinness is a system telling you the truth about its own footing. Chapter 9 explains why the professional duties make that behavior the buyer's concern. An answer assembled confidently out of loosely analogous authority has overstated its support. That overstatement is harder to catch than a gap, because nothing in it reads as missing.

You compare two AI legal research tools by putting the same research questions to both and reading the two memoranda side by side. Fix the question set before either trial begins, so neither vendor sets the test. Weight the criteria the practice actually turns on. Record the result for each system on the same sheet, because a comparison run on different questions is not a comparison.

The weighting belongs to the firm, decided in advance. A practice that lives in state trial courts weights jurisdictional coverage heavily, because a system that cannot reach the authority is not improved by anything else it does well. A practice whose research goes straight into filings weights validation heavily, since an unsupported proposition travels into a brief. A practice that researches novel questions weights the reading depth and the graph connections, which are what reach authority no query names. Deciding the weights in advance keeps the decision from following whichever system made the better first impression.

Compare the work product rather than the meetings. Two memoranda on the same question are the only artifacts a firm can actually set beside each other. They also carry what a meeting passes over, which is whether the controlling authority is there, whether the negative treatment surfaced, whether each proposition traces to its support, and whether the analysis reads like something a lawyer would sign.

What to compareHow to score it
The controlling authorityPresent, reached as a close cousin, or absent
Negative treatmentSurfaced with the treating decision, flagged only, or missed
A traced propositionSource passage in seconds, in minutes, or not traceable
Authority reached without a queryNamed and explained, claimed, or absent
Behavior on the thin questionReports the thinness, or answers confidently anyway
The vendor's account of its failuresSpecific and mechanical, general, or none offered

Resist the feature comparison. A longer capability list is not a better research system. The features that differ between two systems are rarely the ones that decide whether a memorandum is right. The six rows above are all outcomes, which is why they compare.

When the two systems finish close, the tie-breaker is what each vendor said about its own failures. That answer predicts what the relationship looks like after the contract is signed, at the moment something goes wrong and the firm needs a straight account of why.

No, you do not need to replace Westlaw or Lexis to use AI legal research. Agentic legal research is a different category of system, so adopting one does not displace the search-era platforms a firm already licenses. The two coexist from the first day. Each firm decides over time what it still needs from each.

The incumbent platforms, Westlaw and LexisNexis, remain what they have been, the tools every lawyer in the firm already knows, with their citators, their secondary sources, and their practice materials. The agent does work those platforms were never built to do. It conducts the research itself over curated primary law, delivers a validated legal research memorandum instead of a result list, and works from the facts and the record of the matter rather than from a query.

Coexistence is also the practical adoption posture. Nothing has to be migrated. No one has to be retrained off a familiar platform to start. A pilot that disappoints costs the firm the hours it put into the pilot and nothing more. Running both systems is how the comparison gets made at all, since the baseline a firm measures against is the research it does today.

The hours the firm spends inside a search platform change over time, not the subscription itself. That renewal question gets answered later, with usage data the firm will have collected by then. The adoption decision comes first. Nothing about it waits on the renewal. A vendor that frames the choice as a rip and replace, in either direction, is asking you to take on risk the technology does not require.

A law firm should pilot an AI legal research tool on real matters, with a defined group of lawyers, against a baseline recorded before the pilot starts. Choose research tasks the firm has already completed the old way. Run those same tasks through the agent. Compare the memoranda for completeness, for authorities the earlier research never reached, and for whether the quality holds across researchers of different experience.

Two vendor claims call for a direct test. The first is time. The claim is that research time on a matter compresses, because the memorandum arrives written and already checked instead of accumulating researcher hours. Measure the elapsed time and the lawyer's review together, since a compression that review consumes is not one. The second claim is consistency, that every matter's research starts from the same reading depth over the same curated law, which would turn research quality from a property of the individual researcher into a property of the system. Test that one by putting the same question to researchers of different experience. The cost side of the comparison runs on the firm's own numbers, in hours billed, hours written off, and days elapsed from question to memorandum, which Chapter 10 works through.

Design the comparison so it can surprise you. Have a partner read two memoranda on the same question without being told which one the agent produced. A blind read is the only version of this comparison that carries any weight, since a reader who knows which memorandum came from the software grades it against a different standard in both directions. Use enough matters to see a range, because a single impressive result establishes nothing about the next question.

Choose the pilot group for judgment rather than enthusiasm. Include the associates who do the research, a partner who will judge the work product against what the firm expects, the law librarian, and at least one lawyer who expects the tool to fail. Every memorandum in the pilot stays under a lawyer's review before any of it becomes advice, which is the supervision Chapter 9 sets out.

Keep the notes. The same file that supports the adoption decision documents the firm's supervision of the tool, the practical evidence that reasonable efforts were made before the firm relied on the technology.

An AI legal research pilot should run until it has produced enough work product to decide on, which is a milestone rather than a number of weeks. Set the end at the point where the pilot will have delivered three or four distinct research tasks across more than one matter, with more than one researcher involved. Decide that end before the pilot starts.

A calendar sets the wrong condition, because research arrives at points in a matter rather than on a schedule. A six-week pilot on a quiet docket produces two memoranda and no comparison, while the same six weeks across a motion cycle and a summary judgment produces more than the firm can review. Chapter 10 sets out the points where a litigation matter actually needs research. Those points are what a pilot has to cover.

Name the milestone in terms of what has to exist at the end. A memorandum on identified issues, a claims analysis at intake, and a research analysis of an opposing brief cover the range of work product. At least one question should be one where the law is thin, since how a system behaves without controlling authority is the finding a short pilot most often misses. At least two researchers of different experience should have run something, because consistency is one of the two claims under test.

The failure mode is the pilot that never ends. A pilot without a decision date becomes adoption by default, running on real matters without the written policy, the supervision protocol, or the decision that adoption is supposed to involve. Extend once if the milestone has not been reached, for a stated reason, in writing, with a new end.

A law firm should adopt AI legal research in stages, starting from the pilot's results and a written policy on permissible use. Begin with the research tasks that produce a defined work product. Keep every memorandum under a lawyer's review before it becomes advice. Expand toward the workflows where research connects to the rest of the matter once the review record supports the expansion.

The written policy comes first because the professional duties point there. Chapter 9 sets out those duties in full. The firm's managing lawyers establish the policy on permissible use. Training the lawyers and staff who will work with the tool follows from the same supervisory responsibilities. An adoption plan naming who reviews what, at which point, is the operational form of the oversight obligation, which grows with the system's autonomy. A firm that adopts without a policy is harder pressed to show it took reasonable efforts.

Sequence the work by the definiteness of the work product. A legal research memorandum on identified issues is the natural starting point, since the question is framed and the deliverable familiar. A claims analysis of a fact pattern at intake follows, then a research analysis of an opposing brief. Each step moves the research earlier in the matter, where it does more good. Each step also gives the firm another work product to judge before it commits to the next.

Revisit the known limits with the actual practice mix in hand. Jurisdictional coverage matters most where the practice lives in state trial courts or before an agency. Thin-authority questions call for a work product that reports the thinness. The validation audit Chapter 8 describes reaches the propositions a memorandum makes, so a lawyer's judgment about what the research should have addressed remains the part no system supplies.

A solo or small firm can use AI legal research. The evaluation changes more than the answer does. A firm without associates has no research hours to redeploy, so the case for adopting rests on reaching a depth of research that was previously out of reach. Jurisdictional coverage carries more weight, because a small practice is usually concentrated in a few courts.

The pilot design in this chapter assumes a group, though a solo runs a cleaner version of it. Compare the agent's memorandum against research the same lawyer did herself on a question she remembers. One reader graded both, which removes the variable a firm's blind read is built to control. Two or three such comparisons decide the question.

The supervision obligation does not shrink with the firm. It concentrates. No colleague reviews the work product, so the review of every memorandum before it becomes advice rests entirely on one lawyer. The managerial duties that occupy a firm's adoption plan matter less here. The personal duty of competence matters more. Chapter 9 covers what it asks.

Confidentiality diligence is the same work, harder to do alone. A solo has no technology function to read a certification report, which makes the plain-terms answers the ones to insist on. Ask what is in the scope of the certification, ask for the training-data clause, and ask what happens to the data when the subscription ends. The answers still belong in writing.

The depth argument lands hardest where the alternative was doing the reading yourself at night, on a matter whose budget never supported it. That is the honest case for a small practice, a case about the research rather than about the technology. What a research task costs, and how it is priced, is Chapter 10's subject.

A profession bound by duties of competence and supervision should adopt a consequential technology deliberately. The vendors worth working with will support exactly that, and will hand you the evidence to conduct the evaluation rather than ask you to accept the description.