Sample the Documents Before Calling a Search Term Burdensome

In re TikTok, Inc., Minor Privacy Litigation, No. 2:25-ml-03144-GW (RAOx) (C.D. Cal. Aug. 19, 2026), full opinion (PDF)

The defendants in the TikTok minor privacy multidistrict litigation contend that their systems cannot de-duplicate or thread email. By their own account they have said so to the plaintiffs on numerous occasions. The same defendants resisted the plaintiffs' proposed search terms as burdensome, pointing at the hit counts those terms returned. Because they make that contention themselves, Magistrate Judge Rozella A. Oliver reasoned, they cannot rest a showing of burden on the count alone. A random sample would let them confirm that a term really is capturing a disproportionately high number of unresponsive documents, she stated.

The August 19 order resolves neither the search term dispute nor the custodian dispute. It sends the parties back to meet and confer. The defendants are told that a hit count their own platform inflates is not by itself a burden showing. The plaintiffs are told that a proposed custodian needs a reason beyond that person's knowledge of the subject.

What happened

The parties briefed their ESI custodian and search term disputes ahead of a discovery conference. In Judge Oliver's view the meet and confer efforts to that point had not been sufficient. She vacated the conference and directed the parties to confer again before the next one.

The court's analysis

Defendants "cannot point solely to a purportedly high hit number count to establish burden", Judge Oliver wrote, given their contention that their systems do not allow for de-duplication and threading. A count that includes duplicate and unthreaded documents is not sufficient to make the showing, in her view. A term may hit a high percentage of directly relevant documents while the raw count runs high because of the inability to de-duplicate and thread. In that situation, the court explained, it would likely find the plaintiffs entitled to those responsive and relevant documents.

The defendants had offered no compromise on determining whether the terms were hitting mostly responsive or non-responsive documents, the court noted. Judge Oliver directed the plaintiffs, for their part, to consider the compromises the defendants had flagged on connectors and proximity searches. The order gives the defendants a choice on Mandarin. They must either run the search terms in Mandarin for all custodians or identify the parts of the business that conducted their activities entirely in English.

"Generally the selection of ESI custodians is left to the producing party", the order states. A proposed custodian's knowledge of the subject is not enough on its own, the court explained, because another individual may have custody of overlapping responsive documents. For each proposed custodian, the plaintiffs should be prepared to explain why that person is likely to hold documents other custodians would not have, the order adds. On the other side, the court stated, the defendants should be prepared to say why specific agreed-upon custodians would have the information the plaintiffs seek. They "may not simply point to Plaintiffs' lack of knowledge, given that much of this knowledge is only within Defendants' possession", the court wrote.

In the court's view, at least some of the proposed custodians should be added, though not to the extent the plaintiffs requested. A list she wrote herself, Judge Oliver noted, was unlikely to match the custodians the parties would have agreed on, given their greater familiarity with the case.

Why it matters

A limitation in a party's review platform is a fact about that platform. It is not a fact about the requesting party's search terms. Counsel who has described the limitation to an opponent at length has already shown why the raw count overstates the review.

Sampling turns a suspicion into a number. A random sample review of the documents a term captures produces a responsiveness rate the other side has to engage with. It is available long before the dispute reaches a judge.

Custodian arguments run on asymmetric knowledge. That asymmetry cuts against the party holding the information. A requesting party still has to identify the documents only its proposed custodian is likely to hold. A producing party cannot simply answer that its opponent could not have known which documents to ask for.

A court deciding a custodian dispute works from the briefs alone. Parties who know which of their people hold which documents can build a better list than the court will build for them.

The full opinion is available as a PDF.

BlogeDiscovery Case in Focus