Using Differences in State Law to Test Whether LLMs Reason or Remember

Document Type

Article

Publication Date

9-13-2026

Abstract

When a large language model correctly applies a legal rule to facts, it's hard to tell whether the model has reasoned its way to the answer or has followed a context-specific pattern learned from its training data. To disentangle these two accounts, we exploit state-by-state legal variation in the United States, treating a legal rule's prevalence across states as a proxy for its prevalence in LLMs' training data. Legal rules can be placed along a spectrum: majority rules adopted by most states, minority rules adopted by few states, and hypothetical rules adopted by none. Drawing upon the legal topics of comparative fault, noncompete agreements, and hearsay, we evaluate LLMs' performance applying a majority rule, a minority rule, and a hypothetical rule of similar complexity. If a model's process of applying law to facts approximates legal reasoning, its accuracy should be consistent across all three rule conditions. But if a model relies upon context-specific patterns internalized during training, the model's performance should vary based on the prevalence of the rule within its training data. In two of three legal domains, we find that models perform best on majority rules, worst on minority rules, and at an intermediate level on hypothetical rules. In the remaining legal domain, models consistently and accurately apply all three rules. These results suggest that LLM performance at legal reasoning tasks is shaped, in part, by the prevalence of similar legal rules in models' training data. Future research can adopt this method of using state-by-state variation in legal rules as a principled way to distinguish legal reasoning from context-specific pattern matching.

Find on SSRN

Please note license information on SSRN. The file on SSRN may not be the final published version of this work. 

Share

COinS