Research note 003

Why entity resolution matters when measuring law firms

AI systems rarely name a law firm the same way twice. Measurement depends on deciding which mentions refer to the same firm — and which do not.

Name variants

A single firm may appear under its full legal name, an abbreviation, a trading name, a former name, with “&” or “and”, with or without a suffix such as LLP or Lawyers, or as a variant the AI has generated itself. Counting text strings rather than entities splits one firm's visibility across several records.

Separate firms with similar names

The reverse problem is just as serious. Unrelated firms in different cities — or the same city — can share surnames or similar names. Merging them falsely inflates one firm's visibility and erases another's.

Preserving the original language

FirmRanker keeps the exact name the AI used alongside the canonical entity it is resolved to. Resolution happens in an analytical layer; the original observation is never rewritten. If a resolution is later found to be wrong, it can be corrected without losing the evidence.

Resolution states

  • Confirmed — Matched to a canonical firm with sufficient certainty.
  • Probable — Likely match, pending review.
  • Unresolved — Cannot yet be matched — including firms that may not exist.
  • Split — A record previously treated as one firm is separated into two.
  • Merge — Records previously treated separately are combined.
  • Human approval — Ambiguous cases are decided by a reviewer, not guessed.

What we still don't know

How often AI systems name firms that cannot be matched to any real entity, and whether that rate differs between systems, is an open question we intend to measure rather than assume.