Study · Validation
System Validation 001
A controlled Sydney public-liability dataset used to validate FirmRanker's research pipeline before larger benchmark studies.
- Study ID
- SV-001
- Market
- Sydney, NSW, Australia
- Practice area
- Personal Injury
- Prompt archetypes
- 10
- AI systems
- 4
- Observations
- 40
- Status
- Validation
- Methodology
- v1.0
Why this study exists
System Validation 001 is not designed to establish which Sydney law firm is “best” or to publish an AI visibility ranking.
Its purpose is to test whether FirmRanker can reliably preserve AI answers, identify firms, resolve name variants, classify response treatment, capture sources and support human review.
The study is therefore research infrastructure rather than a market benchmark.
Study design
10 unbranded prompt archetypes → × 4 AI platforms → × 1 repetition → = 40 observations
Model identity returned by each provider is recorded with every observation.
- GPT-5 with web search
- Gemini 2.5 Flash with Google Search
- Claude Sonnet with web search
- Perplexity Sonar Pro
The original answer is never rewritten
FirmRanker preserves for every observation:
- Exact prompt
- AI provider
- Returned model
- Timestamp
- Raw answer
- Raw source URLs
- Searches, where available
- Later corrections may change — Extraction, canonical entity, response treatment, position and source classification.
- They do not change — The original observation.
Human QA
Automated extraction → Frozen machine result → Assisted review → Final human review → Accuracy vs frozen output
Automated extraction produces the original machine result, which is frozen. Assisted review may propose corrections. Final human review creates a gold-standard comparison layer, and machine accuracy is measured against the frozen original output.
Current status
Status: Validation in progress.
Firm-level outputs remain preliminary until entity resolution and human QA are complete. No firm-level results, league tables or accuracy metrics are published from this dataset.
Limitations
- Only one market
- One practice area
- One repetition per prompt/model combination
- Machine extraction pending and subject to QA
- AI outputs may vary
- Provider behaviour may change
- Model versions change
- Source retrieval differs by provider
- Results should not be generalised beyond tested conditions