Research report • September 2026
Artificial Authority
A practical assessment of today’s leading AI models across pensions, tax, debt, savings and examining where their guidance helps, where it fails, and what responsible use requires.
Summary
Domain team tested 18 of the most popular AI models including ChatGPT, Gemini, Claude and CoPilot asking 121 questions about financial advice. On average, the models made mistakes 57% of the time. This research shows the calibre of advice given from concerning mistakes to consumer protection failure. It explores if a shift from excessive regulation to putting guardrails on LLMs will narrow the advice gap quicker while protecting people from unregulated AI?



