Risk

Fake citations are now a costs order, not an embarrassment

By Michael Nadalin, Founder, Market Lead · 28 July 2026 · 6 min read

In August 2025 a senior Victorian lawyer stood up in the Supreme Court and apologised to a judge in a murder case. The submissions filed on behalf of his client contained quotes from a speech to state parliament that were never said, and citations to Supreme Court judgments that did not exist. The judge's associates had gone looking for the cases, could not find them, and asked for copies. There were none to give.

That story gets told as a cautionary tale about one lawyer. It is not. As of April 2026, researchers tracking this had documented 1,313 court proceedings worldwide in which AI-generated fabrications were put before a court or tribunal. In 496 of them, the person responsible was a licensed lawyer.

What has changed in Australia is not the frequency. It is the consequence. Courts have moved from raised eyebrows to costs orders. In one Federal Court matter, indemnity costs were ordered against the firm. That is the part worth paying attention to, because it reframes hallucination from a professional embarrassment into a quantifiable liability that sits on your practice, not on the software vendor.

The escalation nobody planned for

The trajectory here is predictable in hindsight. Courts extend goodwill to a genuinely new technology, then run out of patience once the same failure keeps arriving. Australia is now firmly past the goodwill stage.

How the response hardened
Stage 1
Judicial commentary. Courts note AI use and remind practitioners of their existing verification duties. No penalty.
Stage 2
Practice notes and disclosure requirements. Several Australian courts require parties to declare AI use in preparing material.
Stage 3
Named criticism on the record. The lawyer is identified in a published judgment, which is permanent and searchable.
Stage 4
Referral to the regulator. Conduct is escalated to the relevant legal services commissioner.
Stage 5
Indemnity costs against the firm. The financial consequence lands on the practice.
The direction of travel is one way. Each stage assumes practitioners were on notice from the stage before, which means the excuse available to you gets weaker over time, not stronger.

Why 'just check the citations' keeps failing

Every firm that has been burned says the same thing afterwards: we should have checked. And every firm already had a policy that said to check. The policy was not the missing piece.

The reason verification-after-the-fact fails is that it asks a human to disprove something that has been written to be believable. A fabricated citation from a general-purpose model is not obviously wrong. It has a plausible party name, a plausible year, a plausible court, and a plausible proposition attached to it. It looks exactly like the thirty real citations around it. There is no visual signal telling you which one to check, so checking properly means checking all of them, which costs more time than the AI saved.

So the rational thing happens. Under deadline, a junior spot-checks a sample. The sample comes back clean. The one fabrication in the document was not in the sample.

What firms believe vs what actually happens
The model will say when it is unsure
Fluency and confidence are the same in the output. A fabrication reads exactly as steadily as a real holding.
Fabrications look obviously wrong
They are built from real-looking components: real courts, plausible years, plausible party names.
We verify everything before filing
Under deadline, sampling replaces verification. A sample of 5 from 30 misses the one that matters 83% of the time.
The junior will catch it
The junior has the least context to know which propositions are unusual enough to warrant a look.
Our AI policy covers us
Only 21% of firms have a formal AI policy at all, and a policy is not a control. It allocates blame after the fact.

The exposure is bigger than the tool you approved

There is a second problem sitting underneath this one, and it is worse because it is invisible on your risk register.

According to Clio's 2025 Legal Trends Report, 46% of legal professionals are using generic, non-legal AI tools in their work, and 31% personally use generative AI, while only 21% of firms have any formal AI adoption policy. Read those numbers together and the picture is clear: in most firms, AI is already in the workflow, it arrived without approval, and nobody knows which matters it has touched.

That is not a training problem. It is a supply problem. If the sanctioned tool is slow, expensive, or gated behind a request form, people will use the one in the browser tab they already have open.

The adoption gap, Clio 2025 Legal Trends Report
Using generic non-legal AI tools
46%
Personally use generative AI
31%
Firms with a formal AI policy
21%
The gap between the first bar and the last is the shadow AI problem: usage is running roughly twice as far ahead of governance.

Confidentiality is the quieter and more expensive failure

Fabricated citations get the headlines because they are visible. The confidentiality failure is worse precisely because nothing happens when it goes wrong. There is no judgment, no costs order, no moment of discovery. You simply do not find out.

A US federal court in the Southern District of New York has already ruled that documents produced using a publicly available AI tool were not protected by attorney-client privilege or the work product doctrine, on the basis that using a consumer-grade platform compromised confidentiality. The reasoning is not exotic. Putting privileged material into a third-party consumer service introduces a third party. That is the whole mechanism of waiver.

The Australian position is not settled in the same terms, but the risk shape is identical, and no practitioner wants to be the test case that settles it.

Where client material actually ends up, worst to safest
Free consumer chatbot, personal account
Inputs may be retained and used for training. No firm control, no audit trail, no record of which matters were exposed.
Highest risk
Consumer chatbot, paid tier
Training is usually off by default, but retention, jurisdiction and access remain the vendor's terms, not yours.
High risk
General enterprise AI, no legal controls
Contractual protection improves. Nothing stops a lawyer pasting an unredacted brief into it.
Moderate
Legal AI with retrieval over real sources
Answers are grounded in actual legislation and case law, so fabrication risk drops sharply. Confidentiality still depends on the vendor.
Lower
Legal AI with anonymisation before the model
Client-identifying detail is stripped before anything leaves. The question still gets answered; the identifying material never travels.
Lowest
Heat shows exposure, not popularity. The top row is the most commonly used option in most firms, which is the problem.

What actually prevents this

The fix is not more diligence. Diligence is already maxed out; that is why sampling happens. The fix is changing the shape of the work so that verification is cheap instead of expensive.

Two structural changes do most of the work. First, ground the answer in retrieved source material rather than in a model's memory, so every proposition arrives already attached to the legislation or judgment it came from. A citation you can click is a citation you can check in four seconds. Second, strip client-identifying information before the query leaves your environment, so the confidentiality question stops depending on how carefully each person phrased their prompt.

A workflow where verification is cheap
1
Anonymise on the way in
Names, matter numbers and identifying facts are stripped before the query leaves your environment. Confidentiality stops depending on prompt discipline.
2
Retrieve, then answer
The system searches real Australian legislation and case law first, then answers from what it found. It is not recalling from training data.
3
Bind every claim to a source
Each proposition carries the provision or judgment behind it. Nothing appears without something to click.
4
Verify by reading, not by trusting
Checking means opening the source, not deciding whether the sentence feels right. This is the step that scales.
5
Keep the trail
What was asked, what was retrieved, what was relied on. This is what you need if the question is ever put to you.
Each step removes a category of failure rather than adding a review stage on top. Adding review stages is what firms try first, and it is what breaks under deadline.

The uncomfortable summary

AI is not the risk. Unsourced AI is the risk. A model that answers from memory will occasionally invent something, and the invention will be well-written, because writing well is the only thing it was optimised to do.

Australian courts have now priced that behaviour. The price is indemnity costs, a named judgment, and a regulator referral, in whatever order applies. Meanwhile, roughly half your profession is already using consumer tools that were never designed to carry privileged material, mostly without anyone having approved it.

The firms that get through this are not the ones that ban AI. Banning it just moves it into the browser tab you cannot see. They are the ones that make the safe path the fastest path.

Frequently asked questions

Can I be sanctioned in Australia for filing AI-generated fake citations?

Yes. Australian courts have moved past warnings. Practitioners have been named in published judgments, referred to regulators, and in at least one Federal Court matter indemnity costs were ordered against the firm. The duty to verify what you file has not changed because the drafting tool changed.

Does using ChatGPT for legal work waive privilege?

It creates a real risk. A US federal court in the Southern District of New York found that documents produced with a publicly available AI tool were not protected by attorney-client privilege or work product, because using a consumer platform compromised confidentiality. The Australian position is not settled in identical terms, but putting privileged material into a third-party consumer service introduces a third party, which is the mechanism of waiver.

How is retrieval-based legal AI different from a general chatbot?

A general chatbot answers from patterns learned in training, so a citation is generated text that looks like a citation. A retrieval-based legal tool searches actual legislation and case law first, then answers from the documents it found, and shows you those documents. The difference is whether there is a real source sitting behind the sentence.

What is PII anonymisation and why does it matter for legal AI?

It strips client-identifying information such as names, matter numbers and identifying facts from a query before it leaves your environment. The legal question still gets answered, because the law does not depend on who your client is, but the identifying material never travels to a third party.

Check your understanding

Two quick questions

No score is stored. Pick an answer to see why it is right.

  1. 1Why does 'we verify every citation before filing' keep failing in practice?

  2. 2According to Clio's 2025 Legal Trends Report, what share of firms have a formal AI adoption policy?

Research Australian law without handing over client data

Legal Brain searches Australian legislation and case law, shows you the source behind every answer, and anonymises client-identifying detail before anything reaches a model.

Request early access