Why legal AI demos look perfect and your first matter does not
The demo runs about forty minutes. Someone from the vendor asks the tool a question about unfair dismissal, or a retail lease point, or the test for unconscionable conduct. The answer comes back in twelve seconds, properly structured, with citations that check out when you click them. Everyone in the room nods and the practice manager books the trial.
Three weeks later a partner drops a real file on your desk. Nine hundred pages of scanned correspondence, a deed with two undated variations, an email export with the attachments stripped, and a client who has now given you three versions of the same phone call. You ask the tool the question you actually need answered, and the result is thin, or generic, or confidently wrong on the one point you happen to know cold. Nothing broke. The demo was simply a different kind of task.
This is not an argument that legal AI is useless. Plenty of Australian firms are getting real hours back from it. It is an argument that a demo is a controlled experiment and your matter is not, that the difference is structural rather than bad luck, and that you can test for it in about a fortnight before any money changes hands.
The demo question was chosen because it works
Vendors are not usually lying to you. They are doing what anyone does with a live audience: picking the question they have already run. That question tends to sit in a settled area of law, in a jurisdiction the product indexes well, on authorities that are heavily cited and therefore heavily represented in whatever corpus sits behind the tool. Ask about the elements of a cause of action that has been litigated for forty years and almost any competent system will look sharp.
The presenter has also done something you will not do at your desk. They typed a tidy, pre-edited factual summary into the box. No irrelevant background, no contradictory instructions, no client insisting the important thing is a comment made at a barbecue in 2019. The input was already the product of legal judgement, and that judgement is the expensive part.
The honest read of a good demo is narrow: this tool can retrieve and summarise well-trodden law when someone hands it a clean question. That is genuinely worth something. It is just not the same claim as it can handle my matters.
Your file is the variable the demo removed
The first real matter is where the input problem shows up. Australian practice runs on scanned annexures, photographed documents sent by clients, PDF bundles produced by the other side, and email threads where the operative sentence sits three replies down inside a quoted block. Optical character recognition on a faxed 1998 variation is not going to be clean, and a tool cannot reason about a clause it read as noise.
Then there is the shape of a real question. Demos answer what does the law say. Practice asks does this clause, on these facts, in this jurisdiction, give my client a position worth funding, and what does that turn on. The second question requires you to already know which facts matter, which is the judgement you were hoping to buy.
The result is a predictable first fortnight. The tool underdelivers on the thing you tested it for, and quietly overdelivers on things you did not think to try: summarising a long affidavit, producing a chronology from correspondence, first-drafting a letter of advice you then rewrite in your own voice.
The four places the gap opens, ranked by what they cost you
Not every failure carries the same risk. A clumsy draft costs you twenty minutes. A confected authority in a submission costs you considerably more, and several Australian courts have now issued practice notes and guidance requiring disclosure or verification of generative AI use in material filed with them. Your professional obligations do not shift because software produced the paragraph.
Sort the failure modes before you evaluate any product, because the mitigation for each one is different. Two of these are solved by process. One is solved by choosing a tool with the right corpus. One is not solved at all, and simply has to be checked every single time.
Run the trial on matters you have already closed
The single change that makes a trial informative is to stop testing on live work. On a live matter you cannot tell whether the tool was right, because you do not yet know the answer. On a matter you closed eighteen months ago you know the answer, you know which authority turned it, and you know which fact the other side eventually conceded. That gives you a marking key.
Pick matters that are ordinary rather than exotic. The point is to measure the tool against your actual mix of work, not against the hardest thing you have ever run. Include at least one where the answer was genuinely uncertain, because how a tool behaves when there is no clean answer tells you more than ten questions it can hit.
Two weeks of this is usually enough to make the decision. Four weeks is comfortable. Six months of a paid pilot that nobody owns is how firms end up renewing a product no one uses.
- Three to five closed matters across your real practice mix
- At least one where the outcome turned on an unusual or recent authority
- At least one with a genuinely messy document bundle, including scans
- Write down the correct answer for each before you touch the tool
- Ask the question the way you would ask a junior, not the way a demo phrases it
- Give it the raw bundle, not a summary you have tidied up
- Log every citation and open every one of them
- Note what it declined to answer, that behaviour matters
- Correct, incomplete, or wrong, on each matter
- Separate research failures from document reading failures
- Time the verification, that time is part of the cost
- Ask two other people in the firm to repeat one matter each
- Hours actually saved after verification, not hours theoretically saved
- Where client data is stored and who can access it
- Whether it fails loudly or fails quietly
- Whether anyone other than the enthusiast in the office would use it
Five questions that change what a demo tells you
You can turn a sales demo into a useful test by taking the wheel for ten minutes. Ask to type your own question. If the answer is that the environment is not set up for that today, you have learned something.
A straight answer to these sounds specific and slightly unflattering. A vendor who says the product handles everything is either not close to their own product or is hoping you will not check during the trial.
What good actually looks like once the shine comes off
A tool that earns its licence fee in an Australian practice usually looks unglamorous. It gets you a defensible starting point on a research question in ten minutes instead of ninety, it builds chronologies from correspondence without complaint, it drafts the boring two thirds of an advice, and it tells you plainly when it has nothing. You still read the authorities. You still form the view. The saving is in the approach work, not the judgement.
That is the standard worth holding vendors to, including us. Legal Brain is built for Australian practitioners and we would rather you ran the closed matter test above and found the edges yourself than took a polished demo at face value, because you will find them in month two regardless.
The first real matter going sideways is not a sign you chose badly. It is the first honest data you have had. Treat it as the beginning of the evaluation rather than the end of it, work out whether the failures were research, coverage or document reading, and decide from there. A tool that fails in a way you can predict and design around is worth considerably more than one that dazzles for forty minutes and then quietly hands you a citation nobody can find.
Frequently asked questions
Why does legal AI perform worse on my matters than it did in the demo?
Demos use a pre-selected question in a settled area of law, with a clean factual summary typed in by someone who already knows the answer. Your matter arrives as a messy bundle with contradictory instructions and a question that turns on specific facts. The model has not changed, the input and the difficulty have. Most of the drop off happens before the tool does any reasoning at all.
Can legal AI still invent case citations?
Yes, particularly when your question sits at or beyond the edge of what the tool has indexed. Systems that retrieve from a defined legal corpus and link to source reduce the risk considerably, but they do not remove your obligation to open and read every authority before it goes anywhere near a court or a client. Several Australian courts have issued guidance or practice notes on the use of generative AI in filed material, and the responsibility sits with the practitioner who signs.
How long should a legal AI trial run before I decide?
Two to four weeks is normally enough if you test properly. Use three to five closed matters where you already know the correct outcome, run them cold, and record where the tool was right, incomplete or wrong. Open pilots with no defined test set and no owner tend to drift into renewal without anyone ever forming a view.
What is the single most useful thing to ask in a legal AI demo?
Ask to type your own question, using a real point from your practice area, and then open every citation it produces in front of the presenter. If that is not possible in the session, ask them to demonstrate the product failing on a question they know it handles badly. How a tool behaves when it has nothing tells you more than a dozen questions it answers well.
Two quick questions
No score is stored. Pick an answer to see why it is right.
-
1A legal AI demo answers a research question flawlessly in twelve seconds. What is the most useful thing to ask next?
The demo question was chosen because it works, so repeating it tells you nothing new. Typing your own question, ideally something narrow, recent or jurisdiction specific, tests the edge of the corpus, which is where invention and coverage gaps appear. Opening the citations in the room turns a claim into evidence.
-
2Which file is the best test case for a legal AI trial?
A closed matter gives you a marking key. You already know the answer, the authority that decided it and the facts that mattered, so you can score the tool as correct, incomplete or wrong. A live matter cannot be marked because you do not yet know the answer, and a vendor hypothetical was built to be answered.
Research Australian law without handing over client data
Legal Brain searches Australian legislation and case law, shows you the source behind every answer, and anonymises client-identifying detail before anything reaches a model.