
A research assistant grounded in your own precedents and knowledge
A senior associate wants to know how the firm usually drafts a limitation of liability clause for a software supply agreement. The answer is in the firm somewhere, either in the precedent bank or in a file note a partner wrote three years ago. A general-purpose chatbot will write a confident clause from its training data. The associate needs the firm’s own clause and a link to the document it came from.
This post is part of our Practical AI for Law Firms series. It covers how we would build that assistant. The design uses retrieval-augmented generation (RAG) over the firm’s own material. Answers cite the source document, the assistant refuses when the corpus has nothing, and access control works at matter level so information cannot cross an ethical wall.
What grounding means in practice
A grounded assistant answers only from the documents it retrieved for the current question. Every claim in the answer points to a passage in a specific document, and the reader can open that document and check it.
The open-source ai-advisory-board project has a clean version of this pattern. A recent change added a Materials sidebar with versioned source citations, which ties each answer to the exact version of the document it used. Versioning matters for a precedent bank because precedents get revised after a judgment or a change in legislation. A citation to “the indemnity precedent” tells the reader much less than a citation to version 4, approved in March, which they can compare with the current version.
Refusal is the other half. If retrieval returns nothing relevant, the assistant says so and stops. It doesn’t fall back on general knowledge, because that is where invented authorities come from.
The failure mode grounding prevents
Courts in Australia and overseas have dealt with submissions that cited cases that don’t exist. General-purpose AI tools produced the citations and nobody checked them before filing. Several Australian courts have since issued practice notes or guidance on using generative AI in proceedings. Any firm planning an assistant should read the one for its jurisdiction before it designs anything.
These incidents have one cause. Someone asked a model for authority, and it produced text that looked like authority with no document behind it. A grounded design closes that path. The model sees only the retrieved passages. The system attaches citations from retrieval metadata, and a post-check drops any reference that doesn’t match a passage in the retrieved set.
The lawyer still reviews the answer before relying on it. That review is quick, because each citation opens the paragraph it came from.
An architecture sketch
The components are ordinary. The design work is in deciding where each control sits.
- Ingestion: connectors pull from the document management system (DMS), the precedent bank and the knowledge wiki. The connectors are read-only. The same PR adds an optional read-only local MCP reader that works the same way: the assistant can read documents but has no way to change them.
- Chunking and metadata: each document is split into passages. Every passage carries its document ID, version, practice group, matter number where there is one, and the access groups copied from the DMS.
- Index: a vector index combined with keyword search, hosted in an Australian cloud region. Legal text needs the keyword side because embeddings handle defined terms and clause numbers poorly.
- Retrieval: every query runs under the user’s own identity, and the index applies the access filter inside the search call.
- Generation: the model receives the retrieved passages. Its instructions are to answer only from them and to say plainly when they don’t cover the question.
- Citation check: a post-processing step confirms that every cited passage ID was in the retrieved set. It removes or flags anything that wasn’t.
- Audit log: every query is logged with the user and the passages returned, so the firm can reconstruct any answer later.
The agentic-graphrag project made a similar choice when it grounded its query generation in the engine’s resolved schema in place of a hardcoded description, with one bounded retry. The same change treats the caller’s base scope as an authorisation boundary. Filters on an individual call can narrow that scope but never widen it, and the system refuses requests outside it. Matter access needs exactly that rule.
In outline, the request path looks like this:
def answer(question, user):
scope = access_scope_for(user) # DMS groups plus the wall register
passages = index.search(
question,
filter=scope, # enforced by the index at query time
top_k=12,
)
relevant = [p for p in passages if p.score >= MIN_RELEVANCE]
if not relevant:
return Refusal("No matching precedent or know-how in the firm's library.")
draft = llm.generate(question, context=relevant, rules=GROUNDED_ONLY)
allowed = {p.id for p in relevant}
cited = [c for c in draft.citations if c.passage_id in allowed]
log_query(user, question, relevant)
return Answer(draft.text, citations=cited)
Ethical walls belong in the retrieval layer
The firm’s DMS already records who can open which matter, and the assistant has to inherit those permissions exactly. Some information barriers are recorded outside the DMS, such as a conflicts register entry that walls a team off from a former client’s files. The assistant has to honour those as well.
The filter has to run inside the index query. Suppose walled documents reach the model’s context and you try to strip them from the answer afterwards. The model can still paraphrase what it read. When the index applies the filter, walled material is never retrieved, so it has no route into an answer.
Some practical points from designing this:
- Sync permissions often. A wall put up this morning must apply to this afternoon’s queries. Permission changes should reach the index as events, with a scheduled full reconciliation as a backstop.
- Decide carefully what goes into the shared corpus. A precedent that has been cleaned of client detail and approved for firm-wide use belongs there. A raw advice on a matter file stays in that matter’s scope, even if it is the best example of the point.
- Keep refusals identical. The user should get the same message whether nothing exists or they lack access, so a refusal doesn’t reveal that a walled document exists.
- Test the walls. Write a set of questions that must return nothing for a walled user, and run it on every index rebuild.
Refusal is a written policy
Where you draw the refusal line is a choice, and teams do change it. The argus project, a financial research assistant, recorded a decision to reverse its earlier refusal of forward-looking questions. It now answers them as grounded scenarios built from cited inputs. For a law firm we would start stricter and answer only from the corpus. Practice groups can then loosen the rule for particular question types, with each change written down and approved by a named partner.
Where the documents live
Precedents and know-how are among a firm’s most sensitive assets. In September a widely discussed report alleged that OpenAI took mathematicians’ private research from their Codex chats. The allegation is unverified. It still shows why firms are wary of pasting confidential material into hosted AI tools.
For this assistant we would keep the index and document store in the firm’s own cloud tenancy, in an Australian region. The model endpoint should be under enterprise terms that exclude training on inputs, and those terms should be confirmed in writing before any client material is indexed. The firm owes clients the same confidentiality in the index as it does in the DMS.
Limitations and costs
- Corpus quality sets the ceiling. If the precedent bank holds conflicting versions of a clause with no status marked, the assistant will retrieve all of them. Most of the effort in a first project goes into marking which precedent is current and which are superseded.
- The assistant doesn’t research external law. Case law and legislation still come from the firm’s subscription services. The assistant reports what the firm has written and done before.
- Retrieval sometimes misses. If a question is phrased differently from the precedent, the search can fail to find it, and the assistant refuses even though the answer exists. Reviewing refusals regularly shows where the corpus or the search needs work.
- Model usage is usually the smaller cost. Ingestion, permission sync and evaluation take most of the effort, and that work continues after launch because the corpus keeps changing.
A sensible first project covers one practice group’s precedent bank under its existing permissions. That group’s lawyers write the test questions. The assistant then runs under supervision for a few weeks before anyone else gets access.
PicNet builds production AI systems for Australian organisations. If a research assistant grounded in your own precedents is on your list, talk to us about what a first project could look like.