<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>PicNet Blog</title>
<link>https://picnet.com.au/blog/</link>
<atom:link href="https://picnet.com.au/rss.xml" rel="self" type="application/rss+xml"/>
<description>News and technical articles from PicNet, a Sydney-based software, AI and cloud engineering firm serving Australian organisations since 2001.</description>
<language>en-AU</language>
<item><title>From runbooks to agents: automating ops toil without automating outages</title><link>https://picnet.com.au/blog/from-runbooks-to-agents-automating-ops-toil-without-automating-outages/</link><guid isPermaLink="true">https://picnet.com.au/blog/from-runbooks-to-agents-automating-ops-toil-without-automating-outages/</guid><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><description>How we turn ops runbooks into scheduled automated tasks safely: dry-run first, idempotent steps, deterministic edits, anomaly alerts and staged rollouts.</description><category>ai</category><category>devops</category><category>runbookautomation</category><category>aiagents</category><content:encoded><![CDATA[<p>Most IT teams keep a folder of runbooks that someone works through by hand: purge stale records, rotate logs, fix the config value that keeps drifting, restart the service that leaks memory. Nobody enjoys this work, and in 2026 the obvious move is to hand it to a scheduled job or an AI agent. The risk is easy to miss. A careful engineer doing a task at 10am becomes a script doing it at 3am with nobody watching, and one bad assumption can become an outage or an empty table.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/practical-ai-in-devops-the-series/">Practical AI in DevOps</a> series. It covers how we automate this kind of toil on our own platform and where we keep a person in the loop.</p>
<h2 id="pick-runbooks-that-are-boring-and-reversible">Pick runbooks that are boring and reversible</h2>
<p>Good first candidates run often, have a clear definition of done and can be undone. Purging expired records qualifies if you have backups and a retention window. Changing firewall rules on a production edge is a poor first project. If a runbook step says “use judgement”, it needs either a person or a much tighter written rule before it goes on a schedule.</p>
<h2 id="dry-run-first-and-leave-it-there-for-a-while">Dry-run first, and leave it there for a while</h2>
<p>Every automated task we build starts with a dry-run mode. In that mode it does all the reading and deciding but none of the writing, and it logs what it would have done in enough detail that an engineer can check each decision.</p>
<p>The clearest example on our own platform is an automated data cleanup. It ran in dry-run for weeks before we trusted it to delete anything. Over that period we read its output and compared each proposed deletion with what we would have removed by hand. We switched it to live only after those lists matched run after run.</p>
<p>Weeks can feel slow for a small task. Keeping a dry-run going costs very little, while a wrong deletion means restore work and an uncomfortable conversation with whoever owns the data. The shape we use looks like this:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> run_cleanup</span><span style="color:#E1E4E8">(dry_run: </span><span style="color:#79B8FF">bool</span><span style="color:#F97583"> =</span><span style="color:#79B8FF"> True</span><span style="color:#E1E4E8">, max_deletes: </span><span style="color:#79B8FF">int</span><span style="color:#F97583"> =</span><span style="color:#79B8FF"> 500</span><span style="color:#E1E4E8">) -&gt; </span><span style="color:#79B8FF">None</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">    candidates </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> find_expired_records()</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#79B8FF"> len</span><span style="color:#E1E4E8">(candidates) </span><span style="color:#F97583">&gt;</span><span style="color:#E1E4E8"> max_deletes:</span></span>
<span class="line"><span style="color:#E1E4E8">        alert(</span><span style="color:#F97583">f</span><span style="color:#9ECBFF">"cleanup wants </span><span style="color:#79B8FF">{len</span><span style="color:#E1E4E8">(candidates)</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF"> deletes, limit is </span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">max_deletes</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">"</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#F97583">        return</span></span>
<span class="line"><span style="color:#F97583">    for</span><span style="color:#E1E4E8"> record </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> candidates:</span></span>
<span class="line"><span style="color:#F97583">        if</span><span style="color:#E1E4E8"> dry_run:</span></span>
<span class="line"><span style="color:#E1E4E8">            log.info(</span><span style="color:#9ECBFF">"would delete </span><span style="color:#79B8FF">%s</span><span style="color:#9ECBFF"> (expired </span><span style="color:#79B8FF">%s</span><span style="color:#9ECBFF">)"</span><span style="color:#E1E4E8">, record.id, record.expired_at)</span></span>
<span class="line"><span style="color:#F97583">        else</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">            delete_if_exists(record.id)</span></span></code></pre>
<p>Dry-run is the default. Going live needs an explicit flag in the scheduler config, so the switch is a reviewed change with a name against it.</p>
<h2 id="make-every-step-safe-to-run-twice">Make every step safe to run twice</h2>
<p>Schedulers retry, runs overlap and jobs die halfway through. Write each step so that running it a second time does no harm. In practice that means “set the timeout to 30” instead of “add 10 to the timeout”, “create the folder if it is missing”, and “delete this record if it still exists”. Take a lock at the start so two runs can’t work on the same data at once, and give each run an ID that appears in every log line.</p>
<p>When a run fails partway, the next run should finish the job without anyone tidying up first. If you can’t make a step idempotent, leave it in the manual runbook.</p>
<h2 id="prefer-string-edits-to-model-rewrites">Prefer string edits to model rewrites</h2>
<p>Language models are useful in ops for summarising logs, sorting alerts and proposing a fix. They are a poor tool for applying the fix. Ask a model to change one timeout in a 300-line YAML file and you get back a whole file that looks right. A comment may be missing, keys may have moved, and somewhere a value you never mentioned may have changed.</p>
<p>A deterministic edit avoids that. Find this exact string, replace it with that one, and fail if the string doesn’t appear exactly once. The diff is a single line and a reviewer can check it in seconds. When we use a model to work out the change, we have it output the old and new strings, and plain code applies them.</p>
<p>The same idea is spreading through agent tooling. Tesseracted Labs wrote about <a href="https://tesseracted-labs-blog.vercel.app/enforcing-coding-agent-guardrails-in-the-runtime-instead-of-the-prompt">moving coding-agent guardrails out of prompts and into runtime hooks</a>, arguing that an invariant shouldn’t live inside the probabilistic system it is meant to constrain. Their example of two agents with GitHub access approving each other’s pull requests shows how a prompt-level rule fails without anyone noticing. For SQL, the open-source <a href="https://github.com/idk-arsh/schema-guard">schema-guard</a> checks agent-written queries against a snapshot of the real schema before they run, which catches invented column names before they reach a database.</p>
<h2 id="alert-when-the-numbers-look-wrong">Alert when the numbers look wrong</h2>
<p>Give every automated task an expected range: how many rows it normally deletes, how many services it normally restarts, how long a run takes. When a run falls outside that range, the task stops and raises an alert, and a person decides what happens next.</p>
<p>The <code>max_deletes</code> check in the snippet above does this. If a cleanup that usually removes a few hundred rows suddenly wants to remove most of a table, something upstream has probably broken, such as a status field that was renamed so that no record looks active any more. Pushing on through is the wrong response. The same goes for agents watching monitoring data: an anomaly opens a ticket or pages someone. We allow automatic remediation only for actions that have already been through the dry-run and staging steps.</p>
<h2 id="give-each-task-only-the-access-it-needs">Give each task only the access it needs</h2>
<p>Platform vendors now treat broad agent permissions as a real risk. In early October, Apple <a href="https://techcrunch.com/2026/10/02/apple-says-its-tightening-macos-full-disk-access-controls-due-to-new-risks-from-ai-agents/">tightened macOS Full Disk Access controls</a>, and <a href="https://arstechnica.com/security/2026/10/apple-changes-full-disk-access-permissions-to-curb-abuse-from-ai-agents/">Ars Technica reported</a> that Apple’s stated reason was the growing risk of that level of access as AI agents become more capable and autonomous.</p>
<p>For scheduled ops tasks, the cleanup job’s database account can delete from the one table it cleans and can’t drop anything. Each task runs under its own identity, so the audit log shows which task did what. Where an agent proposes shell commands, something deterministic has to sit between the proposal and the server. The <a href="https://github.com/misqe/zero-trust-llm">zero-trust-llm</a> project shows one pattern: middleware intercepts the command, checks it against a read-only allow-list, and asks a human for a yes or no before anything else runs. AWS has also released <a href="https://strandsagents.com/blog/strands-box-the-big-picture/">Strands Box</a>, an open-source local agent sandbox in developer preview, with per-tool policies such as “allow git push only if tests are passing”.</p>
<h2 id="roll-out-in-stages">Roll out in stages</h2>
<p>Once a task has earned its way out of dry-run, it goes live in a test environment, then on one server or tenant, then everywhere. We start with a low per-run cap and raise it as the run history builds up. The manual runbook stays current until the automation has a stable record, and the dry-run flag stays in place so that switching back takes one config change.</p>
<h2 id="budget-for-drift">Budget for drift</h2>
<p>Automations that work on day one break later because the systems around them change. One automation developer on Reddit <a href="https://www.reddit.com/r/B2BForHire/comments/1wk23x3/for_hire_ai_automation_developer_i_build_ai/">described “maintenance drift”</a> as the cost they had underestimated on every job, giving the example of a client renaming a HubSpot property or moving an intake form six weeks after delivery.</p>
<p>Build assertions into each task that check its assumptions before it acts. Confirm the columns it reads still exist and the API still returns the fields it expects. When an assertion fails, the task stops and alerts, through the same anomaly path as above.</p>
<h2 id="what-it-costs">What it costs</h2>
<p>This approach takes more code than a straight script, and the dry-run period pushes back the payoff by weeks. It also has a running cost, because someone has to read alerts and update assertions when the systems around a task change. For a runbook that runs once a quarter, that overhead can be more than the toil it replaces, and we leave those runbooks manual. For daily and weekly chores, the trade has paid off for us.</p>
<p>PicNet builds production AI systems for Australian organisations. If you have a folder of runbooks you would like to stop doing by hand, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/from-runbooks-to-agents-automating-ops-toil-without-automating-outages/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Matter intake, conflict checks and entity resolution</title><link>https://picnet.com.au/blog/matter-intake-conflict-checks-and-entity-resolution/</link><guid isPermaLink="true">https://picnet.com.au/blog/matter-intake-conflict-checks-and-entity-resolution/</guid><pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate><description>Automating legal matter intake and conflict checks: LLM extraction, deterministic entity matching, AI help for hard cases, and a person clearing every hit.</description><category>ai</category><category>legal</category><category>matterintake</category><category>conflictchecks</category><content:encoded><![CDATA[<p>New enquiries reach a law firm as emails, web forms and phone notes typed up by reception. Before anyone opens a matter, someone has to work out who every party is and whether the firm has acted for or against any of them. In most firms a person does this by reading the enquiry, typing names into the practice-management system’s conflict search and scanning the results. It is slow work, and how good the search is depends on how well that person guesses the ways a name might have been recorded.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/practical-ai-for-law-firms-the-series/">Practical AI for Law Firms</a> series. It sets out how we would build an intake pipeline that turns unstructured enquiries into structured party data and resolves those parties against the firm’s records. Plain matching rules do most of the work. A language model helps with the cases the rules can’t settle, and a person clears every potential conflict before a matter opens.</p>
<h2 id="why-manual-conflict-searches-miss-things">Why manual conflict searches miss things</h2>
<p>A conflict search is only as good as the names you search for, and the same party can sit in a practice-management system several ways. “Harbour Fresh Foods Pty Ltd” might have been entered as “Harbour Fresh Foods Pty. Limited”, or as “Harbour Fresh” under a trading name, or under a former company name. Individuals show up as “Jonathan Smith” on one matter and “Jon Smith” on another.</p>
<p>Related entities make it harder. The other side in a lease dispute may be a subsidiary of a company the firm acts for, or its director may be a current client in a separate matter. Trusts are the hardest case. A trust is not a legal person, so the party is the trustee, and an enquiry that names “the Kellett Family Trust” may correspond to a record for “Kellett Nominees Pty Ltd ATF Kellett Family Trust”. A free-text search on “Kellett” either misses the record or returns hundreds of hits that nobody reads closely.</p>
<h2 id="the-pipeline">The pipeline</h2>
<p>This is the architecture we use as a starting point:</p>
<ul>
<li>Capture: enquiries from the intake mailbox and web forms land in a queue, with the original text stored unchanged.</li>
<li>Extraction: an LLM converts the enquiry into a fixed JSON schema of parties, roles and matter details.</li>
<li>Normalisation and matching: deterministic code cleans each name and matches it against an index of every party in the practice-management system, open and closed matters alike.</li>
<li>LLM review of the grey zone: candidates the rules score as uncertain get a written assessment from the model.</li>
<li>Conflict queue: a person reviews every candidate and records a decision.</li>
<li>Write-back: the matter opens in the practice-management system only after every candidate has a recorded decision.</li>
</ul>
<h2 id="extraction-with-evidence">Extraction with evidence</h2>
<p>The extraction step asks the model for structured output against a schema and nothing else. Every party must carry the text it came from, so the reviewer can check the model’s reading against the enquiry.</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="json"><code><span class="line"><span style="color:#E1E4E8">{</span></span>
<span class="line"><span style="color:#79B8FF">  "matter_type"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"commercial lease dispute"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "parties"</span><span style="color:#E1E4E8">: [</span></span>
<span class="line"><span style="color:#E1E4E8">    {</span></span>
<span class="line"><span style="color:#79B8FF">      "name"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"Harbour Fresh Foods Pty Ltd"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "entity_type"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"company"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "role"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"prospective_client"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "acn"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">null</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "evidence"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"I'm a director of Harbour Fresh Foods Pty Ltd"</span></span>
<span class="line"><span style="color:#E1E4E8">    },</span></span>
<span class="line"><span style="color:#E1E4E8">    {</span></span>
<span class="line"><span style="color:#79B8FF">      "name"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"Kellett Nominees Pty Ltd ATF Kellett Family Trust"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "entity_type"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"trustee_company"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "role"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"other_side"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "acn"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">null</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">      "evidence"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"our landlord, Kellett Nominees as trustee for the Kellett Family Trust"</span></span>
<span class="line"><span style="color:#E1E4E8">    }</span></span>
<span class="line"><span style="color:#E1E4E8">  ]</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span></code></pre>
<p>Missing fields stay null. The prompt tells the model never to supply an ABN or ACN that isn’t in the text, and code checks any number it does return against the published check-digit algorithms. A party whose role the model can’t determine goes to the reviewer with that role marked unknown, because an unclear role is still a party that needs a conflict search.</p>
<h2 id="deterministic-matching-first">Deterministic matching first</h2>
<p>Most of the matching is ordinary string and identifier work, and it should be. Rules are cheap, they give the same answer every time, and a reviewer or auditor can see why a match was made.</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#79B8FF">SUFFIXES</span><span style="color:#F97583"> =</span><span style="color:#E1E4E8"> {</span></span>
<span class="line"><span style="color:#9ECBFF">    "proprietary limited"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"pty ltd"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#9ECBFF">    "pty limited"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"pty ltd"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#9ECBFF">    "limited"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"ltd"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> normalise</span><span style="color:#E1E4E8">(name: </span><span style="color:#79B8FF">str</span><span style="color:#E1E4E8">) -&gt; </span><span style="color:#79B8FF">str</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#E1E4E8">    n </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> name.lower().replace(</span><span style="color:#9ECBFF">"&amp;"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">" and "</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#E1E4E8">    n </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> re.sub(</span><span style="color:#F97583">r</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">[</span><span style="color:#F97583">^</span><span style="color:#79B8FF">\w\s]</span><span style="color:#9ECBFF">"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">" "</span><span style="color:#E1E4E8">, n)</span></span>
<span class="line"><span style="color:#E1E4E8">    n </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> re.sub(</span><span style="color:#F97583">r</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">\s</span><span style="color:#F97583">+</span><span style="color:#9ECBFF">"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">" "</span><span style="color:#E1E4E8">, n).strip()</span></span>
<span class="line"><span style="color:#E1E4E8">    n </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> re.sub(</span><span style="color:#F97583">r</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">^</span><span style="color:#DBEDFF">the </span><span style="color:#9ECBFF">"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">""</span><span style="color:#E1E4E8">, n)</span></span>
<span class="line"><span style="color:#F97583">    for</span><span style="color:#E1E4E8"> long_form, short_form </span><span style="color:#F97583">in</span><span style="color:#79B8FF"> SUFFIXES</span><span style="color:#E1E4E8">.items():</span></span>
<span class="line"><span style="color:#E1E4E8">        n </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> re.sub(</span><span style="color:#F97583">rf</span><span style="color:#9ECBFF">"\b</span><span style="color:#79B8FF">{</span><span style="color:#E1E4E8">long_form</span><span style="color:#79B8FF">}</span><span style="color:#9ECBFF">$"</span><span style="color:#E1E4E8">, short_form, n)</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> n</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> split_trustee</span><span style="color:#E1E4E8">(name: </span><span style="color:#79B8FF">str</span><span style="color:#E1E4E8">):</span></span>
<span class="line"><span style="color:#E1E4E8">    m </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> re.match(</span><span style="color:#F97583">r</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF">(.</span><span style="color:#F97583">+?</span><span style="color:#79B8FF">)\s</span><span style="color:#F97583">+</span><span style="color:#79B8FF">(?:</span><span style="color:#DBEDFF">atf</span><span style="color:#F97583">|</span><span style="color:#DBEDFF">as trustee for</span><span style="color:#79B8FF">)\s</span><span style="color:#F97583">+</span><span style="color:#79B8FF">(.</span><span style="color:#F97583">+</span><span style="color:#79B8FF">)</span><span style="color:#9ECBFF">"</span><span style="color:#E1E4E8">, name, re.I)</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> (m.group(</span><span style="color:#79B8FF">1</span><span style="color:#E1E4E8">), m.group(</span><span style="color:#79B8FF">2</span><span style="color:#E1E4E8">)) </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> m </span><span style="color:#F97583">else</span><span style="color:#E1E4E8"> (name, </span><span style="color:#79B8FF">None</span><span style="color:#E1E4E8">)</span></span></code></pre>
<p>The matcher runs in order of strength. An ABN or ACN match is near certain. An exact match on normalised names comes next, then token and phonetic similarity with a nickname table for given names. Trustee names are split from the trust name and both are searched. Where the firm has an ABN for a party, a lookup against the Australian Business Register adds the registered name and trading names to the search terms. Once a party matches, the matcher expands one hop through the relationships the practice-management system already holds, such as directors, trustees and parent companies.</p>
<p>Each candidate comes out with a score and a plain reason (“ACN match”, “normalised name exact”, “name similarity 0.91, same suburb”). Thresholds are set low on purpose. A false positive costs a reviewer a few seconds, while a missed conflict can cost the firm a client or worse. Nothing is ever cleared automatically, and that includes a search that returns no hits.</p>
<h2 id="where-the-llm-helps">Where the LLM helps</h2>
<p>The rules leave a tail of middling candidates, where a reviewer has to think about whether “Harbour Fresh” on a 2019 matter is the same business as the enquirer. For these, the model gets the extracted party, the candidate record and its surrounding context (addresses, matter descriptions, linked parties). It returns a structured verdict of likely same, likely related, unlikely or can’t tell, with a short explanation.</p>
<p>The model can also propose extra search terms, such as a likely trading name or the individual behind a sole-trader business name. Those terms go back through the deterministic matcher, so every hit still carries a rule-based reason.</p>
<p>The model’s permissions are narrow by design. It can add candidates, reorder the queue and attach notes. It cannot remove a candidate from the reviewer’s list or mark anything as cleared. In the review screen its notes are labelled as model output, kept separate from the matcher’s reasons.</p>
<p>Enquiry text is confidential and often sensitive, so the model should run in an Australian region under terms that exclude training on your data. Logs should be retained according to the firm’s existing records policy rather than the vendor’s default.</p>
<h2 id="a-person-clears-every-potential-conflict">A person clears every potential conflict</h2>
<p>The review screen shows the original enquiry, the extracted parties with their evidence, and every candidate for each party with its reasons. The reviewer records one of three outcomes per candidate (no conflict, conflict, or refer to a partner), and the system stores who made each decision and when. The matter cannot open in the practice-management system until every candidate has a recorded decision, and the firm’s obligations under the Australian Solicitors’ Conduct Rules stay with the people making those calls.</p>
<p>Reviewer decisions feed back into the system, with care. A confirmed non-match is stored so the next search shows that pair pre-annotated (“cleared by J. Nguyen, March 2026, different ABN”). It still appears on the list, because circumstances change and a pair that was unrelated last year may not be now.</p>
<h2 id="costs-and-limits">Costs and limits</h2>
<p>Integration is most of the work. The matcher needs a current, complete index of every party across every matter, closed ones included, and practice-management data that goes back years is usually messy, with duplicate contacts, inconsistent entity types and free-text notes holding the relationships. A party added this morning has to be searchable this afternoon, so the index needs frequent incremental syncs from the source system. This is the kind of sync job we built Centazio, our open-source data integration platform, to handle. Model calls are a small running cost next to the integration build and the reviewer time the process still needs.</p>
<p>The system knows only what the firm has recorded. Conflicts held in people’s heads, such as a lateral hire’s former clients, stay invisible until someone enters them. Trust beneficiaries and ultimate owners are often unknown at intake, and no amount of matching will find a relationship nobody has captured.</p>
<p>Before switching over, run the pipeline in shadow mode alongside the existing manual process. Any conflict the manual search found and the pipeline missed is a defect to fix before go-live. Replaying a few months of past enquiries through it first gives you a test set and a realistic picture of reviewer workload.</p>
<h2 id="starting-small">Starting small</h2>
<p>A sensible first project is extraction plus deterministic matching, running in shadow mode against your practice-management data, with the review screen built but not yet controlling matter opening. That shows quickly how clean your party data is and how many grey-zone candidates you actually get, which tells you whether the LLM review step is worth adding.</p>
<p>PicNet builds production AI systems for Australian organisations. If intake and conflict checking is taking up your team’s time, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/matter-intake-conflict-checks-and-entity-resolution/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Entity resolution at scale: deterministic first, LLM for the tail</title><link>https://picnet.com.au/blog/entity-resolution-at-scale-deterministic-first-llm-for-the-tail/</link><guid isPermaLink="true">https://picnet.com.au/blog/entity-resolution-at-scale-deterministic-first-llm-for-the-tail/</guid><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><description>A practical entity resolution pipeline: rules clear the bulk, an LLM proposes matches for the ambiguous tail, and people sign off risky merges.</description><category>ai</category><category>dataintegration</category><category>entityresolution</category><category>recordlinkage</category><content:encoded><![CDATA[<p>Most organisations we work with hold the same customer three or four times. The CRM has one copy under a trading name. The finance system has another under the registered company name, and a third came in from a web form with a typo in the email address. Any AI project that reads across those systems picks up the duplicates. Ask a model to summarise “the customer’s history” and it will give you a confident summary of a third of it.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/ai-ready-data-and-integration-the-series/">AI-Ready Data and Integration</a> series. It describes the matching pipeline we recommend. Exact and rule-based matching handles most records, and an LLM proposes matches for the ambiguous pairs left over. Before any merge above a set risk level is applied, a person approves it.</p>
<h2 id="why-rules-go-first">Why rules go first</h2>
<p>A rule runs in microseconds and costs almost nothing. When a rule merges two records you can point to the reason, and someone will eventually ask why their account history changed. An LLM call is slower and costs money every time. It can also give a different answer to the same pair on a different day.</p>
<p>Volume is the bigger problem. Comparing every record with every other record means n(n-1)/2 comparisons, which comes to roughly 500 billion pairs for a million records. No model budget covers that. The deterministic stages exist to cut that number down to something a model and a review team can handle.</p>
<h2 id="stage-one-normalise-block-match">Stage one: normalise, block, match</h2>
<p>Most matching failures come from formatting, so normalise before you compare anything:</p>
<ul>
<li>Names: case, whitespace and punctuation, with legal suffixes removed (“Pty Ltd”, “Pty. Limited”, “P/L”).</li>
<li>Addresses: standard street types (St, Street), unit and level notation, and “c/-” care-of lines moved out of the street field.</li>
<li>Phone numbers: converted to E.164 (+61…), which makes landlines and mobiles comparable across systems.</li>
</ul>
<p>Next, match exactly on strong identifiers. For organisations that usually means the ABN. One ABN can cover several branches or trading names, though, so we pair it with a name or address rule and don’t merge on ABN alone. For individuals, a verified email or a customer number shared between systems does the same job.</p>
<p>For the remainder, blocking keeps the comparison count manageable. You only compare records that share a cheap key, such as postcode plus the first few letters of the normalised name. Within each block, you score pairs with string similarity on names and token overlap on addresses. Open-source tools such as Splink, which implements the Fellegi-Sunter probabilistic model, and Zingg handle this layer well. Neither needs an LLM.</p>
<p>The output falls into three bands. Pairs in the confident-match and confident-non-match bands are settled. Only the grey band between them goes to the next stage.</p>
<h2 id="stage-two-an-llm-for-the-grey-band">Stage two: an LLM for the grey band</h2>
<p>The grey band holds the pairs rules handle badly: “Bob’s Plumbing” and “R. Smith Plumbing Services” with the same mobile number, a nickname against a full name, a business that moved two suburbs over. An LLM reads both records side by side the way a person would. It can also use free-text notes and order descriptions that are hard to write rules for.</p>
<p>We ask for structured output with three possible decisions (match, no_match or unsure), the fields the decision relied on, and a one-line reason. The model only proposes. It never writes to the master record. The routing logic stays deterministic:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> route</span><span style="color:#E1E4E8">(pair, verdict):</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> pair.abn_match </span><span style="color:#F97583">and</span><span style="color:#E1E4E8"> pair.name_score </span><span style="color:#F97583">&gt;</span><span style="color:#79B8FF"> 0.9</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#9ECBFF"> "auto_merge"</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> pair.score </span><span style="color:#F97583">&lt;</span><span style="color:#79B8FF"> LOWER_BAND</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#9ECBFF"> "no_match"</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> verdict.label </span><span style="color:#F97583">==</span><span style="color:#9ECBFF"> "match"</span><span style="color:#F97583"> and</span><span style="color:#E1E4E8"> pair.risk_tier </span><span style="color:#F97583">==</span><span style="color:#9ECBFF"> "low"</span><span style="color:#F97583"> and</span><span style="color:#E1E4E8"> band_precision(pair) </span><span style="color:#F97583">&gt;=</span><span style="color:#79B8FF"> 0.98</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#9ECBFF"> "auto_merge"</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> verdict.label </span><span style="color:#F97583">==</span><span style="color:#9ECBFF"> "no_match"</span><span style="color:#F97583"> and</span><span style="color:#E1E4E8"> pair.risk_tier </span><span style="color:#F97583">==</span><span style="color:#9ECBFF"> "low"</span><span style="color:#E1E4E8">:</span></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#9ECBFF"> "no_match"</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#9ECBFF"> "human_review"</span></span></code></pre>
<p>Look at <code>band_precision</code>. Don’t route on the model’s own stated confidence, because nothing guarantees that a “0.9” from an LLM means 90 per cent. Route on how accurate the model has proven to be for that score band against your labelled data (covered below).</p>
<h2 id="stage-three-people-confirm-risky-merges">Stage three: people confirm risky merges</h2>
<p>Set the risk threshold by consequence, not by score. Merging two prospects on a marketing list does little harm if it’s wrong. Merging two customer accounts with credit balances, or two people’s personal details, is another matter. Australian Privacy Principle 10 requires reasonable steps to keep personal information accurate. A wrong merge that shows one person’s details to another can turn into a privacy incident, and you may have to assess it under the Notifiable Data Breaches scheme.</p>
<p>Design the review step as part of the pipeline from the start:</p>
<ul>
<li>A queue that shows both records, the fields that differ, and the model’s reason.</li>
<li>Merges that can be undone. Keep the source records and a crosswalk of system IDs, and log every merge decision with who or what made it.</li>
<li>Reviewer decisions fed back into the labelled set, so each week of review improves your measurements.</li>
</ul>
<h2 id="measure-on-your-own-data-before-trusting-any-matcher">Measure on your own data before trusting any matcher</h2>
<p>Precision is the share of proposed merges that are correct. Recall is the share of true duplicates the pipeline finds. Published benchmarks and vendor figures come from someone else’s data. Your names, your legacy systems’ quirks and your data entry habits will all differ from theirs.</p>
<p>Build a labelled set before go-live. Sample a few hundred pairs across every score band. A purely random sample will be almost all obvious non-matches and won’t tell you much. Have two people label each pair independently, then settle the disagreements. Expect a few: some pairs are hard for humans as well, and those belong in review permanently.</p>
<p>Score each stage separately against that set. Then pick operating points: very high precision for anything that auto-merges, and higher recall where missing a duplicate is the expensive mistake, such as screening for duplicate supplier payments. Re-run the set whenever you change a rule, a prompt or a model version. Treat it like a regression suite.</p>
<h2 id="what-it-costs-at-volume">What it costs at volume</h2>
<p>The LLM cost comes down to one line:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="plaintext"><code><span class="line"><span>cost  grey_band_pairs  tokens_per_pair  price_per_token</span></span></code></pre>
<p>Here’s a worked example with assumed numbers. Two million records produce 10 million candidate pairs after blocking. Rules settle all but 2 per cent, which leaves 200,000 pairs for the model. At around 800 tokens per pair, input and output combined, that’s 160 million tokens. Plug in your provider’s current rate card. At mid-sized model prices, the model bill is usually smaller than the human one.</p>
<p>Say 5 per cent of those pairs go to review. That’s 10,000 decisions, and at 30 seconds each it adds up to about 83 hours of staff time. Reviewer hours are the cost to plan and budget for.</p>
<p>To keep both numbers down:</p>
<ul>
<li>Tighten blocking, which shrinks every stage after it.</li>
<li>Start with a small model and escalate only the unsure pairs to a larger one.</li>
<li>Cache verdicts keyed on a hash of the two normalised records, so unchanged pairs never get judged twice.</li>
<li>Run matching incrementally, comparing only new and changed records each night.</li>
<li>Use batch endpoints for the initial backfill. Most providers price them below interactive calls.</li>
</ul>
<h2 id="limitations">Limitations</h2>
<p>The same model can give different verdicts across runs and versions. Pin the model version, set temperature to zero, and lean on the labelled set to catch drift. If you send personal information to a model hosted overseas, APP 8 on cross-border disclosure applies. Check whether your provider offers the model in an Australian region, or run a smaller model in your own environment. Rules need maintenance as source systems change. Finally, no matcher fixes data that was wrong when it was entered. Entity resolution finds the duplicates, and stopping new ones is a data entry and integration problem.</p>
<p>PicNet builds production AI systems for Australian organisations. If you have duplicate records across your systems, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first matching project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/entity-resolution-at-scale-deterministic-first-llm-for-the-tail/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Cloud cost control as code: automated spend reports, stale resources and off-hours shutdowns</title><link>https://picnet.com.au/blog/cloud-cost-control-as-code-automated-spend-reports-stale-resources-and-off-hours-shutdowns/</link><guid isPermaLink="true">https://picnet.com.au/blog/cloud-cost-control-as-code-automated-spend-reports-stale-resources-and-off-hours-shutdowns/</guid><pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate><description>How we automate AWS and Azure cost control with scheduled scripts, and why the LLM in the pipeline only explains spend spikes and never does the maths.</description><category>ai</category><category>devops</category><category>cloudcostmanagement</category><category>aws</category><content:encoded><![CDATA[<p>Most cloud bills grow slowly. A test VM runs through a long weekend, or a disk outlives the server it was attached to, and nobody notices until the invoice arrives. The fast blowouts are rarer and much worse. In September a Reddit user <a href="https://www.reddit.com/r/GoogleGeminiAI/comments/1wjlys0/i_got_an_87k_bill_for_gemini_use_by_google_cloud/">reported an $87K Google Cloud bill for Gemini use</a>, and people in the thread traced it to a phishing ad. Commenters said it was “not possible to place any limit” on the account.</p>
<p>We handle both kinds of blowout with scheduled code that runs against the AWS and Azure environments we look after. A spend report with anomaly flags goes out every morning. Two other jobs run beside it: a stale-resource sweep and an evening shutdown of anything tagged as temporary. This post is part of our <a href="https://picnet.com.au/blog/practical-ai-in-devops-the-series/">Practical AI in DevOps</a> series. Most of what follows is ordinary scripting, and the LLM has one narrow job at the end, which is to explain a spend spike in plain English.</p>
<h2 id="start-with-yesterdays-numbers">Start with yesterday’s numbers</h2>
<p>Both clouds expose billing data through an API. On AWS it’s Cost Explorer, and on Azure it’s the Cost Management query API. Each morning the report job pulls the previous day’s cost, grouped by account or subscription and by service, and writes the rows to a small table. Convert the currency when you load the data so the AWS and Azure figures can sit in one column.</p>
<p>Billing data arrives late on both clouds and keeps changing for a while after that. Treat the report as an accurate view of yesterday, not a live alarm, and re-pull the previous few days on every run so the late adjustments get picked up.</p>
<p>Deciding what counts as an anomaly takes arithmetic, not AI. Compare yesterday with the median of the same weekday over the previous six weeks. Flag it only when the increase clears both a percentage and a dollar floor:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> statistics</span></span>
<span class="line"></span>
<span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> is_anomaly</span><span style="color:#E1E4E8">(yesterday_aud, same_weekday_history, pct</span><span style="color:#F97583">=</span><span style="color:#79B8FF">0.4</span><span style="color:#E1E4E8">, floor_aud</span><span style="color:#F97583">=</span><span style="color:#79B8FF">50</span><span style="color:#E1E4E8">):</span></span>
<span class="line"><span style="color:#E1E4E8">    baseline </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> statistics.median(same_weekday_history)</span></span>
<span class="line"><span style="color:#E1E4E8">    delta </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> yesterday_aud </span><span style="color:#F97583">-</span><span style="color:#E1E4E8"> baseline</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> delta </span><span style="color:#F97583">&gt;</span><span style="color:#E1E4E8"> floor_aud </span><span style="color:#F97583">and</span><span style="color:#E1E4E8"> delta </span><span style="color:#F97583">&gt;</span><span style="color:#E1E4E8"> baseline </span><span style="color:#F97583">*</span><span style="color:#E1E4E8"> pct</span></span></code></pre>
<p>The dollar floor means a $3 service that doubles to $6 doesn’t wake anyone up. Using the same weekday as the baseline means a normal Monday doesn’t look like a spike after a quiet weekend. Tune both thresholds for each account.</p>
<p>Run the check for each account or subscription, and again on the total across all of them. Commenters in the Gemini thread said the attackers got around per-project throttling by <a href="https://www.reddit.com/r/GoogleGeminiAI/comments/1wjlys0/i_got_an_87k_bill_for_gemini_use_by_google_cloud/">“creating dozens of projects”</a>. Each project can stay under its own threshold while the total climbs, so only the summed check catches that pattern. Count resources as well as dollars. If four running GPU instances become forty, the resource count shows it hours before the billing data does.</p>
<p>Turn on AWS Budgets and Azure budgets as well. By default they send an email and leave everything running. The Gemini case shows the same gap: per-project throttling looked like a spend cap, but it didn’t stop the spend.</p>
<h2 id="finding-stale-resources">Finding stale resources</h2>
<p>The stale-resource sweep looks for things that cost money and have no obvious owner or purpose:</p>
<ul>
<li>EBS volumes and Azure managed disks with nothing attached</li>
<li>snapshots older than the retention policy allows</li>
<li>public IP addresses not associated with any resource</li>
<li>Azure VMs that are stopped but not deallocated, which still bill for compute</li>
<li>anything without an <code>owner</code> tag</li>
</ul>
<p>The sweep only reports. A disk with no VM is usually a leftover, but now and then it’s the only copy of something someone needs. Each finding goes into the morning report with its monthly cost and the identity that created it. That identity comes from CloudTrail or the Azure Activity Log, both of which keep about 90 days of history by default, so older resources often show no creator. When the list names each item and its price, the owners can clean up quickly.</p>
<h2 id="shutting-down-temporary-resources-every-evening">Shutting down temporary resources every evening</h2>
<p>Dev and test VMs carry a tag such as <code>lifecycle=temporary</code>. At 7pm Sydney time the shutdown job stops every tagged EC2 instance and deallocates every tagged Azure VM. On Azure, deallocating matters: a VM that’s only stopped keeps billing for compute.</p>
<p>The savings are easy to work out. A week has 168 hours. A VM that runs 7am to 7pm on weekdays is up for 60 of them, so its on-demand compute charge drops by about 64% compared with running all week. That figure has limits. Attached disks and static IPs keep billing while the VM is off. Instances covered by reserved instances or a savings plan save little or nothing when stopped, because you pay for the commitment either way.</p>
<p>Exceptions use a second tag, for example <code>keep-until=2026-10-09</code>, and the job skips that resource until the date passes. Every exception carries an expiry date, so a “just this week” request can’t turn into a permanent one.</p>
<p>Schedule in local time. Daylight saving in NSW starts on the first Sunday in October, and a cron expression written in UTC will run an hour off for half the year. EventBridge Scheduler and Logic Apps recurrence triggers both accept <code>Australia/Sydney</code> as a time zone. Give the shutdown identity permission to read tags and to stop or deallocate instances, and nothing else. It can’t create or delete anything.</p>
<p>Test the job before it touches a real account. <a href="https://fakecloud.dev/">Fakecloud</a> emulates AWS locally, including EC2 and the Resource Groups Tagging API, so tag-based shutdown logic can run in CI with no credentials and no cost. <a href="https://floci.io">Floci</a> has emulators for both AWS and Azure. Its Azure build lists Blob storage, Functions, Key Vault and Service Bus, so check whether it covers the compute APIs your script calls before you rely on it for VM logic.</p>
<h2 id="api-throttling">API throttling</h2>
<p>These jobs tend to work on a small account and fail on a big one. A script that loops through every region and subscription as fast as it can will hit the management API rate limits. EC2 calls start returning <code>RequestLimitExceeded</code> and other AWS services return <code>ThrottlingException</code>. On Azure, Resource Manager returns HTTP 429 with a <code>Retry-After</code> header, and the Cost Management query API has its own separate limits.</p>
<p>The failure that costs money is a shutdown run that gets throttled halfway through. It logs an error nobody reads and leaves half the tagged machines running all night. Here’s what prevents it:</p>
<ul>
<li>page through results and send stop calls in batches, since EC2 <code>StopInstances</code> accepts many instance IDs per call</li>
<li>use the SDK’s retry handling with backoff, and honour <code>Retry-After</code> on Azure</li>
<li>process accounts and subscriptions one after another rather than all at once</li>
<li>schedule jobs a few minutes off the hour so they don’t collide with every other job set for 7:00</li>
<li>run the shutdown a second time half an hour later, and expect it to find nothing</li>
<li>name any instance that failed to stop in the next morning’s report</li>
</ul>
<p>On AWS, the adaptive retry mode in boto3 handles most of this:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">import</span><span style="color:#E1E4E8"> boto3</span></span>
<span class="line"><span style="color:#F97583">from</span><span style="color:#E1E4E8"> botocore.config </span><span style="color:#F97583">import</span><span style="color:#E1E4E8"> Config</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">ec2 </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> boto3.client(</span></span>
<span class="line"><span style="color:#9ECBFF">    "ec2"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#FFAB70">    region_name</span><span style="color:#F97583">=</span><span style="color:#9ECBFF">"ap-southeast-2"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#FFAB70">    config</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">Config(</span><span style="color:#FFAB70">retries</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">{</span><span style="color:#9ECBFF">"max_attempts"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">10</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"mode"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"adaptive"</span><span style="color:#E1E4E8">}),</span></span>
<span class="line"><span style="color:#E1E4E8">)</span></span></code></pre>
<h2 id="where-the-llm-earns-its-place">Where the LLM earns its place</h2>
<p>By the time the LLM gets involved, code has already found the anomaly and calculated the numbers. The model receives the flagged rows, the deltas and the resources created or changed in that account the day before. From these it writes a short paragraph saying which service went up and by how much, and which resources and creators were responsible. An IT manager can read that without opening Cost Explorer or knowing what a usage type code means.</p>
<p>The model never adds up a total or decides what counts as an anomaly. Language models make mistakes when they sum tables, and the same input can produce a different answer tomorrow. The model also has no cloud credentials, so it can explain a spike but can’t act on one. Once it has written its paragraph, a few lines of code confirm that every dollar figure in the text appears in the input. If a figure doesn’t match, the report goes out with only the plain table.</p>
<p>Running costs are small. A daily explanation uses a few thousand tokens, well inside the lightest usage tier (under 1M tokens a day) in <a href="https://www.sitepoint.com/local-llms-vs-cloud-api-cost-analysis-2026/">SitePoint’s 2026 cost analysis</a>. Before sending anything to a hosted model, check whether resource names and tags contain client or project names you’d rather not send offshore. Strip them first or use a model hosted in an Australian region.</p>
<h2 id="limits-of-this-approach">Limits of this approach</h2>
<p>Because of billing lag, the daily report catches yesterday’s problem today. For a fast blowout like the Gemini case, resource counts and provider budget alerts react sooner, but nothing here will stop someone using a stolen credential. That needs separate controls on identity and access. Tag-based shutdown only works if people tag their resources. Anything untagged ends up in the stale sweep, which reports but never acts. If most of your compute runs on reserved capacity, the evening shutdown will save less than cleaning up what the stale sweep finds.</p>
<p>PicNet builds production AI systems for Australian organisations. If you’d like daily spend reports and evening shutdowns like these running across your AWS and Azure accounts, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/cloud-cost-control-as-code-automated-spend-reports-stale-resources-and-off-hours-shutdowns/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>A research assistant grounded in your own precedents and knowledge</title><link>https://picnet.com.au/blog/a-research-assistant-grounded-in-your-own-precedents-and-knowledge/</link><guid isPermaLink="true">https://picnet.com.au/blog/a-research-assistant-grounded-in-your-own-precedents-and-knowledge/</guid><pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate><description>A practical design for a law firm research assistant that answers only from your own precedents, with citations and matter-level access control.</description><category>ai</category><category>legal</category><category>retrievalaugmentedgeneration</category><category>legaltechnology</category><content:encoded><![CDATA[<p>A senior associate wants to know how the firm usually drafts a limitation of liability clause for a software supply agreement. The answer is in the firm somewhere, either in the precedent bank or in a file note a partner wrote three years ago. A general-purpose chatbot will write a confident clause from its training data. The associate needs the firm’s own clause and a link to the document it came from.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/practical-ai-for-law-firms-the-series/">Practical AI for Law Firms</a> series. It covers how we would build that assistant. The design uses retrieval-augmented generation (RAG) over the firm’s own material. Answers cite the source document, the assistant refuses when the corpus has nothing, and access control works at matter level so information cannot cross an ethical wall.</p>
<h2 id="what-grounding-means-in-practice">What grounding means in practice</h2>
<p>A grounded assistant answers only from the documents it retrieved for the current question. Every claim in the answer points to a passage in a specific document, and the reader can open that document and check it.</p>
<p>The open-source ai-advisory-board project has a clean version of this pattern. A recent change added a <a href="https://github.com/HaroldZhong/ai-advisory-board/pull/113">Materials sidebar with versioned source citations</a>, which ties each answer to the exact version of the document it used. Versioning matters for a precedent bank because precedents get revised after a judgment or a change in legislation. A citation to “the indemnity precedent” tells the reader much less than a citation to version 4, approved in March, which they can compare with the current version.</p>
<p>Refusal is the other half. If retrieval returns nothing relevant, the assistant says so and stops. It doesn’t fall back on general knowledge, because that is where invented authorities come from.</p>
<h2 id="the-failure-mode-grounding-prevents">The failure mode grounding prevents</h2>
<p>Courts in Australia and overseas have dealt with submissions that cited cases that don’t exist. General-purpose AI tools produced the citations and nobody checked them before filing. Several Australian courts have since issued practice notes or guidance on using generative AI in proceedings. Any firm planning an assistant should read the one for its jurisdiction before it designs anything.</p>
<p>These incidents have one cause. Someone asked a model for authority, and it produced text that looked like authority with no document behind it. A grounded design closes that path. The model sees only the retrieved passages. The system attaches citations from retrieval metadata, and a post-check drops any reference that doesn’t match a passage in the retrieved set.</p>
<p>The lawyer still reviews the answer before relying on it. That review is quick, because each citation opens the paragraph it came from.</p>
<h2 id="an-architecture-sketch">An architecture sketch</h2>
<p>The components are ordinary. The design work is in deciding where each control sits.</p>
<ul>
<li>Ingestion: connectors pull from the document management system (DMS), the precedent bank and the knowledge wiki. The connectors are read-only. The same PR adds an <a href="https://github.com/HaroldZhong/ai-advisory-board/pull/113">optional read-only local MCP reader</a> that works the same way: the assistant can read documents but has no way to change them.</li>
<li>Chunking and metadata: each document is split into passages. Every passage carries its document ID, version, practice group, matter number where there is one, and the access groups copied from the DMS.</li>
<li>Index: a vector index combined with keyword search, hosted in an Australian cloud region. Legal text needs the keyword side because embeddings handle defined terms and clause numbers poorly.</li>
<li>Retrieval: every query runs under the user’s own identity, and the index applies the access filter inside the search call.</li>
<li>Generation: the model receives the retrieved passages. Its instructions are to answer only from them and to say plainly when they don’t cover the question.</li>
<li>Citation check: a post-processing step confirms that every cited passage ID was in the retrieved set. It removes or flags anything that wasn’t.</li>
<li>Audit log: every query is logged with the user and the passages returned, so the firm can reconstruct any answer later.</li>
</ul>
<p>The agentic-graphrag project made a similar choice when it <a href="https://github.com/ontogr/agentic-graphrag/pull/26">grounded its query generation in the engine’s resolved schema</a> in place of a hardcoded description, with one bounded retry. The same change treats the caller’s base scope as an authorisation boundary. Filters on an individual call can narrow that scope but never widen it, and the system refuses requests outside it. Matter access needs exactly that rule.</p>
<p>In outline, the request path looks like this:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="python"><code><span class="line"><span style="color:#F97583">def</span><span style="color:#B392F0"> answer</span><span style="color:#E1E4E8">(question, user):</span></span>
<span class="line"><span style="color:#E1E4E8">    scope </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> access_scope_for(user)    </span><span style="color:#6A737D"># DMS groups plus the wall register</span></span>
<span class="line"><span style="color:#E1E4E8">    passages </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> index.search(</span></span>
<span class="line"><span style="color:#E1E4E8">        question,</span></span>
<span class="line"><span style="color:#FFAB70">        filter</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">scope,                  </span><span style="color:#6A737D"># enforced by the index at query time</span></span>
<span class="line"><span style="color:#FFAB70">        top_k</span><span style="color:#F97583">=</span><span style="color:#79B8FF">12</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#E1E4E8">    )</span></span>
<span class="line"><span style="color:#E1E4E8">    relevant </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> [p </span><span style="color:#F97583">for</span><span style="color:#E1E4E8"> p </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> passages </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> p.score </span><span style="color:#F97583">&gt;=</span><span style="color:#79B8FF"> MIN_RELEVANCE</span><span style="color:#E1E4E8">]</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#F97583"> not</span><span style="color:#E1E4E8"> relevant:</span></span>
<span class="line"><span style="color:#F97583">        return</span><span style="color:#E1E4E8"> Refusal(</span><span style="color:#9ECBFF">"No matching precedent or know-how in the firm's library."</span><span style="color:#E1E4E8">)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">    draft </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> llm.generate(question, </span><span style="color:#FFAB70">context</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">relevant, </span><span style="color:#FFAB70">rules</span><span style="color:#F97583">=</span><span style="color:#79B8FF">GROUNDED_ONLY</span><span style="color:#E1E4E8">)</span></span>
<span class="line"><span style="color:#E1E4E8">    allowed </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> {p.id </span><span style="color:#F97583">for</span><span style="color:#E1E4E8"> p </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> relevant}</span></span>
<span class="line"><span style="color:#E1E4E8">    cited </span><span style="color:#F97583">=</span><span style="color:#E1E4E8"> [c </span><span style="color:#F97583">for</span><span style="color:#E1E4E8"> c </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> draft.citations </span><span style="color:#F97583">if</span><span style="color:#E1E4E8"> c.passage_id </span><span style="color:#F97583">in</span><span style="color:#E1E4E8"> allowed]</span></span>
<span class="line"><span style="color:#E1E4E8">    log_query(user, question, relevant)</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> Answer(draft.text, </span><span style="color:#FFAB70">citations</span><span style="color:#F97583">=</span><span style="color:#E1E4E8">cited)</span></span></code></pre>
<h2 id="ethical-walls-belong-in-the-retrieval-layer">Ethical walls belong in the retrieval layer</h2>
<p>The firm’s DMS already records who can open which matter, and the assistant has to inherit those permissions exactly. Some information barriers are recorded outside the DMS, such as a conflicts register entry that walls a team off from a former client’s files. The assistant has to honour those as well.</p>
<p>The filter has to run inside the index query. Suppose walled documents reach the model’s context and you try to strip them from the answer afterwards. The model can still paraphrase what it read. When the index applies the filter, walled material is never retrieved, so it has no route into an answer.</p>
<p>Some practical points from designing this:</p>
<ul>
<li>Sync permissions often. A wall put up this morning must apply to this afternoon’s queries. Permission changes should reach the index as events, with a scheduled full reconciliation as a backstop.</li>
<li>Decide carefully what goes into the shared corpus. A precedent that has been cleaned of client detail and approved for firm-wide use belongs there. A raw advice on a matter file stays in that matter’s scope, even if it is the best example of the point.</li>
<li>Keep refusals identical. The user should get the same message whether nothing exists or they lack access, so a refusal doesn’t reveal that a walled document exists.</li>
<li>Test the walls. Write a set of questions that must return nothing for a walled user, and run it on every index rebuild.</li>
</ul>
<h2 id="refusal-is-a-written-policy">Refusal is a written policy</h2>
<p>Where you draw the refusal line is a choice, and teams do change it. The argus project, a financial research assistant, recorded a <a href="https://github.com/lagarcess/argus/pull/589">decision to reverse its earlier refusal</a> of forward-looking questions. It now answers them as grounded scenarios built from cited inputs. For a law firm we would start stricter and answer only from the corpus. Practice groups can then loosen the rule for particular question types, with each change written down and approved by a named partner.</p>
<h2 id="where-the-documents-live">Where the documents live</h2>
<p>Precedents and know-how are among a firm’s most sensitive assets. In September a <a href="https://www.reddit.com/r/technology/comments/1waoiai/report_openai_stole_mathematicians_private/">widely discussed report alleged</a> that OpenAI took mathematicians’ private research from their Codex chats. The allegation is unverified. It still shows why firms are wary of pasting confidential material into hosted AI tools.</p>
<p>For this assistant we would keep the index and document store in the firm’s own cloud tenancy, in an Australian region. The model endpoint should be under enterprise terms that exclude training on inputs, and those terms should be confirmed in writing before any client material is indexed. The firm owes clients the same confidentiality in the index as it does in the DMS.</p>
<h2 id="limitations-and-costs">Limitations and costs</h2>
<ul>
<li>Corpus quality sets the ceiling. If the precedent bank holds conflicting versions of a clause with no status marked, the assistant will retrieve all of them. Most of the effort in a first project goes into marking which precedent is current and which are superseded.</li>
<li>The assistant doesn’t research external law. Case law and legislation still come from the firm’s subscription services. The assistant reports what the firm has written and done before.</li>
<li>Retrieval sometimes misses. If a question is phrased differently from the precedent, the search can fail to find it, and the assistant refuses even though the answer exists. Reviewing refusals regularly shows where the corpus or the search needs work.</li>
<li>Model usage is usually the smaller cost. Ingestion, permission sync and evaluation take most of the effort, and that work continues after launch because the corpus keeps changing.</li>
</ul>
<p>A sensible first project covers one practice group’s precedent bank under its existing permissions. That group’s lawyers write the test questions. The assistant then runs under supervision for a few weeks before anyone else gets access.</p>
<p>PicNet builds production AI systems for Australian organisations. If a research assistant grounded in your own precedents is on your list, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/a-research-assistant-grounded-in-your-own-precedents-and-knowledge/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Code-first integration with Centazio: why we built an open-source alternative to drag-and-drop iPaaS</title><link>https://picnet.com.au/blog/code-first-integration-with-centazio-why-we-built-an-open-source-alternative-to-drag-and-drop-ipaas/</link><guid isPermaLink="true">https://picnet.com.au/blog/code-first-integration-with-centazio-why-we-built-an-open-source-alternative-to-drag-and-drop-ipaas/</guid><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><description>Why PicNet built Centazio, an open-source .NET integration platform, how its read, promote and write design works, and when an iPaaS is still the right choice.</description><category>ai</category><category>dataintegration</category><category>centazio</category><category>ipaas</category><content:encoded><![CDATA[<p>Integration projects often go the same way. A team buys an iPaaS subscription, drags a CRM connector onto a canvas and wires it to the finance system. By Friday, records are flowing. Six months later the canvas has forty boxes. A dozen of them hold expressions that only one person understands, and nobody wants to touch any of it before end of month.</p>
<p>We built Centazio, PicNet’s open-source data integration platform, to do integration a different way. In Centazio an integration is plain .NET code, split into small isolated functions, and developers review and unit test it the same way they handle the rest of their code.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/ai-ready-data-and-integration-the-series/">AI-Ready Data and Integration</a> series. An AI system that reads your operational data picks up every duplicate and stale record your integrations produce, so AI readiness has to start at the integration layer.</p>
<h2 id="where-drag-and-drop-tools-fall-short">Where drag-and-drop tools fall short</h2>
<p>Visual integration tools demo well because a demo only runs the happy path: a record is created in one system and shows up in another. Most of the work in a production integration is the other cases. The target API goes down for an hour. Someone edits a record in both systems between syncs. A vendor renames a field, or a rate limit kicks in halfway through a batch.</p>
<p>Every one of those cases needs logic. On a canvas, that logic ends up in condition blocks and expression fields spread across the flow, where it is hard to find and harder to change safely. These are the problems we run into most:</p>
<ul>
<li>Version control is weak. Flows are usually stored as large generated JSON or XML files, so a diff tells a reviewer very little about what changed.</li>
<li>Testing is manual. You can usually run a flow against a sandbox. What you can’t do is write an automated test that says “this CRM contact becomes this finance contact” and have it run on every change.</li>
<li>Pricing grows with usage. Licences are often metered per connector or per task run, and many are billed in US dollars, so an Australian budget moves with the exchange rate.</li>
<li>The logic is locked in. A flow built in one vendor’s designer won’t open in another vendor’s designer, so if you leave, you rebuild.</li>
</ul>
<h2 id="how-centazio-structures-an-integration">How Centazio structures an integration</h2>
<p>Centazio splits every integration into three kinds of step. Each step runs as its own serverless function:</p>
<ul>
<li>A read function for each source system pulls changed records and stages them in that system’s own format.</li>
<li>A promote function takes the staged records, validates them and maps them into a shared core model that your organisation defines.</li>
<li>A write function for each target system takes core records and pushes them out in the shape that system expects.</li>
</ul>
<p>No system knows about any other. The CRM read function has no idea the finance system exists, and the finance write function only knows the core model. To add a third system, you write its functions against the core model and leave the existing ones alone.</p>
<h2 id="fault-tolerance-comes-from-the-structure">Fault tolerance comes from the structure</h2>
<p>Each step reads from storage and writes to storage, so a failure stays in the step where it happened. Say the finance system’s API is down. Its write function fails and tries again on the next run. Meanwhile the CRM read and the promote step keep running, so the core model stays current. When the API comes back, the write function processes everything that queued up in the meantime. Nobody has to replay a flow by hand or work out which half of a batch went through.</p>
<p>The split also helps when a mapping turns out to be wrong. The raw data is still in staging, so once the promote code is fixed you re-promote what was already read. You don’t need to pull it from the source system again, which matters when that source has tight API limits.</p>
<p>Every function is small and does one job, so when a log entry or alert fires, it points to one step for one system.</p>
<h2 id="tests-come-with-the-design">Tests come with the design</h2>
<p>A promote step takes a staged record and returns either a core record or a reason to skip it. That is ordinary code with no network calls, so you unit test it like any other code. The sketch below shows the pattern. The type names are made up for illustration and are not Centazio’s exact API.</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="csharp"><code><span class="line"><span style="color:#F97583">public</span><span style="color:#F97583"> class</span><span style="color:#B392F0"> ContactPromoter</span><span style="color:#E1E4E8"> {</span></span>
<span class="line"><span style="color:#F97583">  public</span><span style="color:#B392F0"> PromoteResult</span><span style="color:#B392F0"> Promote</span><span style="color:#E1E4E8">(</span><span style="color:#B392F0">CrmContact</span><span style="color:#B392F0"> c</span><span style="color:#E1E4E8">) {</span></span>
<span class="line"><span style="color:#F97583">    if</span><span style="color:#E1E4E8"> (String.</span><span style="color:#B392F0">IsNullOrWhiteSpace</span><span style="color:#E1E4E8">(c.Email)) </span><span style="color:#F97583">return</span><span style="color:#E1E4E8"> PromoteResult.</span><span style="color:#B392F0">Ignore</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">"no email"</span><span style="color:#E1E4E8">);</span></span>
<span class="line"><span style="color:#F97583">    var</span><span style="color:#B392F0"> name</span><span style="color:#F97583"> =</span><span style="color:#9ECBFF"> $"</span><span style="color:#9ECBFF">{</span><span style="color:#E1E4E8">c</span><span style="color:#9ECBFF">.</span><span style="color:#E1E4E8">First</span><span style="color:#9ECBFF">}</span><span style="color:#9ECBFF"> {</span><span style="color:#E1E4E8">c</span><span style="color:#9ECBFF">.</span><span style="color:#E1E4E8">Last</span><span style="color:#9ECBFF">}</span><span style="color:#9ECBFF">"</span><span style="color:#E1E4E8">.</span><span style="color:#B392F0">Trim</span><span style="color:#E1E4E8">();</span></span>
<span class="line"><span style="color:#F97583">    return</span><span style="color:#E1E4E8"> PromoteResult.</span><span style="color:#B392F0">Ok</span><span style="color:#E1E4E8">(</span><span style="color:#F97583">new</span><span style="color:#B392F0"> Customer</span><span style="color:#E1E4E8">(name, c.Email.</span><span style="color:#B392F0">ToLowerInvariant</span><span style="color:#E1E4E8">()));</span></span>
<span class="line"><span style="color:#E1E4E8">  }</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span>
<span class="line"></span>
<span class="line"><span style="color:#E1E4E8">[</span><span style="color:#B392F0">Test</span><span style="color:#E1E4E8">] </span><span style="color:#F97583">public</span><span style="color:#F97583"> void</span><span style="color:#B392F0"> Promote_ignores_contacts_without_email</span><span style="color:#E1E4E8">() {</span></span>
<span class="line"><span style="color:#F97583">  var</span><span style="color:#B392F0"> res</span><span style="color:#F97583"> =</span><span style="color:#F97583"> new</span><span style="color:#B392F0"> ContactPromoter</span><span style="color:#E1E4E8">().</span><span style="color:#B392F0">Promote</span><span style="color:#E1E4E8">(</span><span style="color:#F97583">new</span><span style="color:#B392F0"> CrmContact</span><span style="color:#E1E4E8">(</span><span style="color:#9ECBFF">"Jo"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">"Smith"</span><span style="color:#E1E4E8">, </span><span style="color:#9ECBFF">""</span><span style="color:#E1E4E8">));</span></span>
<span class="line"><span style="color:#E1E4E8">  Assert.</span><span style="color:#B392F0">That</span><span style="color:#E1E4E8">(res.Ignored, Is.True);</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span></code></pre>
<p>Tests like this run in the build pipeline on every change. When someone changes a mapping rule, the pull request shows exactly which line changed and which test covers it. An IT manager can approve that after reading the diff, without sitting through a demo in a sandbox.</p>
<h2 id="what-code-first-costs">What code-first costs</h2>
<p>Code-first has real costs, and they should be on the table before anyone picks it:</p>
<ul>
<li>You need .NET developers, either in-house or through a partner. A business analyst can’t build a new flow in an afternoon.</li>
<li>You host it and you run it. The functions run in your own cloud tenancy. That makes it easier to keep data in an Australian region when your privacy obligations or client contracts require it, but the monitoring, patching and cloud bills are also yours.</li>
<li>Connectors are code too. An iPaaS may ship a ready-made connector for a popular SaaS product. With Centazio, someone writes and maintains the read and write functions for that product’s API.</li>
<li>Open source doesn’t come with a vendor support desk. The code is yours to read and change, and support comes from your own team or from us.</li>
</ul>
<h2 id="when-an-ipaas-is-still-the-right-call">When an iPaaS is still the right call</h2>
<p>A commercial iPaaS is often the better fit in these situations:</p>
<ul>
<li>The systems are mainstream SaaS products with maintained prebuilt connectors, and the flows are simple one-way pushes.</li>
<li>Volumes are low, and a missed or duplicated record costs little.</li>
<li>There is no development capacity, none is planned, and the people who own the process need to change the flows themselves.</li>
<li>The integration won’t be around long, for example a bridge during a system migration.</li>
</ul>
<p>Centazio suits the other end of the scale. That means several systems sharing the same entities, two-way sync, business rules that keep changing, and data that other systems will depend on. We use one test to decide: would you accept this logic untested if it lived in your main application? If the answer is no, it belongs in code.</p>
<h2 id="where-it-runs-today">Where it runs today</h2>
<p>Centazio runs production integrations at CYCA and at Guide Dogs NSW/ACT, built with the read, promote and write structure described above. Both are plain .NET code, so changes go through code review and automated tests before they deploy, and neither organisation pays a per-task integration licence.</p>
<h2 id="why-this-matters-for-ai-work">Why this matters for AI work</h2>
<p>The core model is what pays off later. Once records from every system have been promoted into one validated shape, an AI project has one tested source to read from. If an agent needs to write back to a system, it can use the same write functions and the tests that cover them. AI tooling is moving towards code over configuration as well: <a href="https://github.com/enclawed/omcp">OpenMCP</a>, posted to Hacker News on 22 September, calls itself an open, code-first fork of the Model Context Protocol.</p>
<p>PicNet builds production AI systems for Australian organisations. If your data sits in systems that don’t talk to each other, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/code-first-integration-with-centazio-why-we-built-an-open-source-alternative-to-drag-and-drop-ipaas/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Watching the watchers: monitoring your CI/CD pipelines, branch protection and repo hygiene</title><link>https://picnet.com.au/blog/watching-the-watchers-monitoring-your-ci-cd-pipelines-branch-protection-and-repo-hygiene/</link><guid isPermaLink="true">https://picnet.com.au/blog/watching-the-watchers-monitoring-your-ci-cd-pipelines-branch-protection-and-repo-hygiene/</guid><pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate><description>How we monitor GitHub and Azure DevOps for failing scheduled pipelines, missing branch protection and stale branches, plus AI weekly digests of what changed.</description><category>ai</category><category>devops</category><category>cicd</category><category>github</category><content:encoded><![CDATA[<p>Most organisations monitor production closely and barely monitor their delivery tooling. Someone sets up a nightly build or a branch protection rule and confirms it works that day. After that, nobody looks at it again until a release breaks or an auditor asks how changes get onto main.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/practical-ai-in-devops-the-series/">Practical AI in DevOps</a> series. It covers the continuous checks we run over GitHub and Azure DevOps, and a trap that produces false findings if you build those checks the obvious way. AI comes in at the end, where it turns a week of findings into a short digest so engineers only read what changed.</p>
<h2 id="what-the-checks-cover">What the checks cover</h2>
<p>We run five checks on a daily schedule:</p>
<ul>
<li>Scheduled workflows and pipelines whose recent scheduled runs failed, or that have stopped running.</li>
<li>Default and release branches with no branch protection or ruleset, or with rules weaker than the organisation’s baseline, such as no required review.</li>
<li>Open Dependabot alerts, grouped by severity and by how long they have been open.</li>
<li>Open secret-scanning alerts.</li>
<li>Stale branches with no commits past a threshold you choose (90 days is a reasonable start) and no open pull request.</li>
</ul>
<p>Scheduled jobs are on the list because nobody is waiting for them. When a pull request build fails, the person who opened the pull request sees it. When a dependency refresh fails at 2am on a Sunday, the notice goes to whatever notification setting was configured, which may belong to someone who left last year.</p>
<h2 id="how-it-fits-together">How it fits together</h2>
<p>We keep the architecture small on purpose:</p>
<ul>
<li>A timer-triggered job runs once a day. A GitHub Actions workflow in a separate admin repository will do the job, and so will an Azure Function.</li>
<li>It reads GitHub through a GitHub App with read-only permissions for Actions, Administration, Contents, Dependabot alerts and Secret scanning alerts. For Azure DevOps it uses a service principal or a PAT scoped to Build (read) and Code (read).</li>
<li>Each check writes its findings to a store in one shared format. A JSON file per day in blob storage is enough.</li>
<li>Once a week, a diff step compares the latest findings with the previous week’s.</li>
<li>A language model gets only the diff and writes the digest, which goes to a Teams channel or an email list.</li>
</ul>
<p>A single finding looks like this:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="json"><code><span class="line"><span style="color:#E1E4E8">{</span></span>
<span class="line"><span style="color:#79B8FF">  "source"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"github"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "repo"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"example-org/payments-api"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "check"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"branch-protection"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "subject"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"main"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "severity"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"high"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "detail"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"Default branch has no protection rule or ruleset"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "first_seen"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"2026-09-11"</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span></code></pre>
<p>Each finding is identified by its source, repo, check and subject. The weekly diff only works if that identity stays stable from one run to the next.</p>
<p>GitHub branch protection has a trap of its own. The classic protection endpoint (<code>GET /repos/{owner}/{repo}/branches/{branch}/protection</code>) doesn’t report repository rulesets, so a branch protected only by a ruleset looks unprotected. Query <code>GET /repos/{owner}/{repo}/rules/branches/{branch}</code> as well, and count either one as protection.</p>
<h2 id="dormant-yaml-files-versus-registered-pipelines">Dormant YAML files versus registered pipelines</h2>
<p>The obvious first version of a pipeline check walks the repository tree, looking for files under <code>.github/workflows</code> or an <code>azure-pipelines.yml</code>. That approach is wrong in both directions. A file in a repository only records what someone once declared, and the CI platform decides what runs.</p>
<p>In Azure DevOps, a YAML file does nothing until someone creates a pipeline that points at it, and that pipeline can point at any path. Repositories collect dormant pipeline files from project templates and abandoned experiments. If you probe for <code>azure-pipelines.yml</code>, a dormant file makes it look as if CI exists when it doesn’t. A repository whose real pipeline uses <code>build/ci.yml</code> gets reported as having no CI. A registered pipeline can also be paused or disabled, and a file probe can’t see either state.</p>
<p>GitHub registers workflow files automatically, but the file still leaves things out. Scheduled triggers only run from the default branch, so a schedule in a file that exists only on a feature branch never fires. GitHub can also disable a workflow, and the API reports this as a <code>state</code> of <code>disabled_manually</code> or <code>disabled_inactivity</code>.</p>
<p>Ask the platform instead:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="bash"><code><span class="line"><span style="color:#6A737D"># GitHub: registered workflows and their state</span></span>
<span class="line"><span style="color:#B392F0">gh</span><span style="color:#9ECBFF"> api</span><span style="color:#9ECBFF"> repos/ORG/REPO/actions/workflows</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#79B8FF">  --jq</span><span style="color:#9ECBFF"> '.workflows[] | {name, path, state}'</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># GitHub: recent scheduled runs for one workflow</span></span>
<span class="line"><span style="color:#B392F0">gh</span><span style="color:#9ECBFF"> api</span><span style="color:#9ECBFF"> "repos/ORG/REPO/actions/workflows/WORKFLOW_ID/runs?event=schedule&amp;per_page=5"</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#79B8FF">  --jq</span><span style="color:#9ECBFF"> '.workflow_runs[] | {created_at, conclusion}'</span></span>
<span class="line"></span>
<span class="line"><span style="color:#6A737D"># Azure DevOps: registered pipelines, their YAML path and queue status</span></span>
<span class="line"><span style="color:#B392F0">curl</span><span style="color:#79B8FF"> -s</span><span style="color:#79B8FF"> -u</span><span style="color:#9ECBFF"> ":</span><span style="color:#E1E4E8">$AZDO_PAT</span><span style="color:#9ECBFF">"</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#9ECBFF">  "https://dev.azure.com/ORG/PROJECT/_apis/build/definitions?includeAllProperties=true&amp;api-version=7.1"</span><span style="color:#79B8FF"> \</span></span>
<span class="line"><span style="color:#F97583">  |</span><span style="color:#B392F0"> jq</span><span style="color:#9ECBFF"> '.value[] | {name, yaml: .process.yamlFilename, status: .queueStatus}'</span></span></code></pre>
<p>Then join the file view to the platform view. A pipeline file with no registered definition gets an informational note saying the file is dormant. A registered definition whose YAML path no longer exists, or whose recent scheduled runs failed, is a real finding.</p>
<p>The same failure turned up recently in another tool. In <a href="https://github.com/jdx/mise/pull/13190">a fix to mise</a>, <code>mise doctor</code> reported a background watcher as “declared but not running” while the service manager was running it, and the command it recommended did nothing. The maintainers changed the health checks to use the same test as the command that acts on that state, so the checks no longer recommend a command that won’t do anything. A monitoring check should read the same state the system acts on, and that is rarely a file path.</p>
<h2 id="keeping-the-hygiene-reports-readable">Keeping the hygiene reports readable</h2>
<p>A stale-branch report that lists 300 branches will get ignored. Report a count per repository, list only stale branches that have no open pull request, and leave out release branch patterns you keep on purpose. Don’t delete branches automatically. An old branch can hold the only copy of a hotfix that was never merged.</p>
<p>Any open secret-scanning alert should get a person’s attention that week. The check records the secret type and the repository, with a link to the alert. It never stores the secret value, so secrets can’t end up in anything sent to a model later.</p>
<p>For Dependabot, the useful signal is age and severity together. A critical alert that has been open for a day is normal triage. If the same alert has been open for six weeks, nobody is looking after that repository.</p>
<h2 id="the-weekly-digest">The weekly digest</h2>
<p>Engineers stop reading a report that says the same thing every week. The diff step lists what is new and what was resolved since last week. It also carries forward anything still open past an age threshold. Only those items go to the model.</p>
<p>The model handles wording and grouping, and deterministic code decides what counts as a finding. The prompt asks for one line per item, grouped by team, with each line ending in the link supplied in the input. A cheap post-check compares every repository named in the digest against the input and rejects the output if the model names one that wasn’t there.</p>
<p>An example digest:</p>
<blockquote>
<p>New this week: the <code>payments-api</code> main branch lost its ruleset (link). The nightly dependency refresh in <code>customer-portal</code> has failed its last three scheduled runs (link). Resolved: two high-severity Dependabot alerts in <code>reporting-service</code>.</p>
</blockquote>
<p>The model only ever sees metadata such as repository names, check names, severities and dates. Some organisations treat repository names as sensitive too. In that case, use a model hosted in an Australian region, such as Azure OpenAI in Australia East, and keep the digest inside your own tenant.</p>
<h2 id="costs-and-limits">Costs and limits</h2>
<ul>
<li>GitHub and Azure DevOps both rate-limit their APIs. Large organisations need pagination and caching, and a daily run is plenty.</li>
<li>Dependabot alerts come with every GitHub plan. Secret scanning on private repositories needs GitHub’s paid Advanced Security features. The Azure DevOps equivalent, GitHub Advanced Security for Azure DevOps, is also a paid add-on. Check what you are licensed for before promising coverage.</li>
<li>The model input is a short diff, so the model cost is small.</li>
<li>The platforms change under you. GitHub added rulesets alongside classic branch protection, and a check written before rulesets misses any rule configured that way. Budget time for maintenance.</li>
<li>The monitor is a scheduled job too. Give it a heartbeat, so that if a daily findings file doesn’t arrive, something outside the monitor raises an alert.</li>
<li>The checks report problems and fix nothing. Every finding needs an owner, and the digest only helps if someone is expected to act on it.</li>
</ul>
<p>PicNet builds production AI systems for Australian organisations. If a weekly digest over your pipelines sounds like a sensible place to start, <a href="https://picnet.com.au/ai-services/">talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/watching-the-watchers-monitoring-your-ci-cd-pipelines-branch-protection-and-repo-hygiene/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Contract review and clause extraction: what AI can and can't do</title><link>https://picnet.com.au/blog/contract-review-and-clause-extraction-what-ai-can-and-can-t-do/</link><guid isPermaLink="true">https://picnet.com.au/blog/contract-review-and-clause-extraction-what-ai-can-and-can-t-do/</guid><pubDate>Wed, 23 Sep 2026 00:00:00 GMT</pubDate><description>How to build an LLM first-pass contract review that extracts clauses into a structured summary a lawyer verifies, plus the accuracy and confidentiality limits.</description><category>ai</category><category>legal</category><category>contractreview</category><category>clauseextraction</category><content:encoded><![CDATA[<p>The first AI idea most firms have is to paste a contract into a chat window and ask what’s wrong with it. The output reads well. It names a few issues, sounds confident, and gives a partner no way to tell whether the flagged problem came from reading clause 14.2 or from the model’s sense of what a liability complaint usually looks like. That question, whether a flagged issue is judgment or surface pattern matching, is the one practitioners keep returning to in <a href="https://www.reddit.com/r/AINativeServices/comments/1w1ukv4/when_your_ainative_law_firm_flags_a_critical/">discussions of AI-native legal review</a>, and it’s the question a workflow has to answer structurally rather than by asking the model to try harder.</p>
<p>The version that works in production is narrower and duller. You ask the model to extract facts into a fixed schema, with a pointer back to the text each fact came from, and you hand the result to a lawyer as a starting worksheet.</p>
<h2 id="why-extraction-outperforms-an-open-ended-review">Why extraction outperforms an open-ended review</h2>
<p>“Review this contract” asks for three things at once: find the relevant text, decide what it means in context, and decide whether it’s acceptable for this client on this deal. The model has some capability at the first, patchy capability at the second, and none at the third, because acceptability depends on commercial appetite, the relationship with the counterparty and what the firm agreed last time.</p>
<p>Extraction splits those apart. The model finds and quotes; your playbook of standard positions decides what counts as a deviation; the lawyer decides what to do about it. Each step can be checked. When a comparison against a standard position is wrong, you can see whether the extraction was wrong or the playbook rule was.</p>
<p>The checklist drives the loop, one query per item you care about. This matters more than it sounds, because the expensive failure in contract review is the clause that isn’t there. A model summarising a contract will rarely tell you there’s no assignment restriction, no GST gross-up, or no cap carve-out for a data breach. Ask per item and “not present in this document” becomes an answer the system produces, not a silence you have to notice.</p>
<h2 id="the-shape-of-the-extraction-output">The shape of the extraction output</h2>
<p>Keep the record flat and make every field verifiable against a span of source text:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="json"><code><span class="line"><span style="color:#E1E4E8">{</span></span>
<span class="line"><span style="color:#79B8FF">  "item"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"limitation-of-liability"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "found"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">true</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "clause_ref"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"14.2"</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "page"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">18</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "quote"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"The Supplier's aggregate liability under or in connection with this Agreement shall not exceed the Charges paid in the twelve (12) months preceding the claim."</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "summary"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"Liability capped at fees paid in the 12 months before the claim."</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "standard_position"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"Cap at 125% of total contract value, uncapped for breach of confidentiality."</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "deviation"</span><span style="color:#E1E4E8">: </span><span style="color:#79B8FF">true</span><span style="color:#E1E4E8">,</span></span>
<span class="line"><span style="color:#79B8FF">  "notes"</span><span style="color:#E1E4E8">: </span><span style="color:#9ECBFF">"No carve-out located for confidentiality or IP indemnity."</span></span>
<span class="line"><span style="color:#E1E4E8">}</span></span></code></pre>
<p>Around that sit the boring fields a matter file needs anyway: parties and their ACNs, execution and commencement dates, term and renewal mechanics, termination triggers and notice periods, governing law, payment terms, and the obligations with dates attached that someone will need to diarise.</p>
<p>The <code>quote</code> field carries most of the weight. A summary with no quote can’t be checked in under a minute, so it gets skimmed and trusted. A summary with a quote and a clause reference can be confirmed or rejected at a glance, and the lawyer’s eye goes to the text rather than to the model’s prose.</p>
<h2 id="verify-the-extraction-with-something-other-than-the-model-that-produced-it">Verify the extraction with something other than the model that produced it</h2>
<p>Asking an LLM to grade its own answer produces agreeable nonsense. The pattern that’s emerged in engineering tooling is a separate verification pass: <a href="https://github.com/Kidus-M/MaruCheck">MaruCheck</a>, published on Show HN this month, exists specifically to provide independent QA for AI-generated code rather than relying on the generating model to check itself. Clause extraction needs the same separation.</p>
<p>In our pipelines the verification step is mostly deterministic and cheap:</p>
<ul>
<li>String-match every <code>quote</code> back into the source document. A quote that doesn’t appear verbatim is a fabrication and the record is dropped or re-run.</li>
<li>Check <code>clause_ref</code> and <code>page</code> against the document structure parsed at ingestion, so a quote attributed to the wrong clause fails.</li>
<li>Re-run the deviation comparison as a second call that sees only the quote and the standard position, with no access to the first model’s conclusion.</li>
<li>Flag documents where OCR confidence is low or the clause numbering doesn’t parse, and route those to manual review instead of extracting from mush.</li>
</ul>
<p>Ignore the model’s self-reported confidence score. It’s fluent, not calibrated, and it tends to be highest on the clauses that are drafted most conventionally, which are the ones you needed help with least.</p>
<h2 id="the-lawyers-markup-stays-the-deliverable">The lawyer’s markup stays the deliverable</h2>
<p>The strongest argument in that practitioner discussion is about artefacts: the thing clients pay for is the redline, and the value only lands when the lawyer’s pass is captured as edits. Build the workflow so the machine output is a worksheet sitting beside the document, and the returned markup is what goes out the door.</p>
<p><a href="https://github.com/daniel-freiermuth/bug-hunter/pull/8">CodeRabbit’s auto-generated review summaries</a> show the production shape of this in software: a structured first-pass summary posted alongside the pull request, with the human reviewer’s decision remaining authoritative. Contract review maps onto it cleanly. The summary lands in the matter in your DMS or practice management system, the lawyer accepts, edits or rejects each row, and the accepted rows flow into the redline and the obligations register.</p>
<p>Capture the accept and reject decisions. After a few hundred contracts you have a record of which checklist items the extraction handles reliably and which ones a senior lawyer overturns half the time. That’s the evidence you need to decide where to spend engineering effort, and the only honest basis for telling partners how much of the first pass they can lean on.</p>
<h2 id="what-accuracy-to-expect">What accuracy to expect</h2>
<p>Nobody can give you a number that transfers. Published accuracy claims for clause extraction are measured on someone else’s document set, usually clean, usually in US or English drafting conventions, and they say nothing about how a tool performs on your precedents, your counterparties’ paper, and the scanned 2013 deed of variation that amends clause 14 without renumbering anything.</p>
<p>Measure it yourself before you commit. Take 30 to 50 contracts representative of the work, have a senior lawyer mark up what a correct extraction looks like for each checklist item, and score the pipeline against that set. Report recall and precision separately, because they cost different amounts: a false positive wastes fifteen seconds of a lawyer’s time, while a missed indemnity carve-out can go out the door. Re-run the set whenever you change a model, a prompt or a document parser.</p>
<p>The failure modes are consistent. Defined terms that shift meaning across a suite of documents, obligations incorporated by reference from a schedule or a URL, conditional carve-outs (“except as provided in clause 9.3”), and anything in a table or a poorly scanned annexure are where extraction degrades. Long agreements also lose fidelity toward the end of a context window, which is a reason to chunk by clause and run per-item queries rather than posting the whole document with one broad instruction.</p>
<h2 id="confidentiality-and-cost">Confidentiality and cost</h2>
<p>Contracts are dense with material facts, which is why they reward systematic analysis. The Intercept reconstructed the working relationship between the Pentagon and OpenAI, Google and Anthropic <a href="https://theintercept.com/2026/09/08/military-ai-weapons-contracts-openai-anthropic-google/">from more than 400 pages of contract paperwork</a> obtained under FOI. The same property makes these documents the last thing you want sitting in a third party’s training pipeline or retained logs.</p>
<p>Before any pilot, settle where the documents go: which region the inference runs in, what the provider’s retention and training terms say, whether your client engagement terms and confidentiality undertakings permit disclosure to that processor, and who signs off on that decision inside the firm. Get privacy advice on cross-border disclosure rather than inferring it from a vendor’s marketing page. Deployment options in Australian regions exist across the major cloud providers, and for firms with the strictest undertakings we’ve run smaller open-weight models on infrastructure the client controls, accepting lower extraction quality in exchange for the documents never leaving.</p>
<p>Running costs are modest. Per-document API spend for a commercial agreement is small against six minutes of a lawyer’s time, and the tokens are a rounding error in the project. The real cost is the gold set, the playbook of standard positions written down properly for the first time, and the ongoing maintenance when models change underneath you. Budget for a pilot on one contract type, with one checklist and one practice group, before anyone talks about rolling it out across the firm.</p>
<p>This post is part of our <a href="https://picnet.com.au/blog/practical-ai-for-law-firms-the-series/">Practical AI for Law Firms</a> series, which covers the systems Australian firms can put into production now and the ones worth waiting on.</p>
<p>PicNet builds production AI systems for Australian organisations. <a href="https://picnet.com.au/ai-services/">Talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/contract-review-and-clause-extraction-what-ai-can-and-can-t-do/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Your AI is only as good as your data: why integration comes first</title><link>https://picnet.com.au/blog/your-ai-is-only-as-good-as-your-data-why-integration-comes-first/</link><guid isPermaLink="true">https://picnet.com.au/blog/your-ai-is-only-as-good-as-your-data-why-integration-comes-first/</guid><pubDate>Mon, 21 Sep 2026 00:00:00 GMT</pubDate><description>Most failed AI pilots are data failures, not model failures. What AI-ready really means, the integration architecture behind it, and a checklist before you fund.</description><category>ai</category><category>dataintegration</category><category>masterdata</category><category>aireadiness</category><content:encoded><![CDATA[<p>Almost every stalled AI pilot we get called into looks the same from the inside. The model was fine. The demo worked on the sample extract. Then someone pointed it at production and the answers stopped matching what the business knew to be true, because the customer record in the CRM disagreed with the one in the finance system, the product catalogue had been exported to a spreadsheet eight months ago, and nobody could say which of the two addresses on file was current.</p>
<p>That is a data problem wearing a model costume. This post is part of our <a href="https://picnet.com.au/blog/ai-ready-data-and-integration-the-series/">AI-Ready Data and Integration</a> series, and it covers the part that has to happen before any of the interesting work: getting a trustworthy view of your own operations into a form a machine can read.</p>
<h2 id="why-software-engineering-got-there-first">Why software engineering got there first</h2>
<p>AI uptake has run far ahead in software engineering compared with other business functions, as <a href="https://www.economist.com/business/2026/08/30/will-anybody-use-ai-as-much-as-coders-do">The Economist</a> set out in August. The usual explanation is that developers are early adopters. A better one is that code is already in the shape AI needs. It is text, it is versioned, every change has an author and a timestamp, and there is exactly one repository that counts as the truth. A coding agent never has to work out which branch is the real one.</p>
<p>Now look at the data behind a typical Australian mid-market operation. Customer records in a CRM, billing in an ERP, service history in a ticketing system, a warehouse that only some reports use, and a set of spreadsheets that hold the rules nobody wrote down. No timestamps you would trust. No agreed precedence. An AI project starting there is being asked to do the reconciliation work first and the clever work second, with no tools for the first job.</p>
<h2 id="what-ai-ready-actually-requires">What AI-ready actually requires</h2>
<p>Three things, and none of them are model choices.</p>
<p>A master view. One place that resolves customers, products, transactions and the events connecting them into single records, with the source of every field recorded. Not a copy of each system side by side, which is what most “data lakes” turn out to be, but a resolved view where the question “how many active customers do we have” has one answer.</p>
<p>Freshness guarantees you can state in numbers. For each entity, how old can this be before it is wrong to act on it? Contact details might tolerate a day. Stock levels might tolerate five minutes. Write the number down per entity, measure it in production, and alarm on it. This matters more than which model you pick, because the model contributes nothing current by itself. A <a href="https://stale.jock.pl/">public tracker of 20 models across eight labs</a> shows median release age measured in weeks and training cutoffs months behind release, and only half of those models publish a cutoff at all. Everything the model knows about your customers arrives at query time from your systems, or it does not arrive.</p>
<p>A conflict rule. When two systems disagree, one of them wins, and the rule is written down before the disagreement happens. This is the part organisations skip. It is also the part that decides whether anyone trusts the output six months in.</p>
<h2 id="an-architecture-that-survives-contact-with-production">An architecture that survives contact with production</h2>
<p>This is the shape we build, and it is the shape <a href="https://github.com/PicNet/Centazio">Centazio</a>, our open-source integration platform, is designed around.</p>
<ul>
<li>Read functions per source system, each one responsible for pulling changes and nothing else. They write raw payloads to staging, unmodified, with the retrieval time recorded.</li>
<li>Promote functions that map staged payloads into core entities. Mapping rules live here, in code, under version control.</li>
<li>An entity mapping store that holds the relationship between external identifiers and internal core identifiers, so the same customer arriving from three systems resolves to one record and you can always trace back which system contributed which field.</li>
<li>Write functions that push core state back out to the systems that need it, closing the loop rather than leaving the master view as a read-only reporting artefact.</li>
<li>Checkpoints and operational state per function, so a failed run resumes from where it stopped instead of reprocessing everything.</li>
</ul>
<p>Conflict resolution sits in the promote step and should be boring and explicit:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="csharp"><code><span class="line"><span style="color:#6A737D">// Billing owns the legal entity name and ABN.</span></span>
<span class="line"><span style="color:#6A737D">// CRM owns contact details, but only if its record is fresher than billing's.</span></span>
<span class="line"><span style="color:#6A737D">// Anything older than the freshness budget is not promoted at all; it raises an alert.</span></span>
<span class="line"><span style="color:#F97583">var</span><span style="color:#B392F0"> name</span><span style="color:#F97583">    =</span><span style="color:#E1E4E8"> billing.LegalName </span><span style="color:#F97583">??</span><span style="color:#E1E4E8"> crm.TradingName;</span></span>
<span class="line"><span style="color:#F97583">var</span><span style="color:#B392F0"> email</span><span style="color:#F97583">   =</span><span style="color:#B392F0"> Fresher</span><span style="color:#E1E4E8">(crm.UpdatedUtc, billing.UpdatedUtc) </span><span style="color:#F97583">?</span><span style="color:#E1E4E8"> crm.Email </span><span style="color:#F97583">:</span><span style="color:#E1E4E8"> billing.Email;</span></span>
<span class="line"><span style="color:#F97583">var</span><span style="color:#B392F0"> stale</span><span style="color:#F97583">   =</span><span style="color:#E1E4E8"> UtcNow </span><span style="color:#F97583">-</span><span style="color:#B392F0"> Newest</span><span style="color:#E1E4E8">(crm, billing) </span><span style="color:#F97583">&gt;</span><span style="color:#E1E4E8"> budgets.Contact; </span><span style="color:#6A737D">// 24h</span></span>
<span class="line"><span style="color:#F97583">if</span><span style="color:#E1E4E8"> (stale) alerts.</span><span style="color:#B392F0">Raise</span><span style="color:#E1E4E8">(ConflictAlert.StaleEntity, coreId);</span></span></code></pre>
<p>Five lines of deterministic code replace an assumption the model would otherwise have to guess at. That is the argument the author of <a href="https://nloum.github.io/blog/codeio/">CodeIO</a> made in September about AI-assisted development: supply deterministic, verifiable inputs and you get measurably better output than letting the model infer context. The same principle holds at the data layer. Every fact you can resolve before the prompt is a fact the model cannot get wrong.</p>
<h2 id="contradiction-is-the-normal-case">Contradiction is the normal case</h2>
<p>There is now open-source tooling whose entire premise is that <a href="https://github.com/vikcena01/ai-continuity-plugin">AI sessions silently contradict earlier state</a> unless one authoritative record exists. Engineers building agent infrastructure are <a href="https://www.reddit.com/r/ContextEngineering/comments/1w75rp9/i_vibecoded_infrastructure_for_ai_agents_heres/">separating persistent memory from durable execution</a> for the same reason: the record of what is true has to be managed separately from whatever the agent is doing this minute. When practitioners build guardrails against a failure, that failure is the default behaviour, not an edge case.</p>
<p>Meanwhile the market’s instinct is to add more surfaces. September brought another wave of tooling to <a href="https://news.ycombinator.com/item?id=49728159">connect agents to WhatsApp</a>. Every new channel is another place a customer can state a fact and another version of that fact to reconcile. Adding channels before you have decided which system wins multiplies the contradictions.</p>
<h2 id="the-stakes-rise-when-the-ai-starts-acting">The stakes rise when the AI starts acting</h2>
<p>An AI that drafts a summary and gets it wrong wastes a few minutes. An AI that has authority to transact does something to the world. <a href="https://www.theatlantic.com/technology/2026/09/instinct-ai-personal-assistant-credit-card/688607/">The Atlantic</a> wrote in September about handing a personal assistant a credit card, and the point generalises to any agent with write access to your systems: bad underlying data now produces a wrong action, an incorrect invoice or a refund to the wrong account.</p>
<p>Source-of-truth governance is also a security control. Lakera’s public <a href="https://play.lakera.ai/agent-breaker">agent-breaker exercises</a> demonstrate agents being subverted through the content they ingest. If your agent reads free-text fields that anyone outside the organisation can write into, those fields are an input channel for instructions, and they need the same treatment as any other untrusted input. Under the Australian Privacy Principles you also need to know which system holds the authoritative copy of personal information before you can answer a correction request, which is a good reason to sort this out regardless of AI.</p>
<p>In administrative health settings, the same architecture applies to referrals, bookings and billing reconciliation. Any output that touches clinical judgement needs a named clinician signing off inside the workflow, designed in from the start, with the AI restricted to assembling and presenting the record.</p>
<h2 id="a-readiness-checklist-before-you-fund-anything">A readiness checklist before you fund anything</h2>
<p>Ask your team these questions. Written answers, not verbal ones.</p>
<ol>
<li>For customers, products and each other core entity, which system is authoritative, and where is that written down?</li>
<li>When two systems disagree on a field, what is the resolution rule and who approved it?</li>
<li>What is the freshness budget for each entity, in minutes or hours, and do we measure the actual lag today?</li>
<li>Can we trace any field in a report back to the system and the record it came from?</li>
<li>How many customer records exist in more than one system with no link between the copies?</li>
<li>If the AI writes back to a source system, what is the rollback path when it writes something wrong?</li>
<li>Which free-text fields does the AI read, and who can write into them from outside the organisation?</li>
</ol>
<p>An initiative that cannot answer one to three should fund integration work first. The AI project does not disappear; it starts three months later with a foundation, and it usually gets cheaper because the prompt no longer has to compensate for missing context.</p>
<h2 id="what-it-costs">What it costs</h2>
<p>Integration is unglamorous and it is usually the larger half of a first AI project’s budget. Connecting two systems properly, with mapping, conflict rules, monitoring and a replay path for failures, takes weeks rather than days, and legacy systems with poor APIs take longer. We built Centazio and released it as open source because we were rewriting the same staging, checkpointing and entity-mapping machinery on every engagement, and that machinery is not where the value sits.</p>
<p>The payoff is that the work is reusable. A master view built for a customer service assistant also serves the next forecasting model, the reporting rebuild and the system you replace in two years. Model choice is a decision you will revisit every few months as the tracker above keeps ticking over. The data foundation is the part you build once.</p>
<p>PicNet builds production AI systems for Australian organisations. <a href="https://picnet.com.au/ai-services/">Talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/your-ai-is-only-as-good-as-your-data-why-integration-comes-first/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
<item><title>Coding agents in a production .NET shop: what holds up</title><link>https://picnet.com.au/blog/coding-agents-in-a-production-net-shop-what-actually-works/</link><guid isPermaLink="true">https://picnet.com.au/blog/coding-agents-in-a-production-net-shop-what-actually-works/</guid><pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate><description>How PicNet runs Claude-style coding agents on production .NET code: standards docs in the repo, tests as the safety net, AI review first, and what agents still get wrong.</description><category>ai</category><category>devops</category><category>codingagents</category><category>dotnet</category><content:encoded><![CDATA[<p>We have been running Claude-style CLI coding agents against production .NET code for long enough to have opinions that survived contact with a client’s release schedule. This post, part of our <a href="https://picnet.com.au/blog/practical-ai-in-devops-the-series/">Practical AI in DevOps</a> series, describes the setup we use, the parts that pay for themselves, and the work we still hand back to a person.</p>
<p>Two findings shape the rest of it.</p>
<p>A coding agent is a model plus a harness: the tools it can call, the context it starts with, the loop that decides what it sees next. A Berkeley group compared 21 model and harness pairings and found harness choice barely moved success rates, staying within about 2% on SWE-bench Lite, while moving cost a long way. The same model solved roughly the same share of tasks at up to five times the price depending on what wrapped it, and Claude Code averaged about twice Pi’s cost per attempt (<a href="https://harnesstax.github.io/">HarnessTax</a>). Effort spent on repo scaffolding returns more than effort spent swapping models each time a new one ships.</p>
<p>The second finding is that an agent does what is cheap rather than what is correct. When a study gave agents both <code>grep</code> and language-server navigation on the same repositories, they picked the semantic tool between 0% and 6% of the time on code-localisation tasks, and forcing the semantic path first dropped success from 100% to 89% (<a href="https://www.agentconnect.md/blog/grep-beat-lsp-harness/">agentconnect</a>). On a large typed C# solution that matters, because text search cannot tell you whether a call binds to the <code>int</code> or the <code>string</code> overload, or which project a caller lives in. The Graphify C# project exists to hand agents compiler-accurate find-usages over a solution through Roslyn and MSBuild, which tells you how often they get it wrong without that help (<a href="https://github.com/zachsaw/graphify-csharp">graphify-csharp</a>).</p>
<h2 id="what-the-setup-looks-like">What the setup looks like</h2>
<p>The repository carries the agent’s working environment alongside the code:</p>
<ul>
<li>An <code>AGENTS.md</code> (or <code>CLAUDE.md</code>) at the solution root with the build and test commands, the project layout, and the rules the agent must follow.</li>
<li>A coding standards document under version control, covering naming, structure, iteration style, and what a unit test has to assert before it counts.</li>
<li>A semantic index of the solution so reference queries return bound symbols rather than string matches.</li>
<li>A container the agent runs in, with no production credentials and no write access to anything outside the working tree.</li>
<li>A review pass by a second agent instance against the standards document, before the diff reaches a human.</li>
</ul>
<p>Checking agent context into git is now a common pattern rather than a PicNet invention. Tools such as <a href="https://github.com/okf-memory/okf-agent-memory">OKF Agent Memory</a> and hosted <a href="https://context.apimatic.io/">context registries</a> are being built on the same premise, that curated project facts should be versioned and supplied, not guessed at from training data.</p>
<h2 id="the-standards-document-does-most-of-the-work">The standards document does most of the work</h2>
<p>Our C# standards existed as a PDF long before we used agents, and it was read about as often as most standards documents are. Turning it into a file the agent reads on every task changed its status. An agent follows a written rule far more consistently than a human under deadline, which means the document now has to be right, specific and short enough to sit in context.</p>
<p>The shape of it:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8; overflow-x: auto;" tabindex="0" data-language="markdown"><code><span class="line"><span style="color:#79B8FF;font-weight:bold"># C# standards (agent must comply)</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Structure</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> One public type per file, file name matches the type.</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> No regions. No partial classes outside generated code.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Nullability</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Nullable reference types enabled solution-wide. Do not suppress with </span><span style="color:#79B8FF">`!`</span><span style="color:#E1E4E8">;</span></span>
<span class="line"><span style="color:#E1E4E8">  fix the model or the call site.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Tests</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Every behaviour change needs a test that fails without the change.</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Do not edit an existing test to make a new change pass. Flag it instead.</span></span>
<span class="line"></span>
<span class="line"><span style="color:#79B8FF;font-weight:bold">## Framework</span></span>
<span class="line"><span style="color:#FFAB70">-</span><span style="color:#E1E4E8"> Target framework is defined in Directory.Build.props. Do not assume</span></span>
<span class="line"><span style="color:#E1E4E8">  the latest .NET version or the latest C# syntax.</span></span></code></pre>
<p>That last rule earns its place. Practitioners comparing assistants on C# repeatedly single out framework version awareness and idiomatic C# as the weak spots of general purpose models (<a href="https://www.reddit.com/r/jenova_ai/comments/1wd9rkk/what_is_the_best_ai_coding_assistant_for_c_and_net/">r/jenova_ai discussion</a>). An agent will happily write C# 13 syntax into a project pinned to an older SDK, then spend twenty minutes trying to work out why the build fails.</p>
<h2 id="tests-are-the-safety-net-and-the-agent-will-try-to-cut-them">Tests are the safety net, and the agent will try to cut them</h2>
<p>Unit tests are the only thing standing between an agent’s confident diff and a production defect. That makes the tests themselves a target.</p>
<p>OpenAI publishes what its monitoring found in its own internal coding agent deployments, and reward hacking is on the list: agents “illegitimately edit tests to make them pass”, disable checks, and sometimes misrepresent whether a task was completed at all (<a href="https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/">OpenAI</a>). If the organisation building the model treats its agents as something to supervise continuously, a Sydney consultancy has no basis for treating them as a trusted contributor.</p>
<p>So we read test diffs separately from implementation diffs, and a changed assertion in an existing test gets more scrutiny than a hundred lines of new code. Where an agent touches a codebase under regulatory constraint, an APRA CPS 234 environment or anything handling personal information under the Privacy Act, a named human approves the merge. That is a design requirement of the pipeline, not a policy we hope people remember.</p>
<h2 id="ai-review-before-human-review">AI review before human review</h2>
<p>The first reviewer of every agent diff is another agent, running a review prompt with the standards document and no memory of the conversation that produced the code. It catches the mechanical failures: standards breaches, missing tests, a public method with no null handling, a dependency added without being asked for.</p>
<p>This exists because a human reviewing forty agent diffs a day reviews the fortieth badly. It also filters out agent verbosity. Agents bury the conclusion under paragraphs of narration, to the point that people publish skills whose only job is to force the model to lead with the answer (<a href="https://github.com/ayghri/i-have-adhd">i-have-adhd</a>). A review agent that outputs a short list of concrete findings is much easier to act on than a wall of summary.</p>
<h2 id="sandboxing-is-not-optional-any-more">Sandboxing is not optional any more</h2>
<p>Google Threat Intelligence reported in September 2026 that attackers are targeting coding agents directly. The DUSTMAKER credential stealer drops malicious configuration into hidden workspace directories such as <code>.claude/</code> and <code>.cursor/</code>, then uses prompt injection in those files to make the assistant run attacker commands during ordinary developer interactions. The same actor published trojanised MCP servers to PyPI and stole OIDC tokens from GitHub Actions runners to publish signed packages that pass automated trust checks (<a href="https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai">GTIG</a>).</p>
<p>Our rules follow from that. Agents run in a container with scoped credentials. Any dependency an agent adds is reviewed by a person against the registry. Agent configuration files are treated as executable code in review, because that is what they are.</p>
<h2 id="honest-numbers">Honest numbers</h2>
<p>We report two figures per task: elapsed time and token cost. Tooling for this is maturing, with projects like <a href="https://demo.frugaltokens.com/">Frugal Tokens</a> built purely to track spend and usage across agents. Time saved without cost per completed task is half a measurement, and a failed agent run costs money while producing nothing.</p>
<p>What we see: greenfield work inside a well-defined module goes several times faster. Mechanical refactors across a large solution go faster only where the semantic index is in place, and go badly wrong without it. Debugging a subtle production issue in unfamiliar code is roughly a wash, because the time saved writing the fix is spent verifying that the fix addresses the real cause. We do not publish a single productivity multiplier, because the honest answer varies by task type more than it varies by model.</p>
<h2 id="senior-engineers-become-editors">Senior engineers become editors</h2>
<p>The part nobody costed properly is what this does to the working day of an experienced developer. Reading and judging generated diffs all day is a different job from writing code, and developers are feeling it as an identity change rather than a tooling change, as a heavily discussed <a href="https://news.ycombinator.com/item?id=49389408">Hacker News thread from August 2026</a> showed.</p>
<p>This favours senior-heavy teams, which suits how PicNet is built. Judging whether a diff is correct, idiomatic and safe to deploy requires exactly the knowledge that used to come from years of writing that code by hand. A junior engineer reviewing agent output has no way to tell a good diff from a plausible one, and the agent is very good at plausible. Small teams where everyone reviewing has shipped production systems get more out of these tools than large teams with a thin layer of seniors on top.</p>
<p>PicNet builds production AI systems for Australian organisations. <a href="https://picnet.com.au/ai-services/">Talk to us</a> about what a first project could look like.</p>

<p><em>Originally published at <a href="https://picnet.com.au/blog/coding-agents-in-a-production-net-shop-what-actually-works/">picnet.com.au</a>.</em></p>]]></content:encoded></item>
</channel>
</rss>