
Cloud cost control as code: automated spend reports, stale resources and off-hours shutdowns
Most cloud bills grow slowly. A test VM runs through a long weekend, or a disk outlives the server it was attached to, and nobody notices until the invoice arrives. The fast blowouts are rarer and much worse. In September a Reddit user reported an $87K Google Cloud bill for Gemini use, and people in the thread traced it to a phishing ad. Commenters said it was “not possible to place any limit” on the account.
We handle both kinds of blowout with scheduled code that runs against the AWS and Azure environments we look after. A spend report with anomaly flags goes out every morning. Two other jobs run beside it: a stale-resource sweep and an evening shutdown of anything tagged as temporary. This post is part of our Practical AI in DevOps series. Most of what follows is ordinary scripting, and the LLM has one narrow job at the end, which is to explain a spend spike in plain English.
Start with yesterday’s numbers
Both clouds expose billing data through an API. On AWS it’s Cost Explorer, and on Azure it’s the Cost Management query API. Each morning the report job pulls the previous day’s cost, grouped by account or subscription and by service, and writes the rows to a small table. Convert the currency when you load the data so the AWS and Azure figures can sit in one column.
Billing data arrives late on both clouds and keeps changing for a while after that. Treat the report as an accurate view of yesterday, not a live alarm, and re-pull the previous few days on every run so the late adjustments get picked up.
Deciding what counts as an anomaly takes arithmetic, not AI. Compare yesterday with the median of the same weekday over the previous six weeks. Flag it only when the increase clears both a percentage and a dollar floor:
import statistics
def is_anomaly(yesterday_aud, same_weekday_history, pct=0.4, floor_aud=50):
baseline = statistics.median(same_weekday_history)
delta = yesterday_aud - baseline
return delta > floor_aud and delta > baseline * pct
The dollar floor means a $3 service that doubles to $6 doesn’t wake anyone up. Using the same weekday as the baseline means a normal Monday doesn’t look like a spike after a quiet weekend. Tune both thresholds for each account.
Run the check for each account or subscription, and again on the total across all of them. Commenters in the Gemini thread said the attackers got around per-project throttling by “creating dozens of projects”. Each project can stay under its own threshold while the total climbs, so only the summed check catches that pattern. Count resources as well as dollars. If four running GPU instances become forty, the resource count shows it hours before the billing data does.
Turn on AWS Budgets and Azure budgets as well. By default they send an email and leave everything running. The Gemini case shows the same gap: per-project throttling looked like a spend cap, but it didn’t stop the spend.
Finding stale resources
The stale-resource sweep looks for things that cost money and have no obvious owner or purpose:
- EBS volumes and Azure managed disks with nothing attached
- snapshots older than the retention policy allows
- public IP addresses not associated with any resource
- Azure VMs that are stopped but not deallocated, which still bill for compute
- anything without an
ownertag
The sweep only reports. A disk with no VM is usually a leftover, but now and then it’s the only copy of something someone needs. Each finding goes into the morning report with its monthly cost and the identity that created it. That identity comes from CloudTrail or the Azure Activity Log, both of which keep about 90 days of history by default, so older resources often show no creator. When the list names each item and its price, the owners can clean up quickly.
Shutting down temporary resources every evening
Dev and test VMs carry a tag such as lifecycle=temporary. At 7pm Sydney time the shutdown job stops every tagged EC2 instance and deallocates every tagged Azure VM. On Azure, deallocating matters: a VM that’s only stopped keeps billing for compute.
The savings are easy to work out. A week has 168 hours. A VM that runs 7am to 7pm on weekdays is up for 60 of them, so its on-demand compute charge drops by about 64% compared with running all week. That figure has limits. Attached disks and static IPs keep billing while the VM is off. Instances covered by reserved instances or a savings plan save little or nothing when stopped, because you pay for the commitment either way.
Exceptions use a second tag, for example keep-until=2026-10-09, and the job skips that resource until the date passes. Every exception carries an expiry date, so a “just this week” request can’t turn into a permanent one.
Schedule in local time. Daylight saving in NSW starts on the first Sunday in October, and a cron expression written in UTC will run an hour off for half the year. EventBridge Scheduler and Logic Apps recurrence triggers both accept Australia/Sydney as a time zone. Give the shutdown identity permission to read tags and to stop or deallocate instances, and nothing else. It can’t create or delete anything.
Test the job before it touches a real account. Fakecloud emulates AWS locally, including EC2 and the Resource Groups Tagging API, so tag-based shutdown logic can run in CI with no credentials and no cost. Floci has emulators for both AWS and Azure. Its Azure build lists Blob storage, Functions, Key Vault and Service Bus, so check whether it covers the compute APIs your script calls before you rely on it for VM logic.
API throttling
These jobs tend to work on a small account and fail on a big one. A script that loops through every region and subscription as fast as it can will hit the management API rate limits. EC2 calls start returning RequestLimitExceeded and other AWS services return ThrottlingException. On Azure, Resource Manager returns HTTP 429 with a Retry-After header, and the Cost Management query API has its own separate limits.
The failure that costs money is a shutdown run that gets throttled halfway through. It logs an error nobody reads and leaves half the tagged machines running all night. Here’s what prevents it:
- page through results and send stop calls in batches, since EC2
StopInstancesaccepts many instance IDs per call - use the SDK’s retry handling with backoff, and honour
Retry-Afteron Azure - process accounts and subscriptions one after another rather than all at once
- schedule jobs a few minutes off the hour so they don’t collide with every other job set for 7:00
- run the shutdown a second time half an hour later, and expect it to find nothing
- name any instance that failed to stop in the next morning’s report
On AWS, the adaptive retry mode in boto3 handles most of this:
import boto3
from botocore.config import Config
ec2 = boto3.client(
"ec2",
region_name="ap-southeast-2",
config=Config(retries={"max_attempts": 10, "mode": "adaptive"}),
)
Where the LLM earns its place
By the time the LLM gets involved, code has already found the anomaly and calculated the numbers. The model receives the flagged rows, the deltas and the resources created or changed in that account the day before. From these it writes a short paragraph saying which service went up and by how much, and which resources and creators were responsible. An IT manager can read that without opening Cost Explorer or knowing what a usage type code means.
The model never adds up a total or decides what counts as an anomaly. Language models make mistakes when they sum tables, and the same input can produce a different answer tomorrow. The model also has no cloud credentials, so it can explain a spike but can’t act on one. Once it has written its paragraph, a few lines of code confirm that every dollar figure in the text appears in the input. If a figure doesn’t match, the report goes out with only the plain table.
Running costs are small. A daily explanation uses a few thousand tokens, well inside the lightest usage tier (under 1M tokens a day) in SitePoint’s 2026 cost analysis. Before sending anything to a hosted model, check whether resource names and tags contain client or project names you’d rather not send offshore. Strip them first or use a model hosted in an Australian region.
Limits of this approach
Because of billing lag, the daily report catches yesterday’s problem today. For a fast blowout like the Gemini case, resource counts and provider budget alerts react sooner, but nothing here will stop someone using a stolen credential. That needs separate controls on identity and access. Tag-based shutdown only works if people tag their resources. Anything untagged ends up in the stale sweep, which reports but never acts. If most of your compute runs on reserved capacity, the evening shutdown will save less than cleaning up what the stale sweep finds.
PicNet builds production AI systems for Australian organisations. If you’d like daily spend reports and evening shutdowns like these running across your AWS and Azure accounts, talk to us about what a first project could look like.