<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Geek Consulting</title><description>New projects and build logs from Tim Moore at Geek Consulting.</description><link>https://geekconsulting.au/</link><language>en-au</language><item><title>My First Claude Code Mod: A Task Board the Agent Built Without Ever Seeing It</title><link>https://geekconsulting.au/writing/first-claude-code-mod/</link><guid isPermaLink="true">https://geekconsulting.au/writing/first-claude-code-mod/</guid><description>The agent that built this panel has never looked at it. For most of the evening, neither of us could say why every button needed pressing twice.</description><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/first-claude-code-mod/cover.jpg&quot; alt=&quot;A cork board of sticky notes on a study wall at night. Ribbons of light carry the notes across the room to a side panel on a monitor, where they land as three cards in green, amber and blue.&quot;&gt;&lt;/figure&gt;&lt;h2 id=&quot;tldr&quot;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;I keep a task board for everything the house and my side projects need doing, and the AI agent I work with keeps it up to date as we go. I read the board but I don’t work it, the agent does, and reading it meant switching to a browser, so the work and the record of the work were always in different windows.&lt;/p&gt;
&lt;p&gt;Now the board sits in a panel beside the conversation. It shows what’s in progress, what’s stuck and what’s waiting. When the agent picks up a task, the panel opens that task by itself and updates while it works. There’s a button to start a fresh session on any task, in the right folder.&lt;/p&gt;
&lt;p&gt;It took an evening and the morning after. Most of it works and I use it. Two parts have only been tested against a pretend version of the board, and I say which further down. It’s on GitHub, and getting it there taught me something about what “scrubbed” means.&lt;/p&gt;
&lt;p&gt;If you’re here for the code, it’s at the bottom.&lt;/p&gt;
&lt;h2 id=&quot;the-board-in-the-other-window&quot;&gt;The board in the other window&lt;/h2&gt;
&lt;p&gt;The tracker is &lt;a href=&quot;https://vikunja.io/&quot;&gt;Vikunja&lt;/a&gt;, self-hosted, and it’s the one piece of the homelab that’s about the work rather than the house. Every job gets a card. The agent creates them, moves them to Doing when it starts, writes what it built into the description when it finishes. The Done column is the history of the whole lab.&lt;/p&gt;
&lt;p&gt;I look at it. I don’t touch it. The agent does that.&lt;/p&gt;
&lt;p&gt;Which is how I want it, but it put the board in the wrong place. It’s a browser tab, and the work happens in Claude Code. When the agent said “moved to Doing”, seeing that meant leaving the conversation and finding the card. I’m the reader of this board, not its editor, and the reading was happening in a different window from the thing I was reading about.&lt;/p&gt;
&lt;p&gt;Claude Code recently grew a way to fix that: &lt;strong&gt;mods&lt;/strong&gt;. A mod is a small plugin of hooks that runs inside the session and can draw its own panel next to the conversation. I asked for one in a single sentence: &lt;em&gt;let’s create a mod to see our tasks when we’re working on them.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I’ve written about building things this way before: &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;a data platform&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-2/&quot;&gt;a dashboard&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-3/&quot;&gt;a car alert&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/spa-heats-on-spare-sunshine/&quot;&gt;a spa that heats on spare sunshine&lt;/a&gt;. Same principle every time: &lt;strong&gt;I supply direction, the agent supplies execution.&lt;/strong&gt; This one tested the first half of that sentence harder than usual, because the thing being built was a user interface, and only one of us could see it.&lt;/p&gt;
&lt;h2 id=&quot;one-sentence-and-a-question-about-lanes&quot;&gt;One sentence, and a question about lanes&lt;/h2&gt;
&lt;p&gt;Before it wrote a line, the agent went and asked the tracker some questions, read-only. 32 projects. 179 open tasks. And one awkward fact: ask Vikunja for a list of tasks and every one comes back saying it’s in bucket zero. The lane a card sits in is only reported when you read a project’s board, one project at a time.&lt;/p&gt;
&lt;p&gt;So it read one board to see what that cost. &lt;strong&gt;1.2 megabytes&lt;/strong&gt;, for a single project, because every card comes with its full description. Thirty-two of those a minute was not a design.&lt;/p&gt;
&lt;p&gt;It tried the other door instead: would the task list accept a &lt;em&gt;filter&lt;/em&gt; on a bucket it refuses to report? It would. One small call per project to learn which bucket is called Doing and which is Blocked, cached for half an hour, then one filtered query per lane.&lt;/p&gt;
&lt;p&gt;That’s the pattern of the whole evening in miniature. It didn’t assume the sensible-looking route worked. It measured it, found it was absurd, and went looking for a better one before building on either.&lt;/p&gt;
&lt;p&gt;The first version was up a few minutes later: a panel with two lists, and a section at the top for any task the session had touched.&lt;/p&gt;
&lt;h2 id=&quot;every-button-needed-two-clicks&quot;&gt;Every button needed two clicks&lt;/h2&gt;
&lt;p&gt;I clicked a task and it opened in the browser, which was the exact thing I was trying to get away from. That got fixed quickly: click a task, the panel shows the task, a Back button returns to the list.&lt;/p&gt;
&lt;p&gt;Then I told it the Back button was unresponsive. It took a few goes to work.&lt;/p&gt;
&lt;p&gt;The agent had a theory, and it was a reasonable one. Back and the background refresh were both writing to the same stored value, so a press landing mid-refresh could be overwritten. It rewrote that, gave Back its own tiny value that nothing else touches, and said plainly that this was a fix for its best guess, not a confirmed cause.&lt;/p&gt;
&lt;p&gt;It didn’t fix it. My next report was more precise: I had to click &lt;em&gt;everything&lt;/em&gt; twice. Tasks, Back, all of it.&lt;/p&gt;
&lt;p&gt;Here is the part I like. The agent could not look at the panel. It tried: it asked for permission to take a screenshot of the Claude app, and the app refused, on the grounds that an agent operating its own window could change its own permissions. Fair enough. So it had no eyes on the thing it was debugging, and one failed theory behind it.&lt;/p&gt;
&lt;p&gt;It didn’t produce a second theory. It built an instrument. The mod got a temporary trace that wrote every press, every focus change and every redraw to a log file with a timestamp, and I was asked to click a task, click Back, and click another task.&lt;/p&gt;
&lt;p&gt;The log said this, 35 seconds in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;35.27 s:&lt;/strong&gt; the focus ring moves to Back. No press.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;35.95 s:&lt;/strong&gt; a press on Back.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The first click never reached the button. It handed the panel the keyboard and stopped there. The second click, 0.7 seconds later, was the press. And every time the view changed, the button that held the focus disappeared with the old view, so the panel dropped the keyboard and the next click was spent getting it back.&lt;/p&gt;
&lt;p&gt;The fix was to move the focus onto the new view’s first button whenever the view swaps. In the same log, after it went in: &lt;strong&gt;fourteen presses in a row, each one landing on a single click.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The trace came out again afterwards and the log file was deleted. One extra click when you first move from the prompt into the panel is still there, and it’s the app’s to decide who holds the keyboard.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; When you can’t see the thing, don’t guess harder. Make it write down what it’s doing, and read that.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;building-a-screen-blind&quot;&gt;Building a screen blind&lt;/h2&gt;
&lt;p&gt;“Fix the UX, it’s pretty ugly so far.”&lt;/p&gt;
&lt;p&gt;That’s a hard instruction to give something that can’t see. The agent’s answer was to stop trusting itself and build a test: mount the panel on a pretend desktop, feed it a pretend tracker with one project and a handful of tasks, press the buttons, and check what was drawn. It doesn’t show what anything &lt;em&gt;looks&lt;/em&gt; like. It does prove the app will accept the layout, that a click opens a task, and that Back comes back.&lt;/p&gt;
&lt;p&gt;It used the same trick to settle a question about colour. Rather than hope a colour name was valid, it built a throwaway mod that drew one box in nine different colour spellings and ran it. Seven were accepted, including one it had made up to see what happened. &lt;code&gt;ansi:blue&lt;/code&gt; was refused. That told it exactly how far the checking went, and it deleted the throwaway.&lt;/p&gt;
&lt;p&gt;Then the test missed something, and it missed it for an instructive reason.&lt;/p&gt;
&lt;p&gt;I’d asked for a way to start a session on a task, with a folder picker, the folders all living on my D: drive. Shortly after that went in, I opened a task and the whole panel went blank.&lt;/p&gt;
&lt;p&gt;The agent’s test still passed. So it stopped testing with pretend folders and fed the test &lt;strong&gt;my real folder list&lt;/strong&gt;. It failed at once, with the app’s own words: a dropdown takes between 1 and 64 options. My D: drive has &lt;strong&gt;80 folders&lt;/strong&gt;. One dropdown over the limit and the app declines to draw &lt;em&gt;anything&lt;/em&gt;, not just the dropdown.&lt;/p&gt;
&lt;p&gt;The picker now offers the 58 most recently changed folders and a box to type any other name. The test keeps a 90-folder case in it permanently.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; A test built on tidy pretend data proves the code works on tidy pretend data. The bug was in how much stuff I actually own.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;the-first-time-it-saw-its-own-work&quot;&gt;The first time it saw its own work&lt;/h2&gt;
&lt;p&gt;Eventually I did the obvious thing and pasted a screenshot into the chat.&lt;/p&gt;
&lt;p&gt;It was the first time the agent had seen its own work, and it found two faults in it before I’d said what I wanted. Titles were being cut off with a third of the row still empty, because it had counted characters as if they were all the same width and the panel uses a proportional font. And the titles didn’t line up on the left, because each one was being centred in its own space.&lt;/p&gt;
&lt;p&gt;I’d only asked for the priority badges to move to the end of the row.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/first-claude-code-mod/list.jpg&quot; alt=&quot;The panel&apos;s list view: filters, a search box, and three cards for Doing, Blocked and To-Do&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;Where it ended up. This is the built-in demo data, not my board; more on why further down.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;After that we did the visual work the way you’d do it with a designer: one change at a time, and I said keep or drop.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A highlight when the pointer is over a row.&lt;/strong&gt; Kept.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Priority and status as filled badges&lt;/strong&gt; instead of coloured letters. Kept, then moved to the end of the row, then made all the same width, because “P1” is narrower than “P4” and a ragged column of badges is worse than none.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A bordered card around each section.&lt;/strong&gt; Kept.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A summary bar across the top&lt;/strong&gt;, showing the proportions of the three lanes. Dropped the moment I saw it. The sections already have counts.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That last one was the agent’s own suggestion, built properly, and it was redundant. It took me about four seconds of looking to know. The agent couldn’t have known at all.&lt;/p&gt;
&lt;h2 id=&quot;scrubbed-published-and-still-showing-my-board&quot;&gt;Scrubbed, published, and still showing my board&lt;/h2&gt;
&lt;p&gt;The next morning I had it moved somewhere permanent and put on GitHub.&lt;/p&gt;
&lt;p&gt;The agent did what it always does before something goes public. It pulled the house out of the code: the server’s address, my hostname, the D: drive and the home-improvement rule all became settings. Then it searched every file in the history for addresses, hostnames, user names, emails and anything shaped like a token. Nothing. It published, and opened a pull request containing the whole repository so a review bot could scan it for personal details as a second opinion. The bot came back with three minor code findings and nothing about privacy.&lt;/p&gt;
&lt;p&gt;A while later I asked one idle question: &lt;em&gt;is my “all but home improvement” filter in the committed code?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The filter wasn’t. But the agent went and looked rather than answering from memory, and came back with something else. The test it had written to check the panel used a pretend tracker, and it had populated that pretend tracker with &lt;strong&gt;three of my real project names&lt;/strong&gt;, and real id numbers off my board. They’d been public since the first push.&lt;/p&gt;
&lt;p&gt;Its search had been for the things that grant access or locate a house. Project names are neither, so the search wasn’t looking for them, and the review bot hadn’t blinked either. They’re not secrets. They are, though, a list of what I’m working on, sitting in a public repository because nobody asked the right question.&lt;/p&gt;
&lt;p&gt;The names are invented ones now. The old ones are still in the history; I chose not to rewrite it.&lt;/p&gt;
&lt;p&gt;The same morning produced the fix for the other half of that problem. I wanted screenshots in the README, and every screenshot I’d taken showed my real tasks. Blurring them would have looked rough and risked missing one. The agent built &lt;strong&gt;a demo mode&lt;/strong&gt; instead: type &lt;code&gt;/vikunja demo&lt;/code&gt; and the panel switches to a small invented tracker held in memory, with made-up projects, people and folders. A test proves it makes no network, disk or process calls while it’s on. Both screenshots in this piece are of that.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/first-claude-code-mod/task.jpg&quot; alt=&quot;The panel&apos;s task view: a title, status and priority badges, a folder picker with a Start session button, the description and comments&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;A task, in demo mode. On my own board that folder picker lists the real folders on my drive.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; A scrub finds what you told it to look for. Decide what counts as private before you search, and include the boring things: names, numbers, the shape of your to-do list.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;smaller-traps-for-the-record&quot;&gt;Smaller traps, for the record&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The validator on my machine was too old.&lt;/strong&gt; The command-line Claude Code on my PATH was 48 builds behind the one inside the desktop app, and rejected the mod’s format outright. The agent found the app’s own copy and validated with that.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;“5 of 6”.&lt;/strong&gt; With nothing filtered, the Doing section read “5 of 6”. The sixth was a home-improvement task the panel hides by default. A count that makes you ask what’s missing, when you haven’t hidden anything yourself, is a bug. Hidden projects are now out of both numbers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vikunja has no custom fields.&lt;/strong&gt; A task’s folder is stored as a label on the task, &lt;code&gt;folder:&lt;/code&gt; and then the path. It shows up as a chip in Vikunja itself, and anything else that reads the task can see it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Updating a task replaces the whole task.&lt;/strong&gt; Send Vikunja only a new description and it resets the rest, assignees included. The mod never updates a task. It only adds or removes that one label.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The mod can make a folder only by writing a file into it.&lt;/strong&gt; So a folder created from the panel starts life with one empty &lt;code&gt;.gitkeep&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Git Bash ate a slash.&lt;/strong&gt; Checking that a fresh session loaded the mod, the agent ran the &lt;code&gt;/vikunja&lt;/code&gt; command from a shell that helpfully rewrote it into a Windows path. The fresh session received a file path, shrugged, and did something else entirely. Twice, before the shell was told to leave it alone.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A clear button for a toggle is silly.&lt;/strong&gt; A Show/Hide switch displays its own state. It doesn’t need a Clear next to it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The wrong number.&lt;/strong&gt; Every task has two: the one on its card, and an internal one that only appears in its web address. The agent had been naming tasks by the internal one, which I can’t find on the board. It now uses the number I can see, and that’s written down where every future session will read it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A link can’t have a keyboard shortcut.&lt;/strong&gt; So “Open in Vikunja” became a button that asks the operating system to open the page. The app then drew its own little key badge beside it, which made the “(o)” in the label redundant within the minute.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One review an hour.&lt;/strong&gt; The review bot’s free allowance stalled two pull requests with every finding fixed and nothing left to approve them. It’s been replaced with a secret scanner and a review that runs on the Claude subscription I already pay for.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-it-looks-like-now&quot;&gt;What it looks like now&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Three lanes&lt;/strong&gt; in cards: Doing, Blocked, To-Do, grouped by project, with a priority badge and who it’s assigned to. To-Do stops at 80 rows and tells you how many more there are.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filters&lt;/strong&gt; for who, which project, what priority and what’s due (overdue, this week), and a Show/Hide for Blocked.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find #&lt;/strong&gt;: type the number on a card and it finds the task in any project or lane, finished ones included.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keyboard shortcuts&lt;/strong&gt;: &lt;code&gt;r&lt;/code&gt; to refresh, &lt;code&gt;o&lt;/code&gt; to open the task in the browser, &lt;code&gt;b&lt;/code&gt; to go back.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A demo mode&lt;/strong&gt; with an invented board, for screenshots and for trying it without a server.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Home-improvement projects hidden by default&lt;/strong&gt;, because this panel sits next to code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The task the session is on opens by itself&lt;/strong&gt;, and is re-read every 10 seconds. It only redraws when something changed, so it doesn’t fight your scrolling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A Start session button&lt;/strong&gt; on every task, with its folder.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;About 1,200 lines&lt;/strong&gt; in the main file, a test that runs the panel on two surfaces, and three settings so it isn’t welded to my house.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-honest-part&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;The other pieces ended on what unblocked the thing: the tedious middle could be handed off, or the earlier projects already existed, or the brief was wrong and something that could see the house said so. This one’s different, because nothing was blocked. I asked for it in a sentence and had a working panel in minutes.&lt;/p&gt;
&lt;p&gt;What I noticed instead was &lt;strong&gt;which of my contributions mattered.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I didn’t supply a design. I didn’t know the plugin API existed in any detail, and I couldn’t have told you a dropdown stops at 64. Every technical decision in this, the filtered query, the label instead of a field, the focus fix, was the agent’s, and each one was checked against the real system before it was built on.&lt;/p&gt;
&lt;p&gt;What I supplied was eyes. “It opens in the browser.” “I have to click everything twice.” “It’s pretty ugly.” “Put the badges last.” “That bar is redundant.” “Is my filter in the committed code?” None of those is a specification. Every one of them is something only a person looking at the screen could say, and every one moved the thing further than the paragraph of requirements I might have written instead.&lt;/p&gt;
&lt;p&gt;And the agent’s side of that bargain was knowing what it couldn’t see. It said so each time: &lt;em&gt;I haven’t seen it draw. This is a fix for my best guess. I’m least sure about the borders.&lt;/em&gt; When it was refused a screenshot it didn’t pretend; it built a trace, then a test, then asked me for a picture. I’d take that over confidence.&lt;/p&gt;
&lt;p&gt;Now the part that isn’t finished. Two things in the mod have only ever run against the pretend tracker: &lt;strong&gt;writing the folder label to the real one&lt;/strong&gt;, and &lt;strong&gt;Start session&lt;/strong&gt;, which hands off to the desktop app to offer a new session. The test says both do what they should. I’ve used the lanes, the filters, the search and the shortcuts on my real board. I haven’t pressed either of those two buttons in anger. And the new session probably starts in a fresh working copy of the folder, which will suit a git repo and may well fail on a plain one. I don’t know yet.&lt;/p&gt;
&lt;p&gt;There’s also a small, pleasing loop in what this is. The panel exists so I can watch the agent work through the board. The agent built the panel. It was told how the panel looked by the one participant who could see it, and it now sits beside the conversation.&lt;/p&gt;
&lt;p&gt;Build the tool that lets you watch the work. Then be the one who looks.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;under-the-hood-if-you-want-to-build-one&quot;&gt;Under the hood: if you want to build one&lt;/h2&gt;
&lt;h3 id=&quot;what-it-runs-on&quot;&gt;What it runs on&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://code.claude.com/docs/en/plugins&quot;&gt;Claude Code&lt;/a&gt; 2.1.286, desktop app. Mods are plugins of &lt;em&gt;function hooks&lt;/em&gt;: one TypeScript module that hooks session events and draws panels. The API is early access and moves between releases.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://vikunja.io/&quot;&gt;Vikunja&lt;/a&gt; v2.4.0, self-hosted. Read over its REST API with a token from the environment.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/netadvanced/vikunja-mcp-ng&quot;&gt;vikunja-mcp-ng&lt;/a&gt; v0.6.0: the MCP server the agent uses to work the board. The mod watches for its tool calls to know which task the session is on.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;decisions&quot;&gt;Decisions&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Read the tracker directly, not through the agent.&lt;/strong&gt; The panel polls on a timer and costs no model calls.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find lanes by bucket name, then filter tasks by bucket id.&lt;/strong&gt; A task list reports bucket 0 for everything, but accepts &lt;code&gt;bucket_id in …&lt;/code&gt; as a filter. Reading whole project boards was 1.2 MB a project.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The folder is a label, not a field.&lt;/strong&gt; Vikunja has no custom fields; &lt;code&gt;folder: &amp;lt;path&amp;gt;&lt;/code&gt; is visible in its own UI and to every other client.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Never send a task update.&lt;/strong&gt; It replaces the whole task. Add and remove a label instead.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Write state only when it changed.&lt;/strong&gt; A write redraws the panel; redrawing every 10 seconds fights scrolling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Site details are settings.&lt;/strong&gt; API address and token from the environment; web address, folder root and the hidden project tree from the plugin’s options.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A demo mode instead of blurred screenshots.&lt;/strong&gt; An in-memory tracker with invented data, proven by test to make no outside calls.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Search by the card’s number, not the internal id.&lt;/strong&gt; The number is per project, so one number can list several tasks; the project name tells them apart.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rejected: a custom-drawn board with drag and drop.&lt;/strong&gt; Possible, and a much bigger build than a list you can click.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;gotchas&quot;&gt;Gotchas&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The first click on a panel only gives it the keyboard.&lt;/strong&gt; And a view change drops it again unless you move the focus to an element in the new view.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A dropdown takes 64 options at most.&lt;/strong&gt; One more and the entire tree is refused; the panel goes blank.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The desktop draws text in a proportional font.&lt;/strong&gt; Character-count truncation runs short. Cut generously and let the box clip.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Badges of equal width need a fixed-width box&lt;/strong&gt;, with the label centred in it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Colour values aren’t checked against a palette.&lt;/strong&gt; Any word passes validation; only the format is refused.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The agent can’t screenshot the app it runs in.&lt;/strong&gt; Budget for a trace, a mount test, and a person.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Check which Claude Code your shell finds.&lt;/strong&gt; An older CLI rejects the plugin outright.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A mod makes directories only on the way to writing a file.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A Link can’t carry a hotkey.&lt;/strong&gt; Use a Button and open the page through the host. Don’t put the key in the label: the desktop draws its own badge.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The focus ring is drawn outside its button&lt;/strong&gt;, and the pane’s edge clips it. Leave a cell of padding on the top row.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scrub test data too.&lt;/strong&gt; Fixtures are where real names hide.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;the-focus-fix&quot;&gt;The focus fix&lt;/h3&gt;
&lt;p&gt;The few lines that turned two clicks into one. &lt;code&gt;PANE&lt;/code&gt; is the pane’s id; the key is whichever button the next view leads with.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;ts&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6A737D&quot;&gt;// Denied while the prompt holds the keyboard, which is the person&apos;s to give.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;const&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; keepFocus&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt; =&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; (&lt;/span&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;$&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;key&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;) &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  $.ui.&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;focus&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;({ requestId: &lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;PANE&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, key }).&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;catch&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(() &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&amp;gt;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; {})&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6A737D&quot;&gt;// opening a task&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;await&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; update&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;($, openId, () &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&amp;gt;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; task.id)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;void&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; keepFocus&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;($, &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&apos;back&apos;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6A737D&quot;&gt;// going back to the lists&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;await&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; update&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;($, openId, () &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&amp;gt;&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt; 0&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;void&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; keepFocus&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;($, &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&apos;refresh&apos;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;take-it-with-you&quot;&gt;Take it with you&lt;/h3&gt;
&lt;p&gt;The whole mod is one repository: &lt;a href=&quot;https://github.com/angusmaul/claude-code-vikunja-tasks&quot;&gt;angusmaul/claude-code-vikunja-tasks&lt;/a&gt;. MIT. It needs a Vikunja address and token in the environment, and takes three optional settings. The README says what it reads, the two things it writes, and how to run the validator and the test.&lt;/p&gt;
&lt;p&gt;Used by hand on a real board, in the desktop app on Windows: the lanes, filters, search, task detail, auto-open, live refresh and keyboard shortcuts. Proven only against a fake tracker: saving a folder label, creating a folder, and Start session. Not run at all: the terminal, macOS, Linux.&lt;/p&gt;
</content:encoded><category>claude-code</category><category>ai-agents-in-action</category><category>homelab</category><category>self-hosted</category><category>productivity</category><category>Writing</category></item><item><title>Project: Vikunja tasks for Claude Code</title><link>https://geekconsulting.au/projects/claude-code-vikunja-tasks/</link><guid isPermaLink="true">https://geekconsulting.au/projects/claude-code-vikunja-tasks/</guid><description>A Claude Code mod that puts a Vikunja task board in a panel beside the conversation and opens the task the agent is working on.</description><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>Project</category><category>Open source</category></item><item><title>I Got Into the Spa an Architect and Got Out a Scrum Master. Again.</title><link>https://geekconsulting.au/writing/human-in-the-loop-scrum-master/</link><guid isPermaLink="true">https://geekconsulting.au/writing/human-in-the-loop-scrum-master/</guid><description>After a year of working with an AI agent, the job title I&apos;d give the human in the loop is one I last held fourteen years ago, and I think it only works the way I did it then: by someone who&apos;s also doing real work.</description><pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/human-in-the-loop-scrum-master/1.jpg&quot; alt=&quot;A man at a kitchen table with a laptop looks up at a task board of sticky notes, while forms of blue light hover over the empty chairs. A spa steams in the backyard beyond the glass door.&quot;&gt;&lt;/figure&gt;&lt;p&gt;&lt;em&gt;I had two thoughts in the spa yesterday. The first was that I’m content, which is not a word I’d normally use about work. The second was that I’m a scrum master again, fourteen years after the last time.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;This one isn’t a build log. Nothing broke, nothing shipped, and there’s no config at the bottom. It’s a retrospective on a year of working with Claude more and more, written the day after sitting in &lt;a href=&quot;https://geekconsulting.au/writing/spa-heats-on-spare-sunshine/&quot;&gt;a spa that an AI agent taught to heat itself on spare sunshine&lt;/a&gt;, which is as good a place as any to wonder what my job has become.&lt;/p&gt;
&lt;h2 id=&quot;the-contentment-first&quot;&gt;The contentment first&lt;/h2&gt;
&lt;p&gt;For most of my working life an idea had three possible fates. It got done, which cost an evening or a weekend. It went on a list, which is where ideas go to be felt guilty about; I wrote &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;three articles&lt;/a&gt; about a list like that. Or, most often, it never got as far as the list. Too many unknowns, too much effort, too much complexity, or it needed a language I didn’t write and skills I’d have to go and learn first. Those ideas didn’t wait. They were discarded on the spot and forgotten by the end of the day.&lt;/p&gt;
&lt;p&gt;That third pile is the seed of the contentment, because it’s the one that no longer exists. &lt;a href=&quot;https://geekconsulting.au/writing/salesforce-architect-rust-prs/&quot;&gt;A connector written in Rust&lt;/a&gt; would have gone straight into it a year ago.&lt;/p&gt;
&lt;p&gt;What’s changed is that a small idea now costs a prompt. Not a prompt that builds the thing, usually. A prompt that goes and looks: what’s already there, what the options are, what it would take, and what I’ve got wrong about my own setup. The spa piece is the clearest case I have. The first task in that brief was to discover the current state, read-only, and the first thing discovery found was the Powerwall I’d forgotten to mention.&lt;/p&gt;
&lt;p&gt;So every idea gets a hearing. Some come back as “that’s an afternoon”, some as “that’s a bad idea, and here’s why”, and both answers are worth having. What’s left on the list is there because I decided so, not because I ran out of evenings.&lt;/p&gt;
&lt;p&gt;And sometimes the answer is “yes, and sooner than you think”. The best example I have is last week’s. My 3D printing business had a website that was friendly right up to the quote form, and after that it was an email to me, with every quote, photo, invoice and tracking number handled by hand from whichever inbox the last message landed in. That had been true for as long as the site existed. &lt;a href=&quot;https://geekconsulting.au/writing/order-portal-in-four-days/&quot;&gt;Four days later&lt;/a&gt;, and those four days were a couple of evenings and a few hours on the weekend, it was an order management system: every enquiry an order with its own page, customers signing in by emailed link, quotes accepted with a tick-box, postage and payment built in. Eight milestones, each a pull request, and the first real order went all the way through on launch day.&lt;/p&gt;
&lt;p&gt;I’d have put that project at months, which is to say never.&lt;/p&gt;
&lt;p&gt;That’s the contentment. It isn’t the thrill of going fast. It’s that nothing gets thrown away for being too hard any more. Every one of those ideas can be acted on now, and the only limits left are my own: how much I can hold in my head at once, and how much I want it.&lt;/p&gt;
&lt;h2 id=&quot;then-the-job-title&quot;&gt;Then the job title&lt;/h2&gt;
&lt;p&gt;In July I &lt;a href=&quot;https://geekconsulting.au/writing/salesforce-architect-rust-prs/&quot;&gt;wrote&lt;/a&gt; that when I work with the agent, I’m the product manager, architect and QA, and the agent is the engineering team. I still think that’s true. I also think it left out most of my day.&lt;/p&gt;
&lt;p&gt;Because when I look at what I do between the deciding and the checking, it’s this. I find out why the work has stopped. I log in to the thing the agent can’t log in to. I approve the permission, authorise the connector, restart the service that needs a human at the keyboard. I go and talk to the person who has to agree, whether that’s an open-source maintainer or someone in my own house. I notice the same mistake has happened twice and write the fix down where the next session will read it. I keep things moving.&lt;/p&gt;
&lt;p&gt;None of that is product management or architecture. It’s clearing the path so a team can keep working, and I know the name for the person who does that, because I used to be one. Fourteen years ago, when my first Salesforce project went live, I was its scrum master.&lt;/p&gt;
&lt;p&gt;I wasn’t hired as one. I was voluntold. I’d been the business subject matter expert through the implementation, which meant I knew what the platform was supposed to do for the people using it, and I’d picked up a lot of Salesforce along the way. I’d been a certified admin for about a year, and I was the only one on the team. The platform had gone live, I was moving out of a business and project role into IT, and the team that would deliver on it from then on needed a scrum master. So I did that alongside the admin work and the business questions, because those were my job too.&lt;/p&gt;
&lt;p&gt;It’s an uncomfortably good match for now. Then, as now, I was the one who knew what the thing was for, doing real work on the platform and keeping everyone else’s work moving in the gaps.&lt;/p&gt;
&lt;h2 id=&quot;has-anyone-said-this-already&quot;&gt;Has anyone said this already?&lt;/h2&gt;
&lt;p&gt;I went looking, expecting to find the idea worn smooth. It’s there, but not quite in this shape. What I found falls into three camps.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The developers say orchestrator.&lt;/strong&gt; The popular framing is Addy Osmani’s &lt;a href=&quot;https://addyosmani.com/blog/future-agentic-coding/&quot;&gt;conductor and orchestrator&lt;/a&gt;: a conductor steers one agent in real time, an orchestrator hands work to several and reviews what comes back. His comparison for the second is a tech lead delegating to developers and reading their pull requests. I don’t think that’s a rival to my word. Orchestrating is a fair description of half of what a scrum master does. It’s the half you can see. The other half is the unblocking, and that’s the half the metaphor leaves out.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Scrum people are writing about the scrum master as a separate person.&lt;/strong&gt; A &lt;a href=&quot;https://www.scrum.org/resources/blog/ai-augmented-scrum-framework-when-half-your-team-autonomous-agents&quot;&gt;Scrum.org piece from March&lt;/a&gt; describes a team that’s half agents, with the scrum master watching logs and rate limits so the bots stay unblocked. IBM &lt;a href=&quot;https://www.ibm.com/think/perspectives/scrum-masters-hidden-job&quot;&gt;argued last month&lt;/a&gt; that agents should take over the status-chasing that has quietly eaten the role, leaving the judgement to the human. Both are about what happens to someone whose whole job is Scrum Master. I’ll come back to why I think that’s the wrong starting point.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The framework builders hand the role to a bot.&lt;/strong&gt; There are several projects that run a set of agents as a Scrum team, and the scrum master is usually one of the agents. The nearest thing to my thought is in Michael Bleterman’s &lt;a href=&quot;https://engineeringexec.tech/posts/ai-scrum-can-proven-agile-principles-work-for-agent-teams/&quot;&gt;write-up of AI-Scrum&lt;/a&gt;, which says the human ends up as the “Scrum Master of a virtual team”. That’s very close, and it’s the same phrase I landed on. Where we differ is where the human stands. In his design the agents fill every seat on the team, product manager and QA included, and the human sits outside it: setting the direction for the sprint, reviewing what comes back at the end, adjusting the guardrails. That’s a manager of a team. I’m on the team. I write the requirements, I test what gets built against a real system, and I do the scrum master’s job in between.&lt;/p&gt;
&lt;p&gt;So: raised, and by people who’ve thought about it harder than I did in a spa. What I couldn’t find is anyone arguing it from the seat I’m in, which is one person, one agent or a few, and no org chart anywhere in sight.&lt;/p&gt;
&lt;h2 id=&quot;where-it-fits&quot;&gt;Where it fits&lt;/h2&gt;
&lt;p&gt;The &lt;a href=&quot;https://scrumguides.org/scrum-guide.html&quot;&gt;Scrum Guide&lt;/a&gt; is short and worth rereading with an agent in mind. Four things describe my year better than I expected.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A scrum master who also does the work.&lt;/strong&gt; I’ve never believed in the dedicated scrum master. The role works when it’s held by a member of the team who gets real work done too, and it goes wrong when it becomes somebody’s entire job. That’s what IBM’s piece reads like to me: a description of a role that filled its empty hours with ticket-tidying. The human in the loop can’t go that way. I’m specifying, testing and deciding, which is real work by anyone’s measure, and the scrum master part happens in between because somebody on the team has to do it and I’m the only one who can.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impediments.&lt;/strong&gt; The scrum master is the one who causes impediments to be removed. With an agent, almost everything that stops the work is something only I can remove: a credential, an approval, a decision, a cable. On the Rust contribution, the Windows toolchain ate an afternoon, and no amount of capable code generation was going to install it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Serving, not assigning.&lt;/strong&gt; A scrum master isn’t the team’s boss. The Guide calls them leaders who serve. I don’t tell the agent how to do the work, and the sessions go worse when I try. I make sure it can see the system, has the context, and knows what done looks like.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The retrospective.&lt;/strong&gt; In Scrum the team stops at the end of each sprint to ask what went well, what didn’t, and what to change. With human teams the actions have a way of evaporating by the next sprint, because people carry the lesson in their heads and feel that’s enough. An agent carries nothing in its head overnight. If the lesson isn’t written down where the next session reads it, it didn’t happen. The lessons in my instruction files mostly have a date on them and an incident behind them, which makes them the most rigorous retrospectives I’ve ever kept.&lt;/p&gt;
&lt;h2 id=&quot;where-it-creaks&quot;&gt;Where it creaks&lt;/h2&gt;
&lt;p&gt;Honest version. The analogy strains in two places.&lt;/p&gt;
&lt;p&gt;The Guide gives the requirements to the product owner, not the scrum master. When I said my job included managing stakeholders and requirements, that was the product owner’s half. I hold both, and on a team of people that’s a known hazard: there’s nobody to push back when the one who wants the feature and the one protecting the team are the same person. I haven’t solved that. I’ve mostly been lucky that the agent pushes back for me, as it did over the Powerwall.&lt;/p&gt;
&lt;p&gt;And a scrum master’s team is supposed to get better at running itself. Mine starts from zero every morning. I’m not coaching a team towards independence. I’m its memory.&lt;/p&gt;
&lt;h2 id=&quot;dont-give-the-job-to-the-agent&quot;&gt;Don’t give the job to the agent&lt;/h2&gt;
&lt;p&gt;The full-time scrum master has mostly gone. Capital One removed its whole agile job family in January 2023, about 1,100 roles, &lt;a href=&quot;https://big-agile.com/blog/why-the-scrum-master-role-is-being-cut-and-whats-actually-replacing-it&quot;&gt;saying the work would be folded into engineering&lt;/a&gt;. One writer’s &lt;a href=&quot;https://www.vibhorchandel.com/p/the-scrum-master-job-still-exists&quot;&gt;check of job listings this August&lt;/a&gt; found the standalone title close to absent in Toronto and London, bundled with other skills in the United States, and surviving mainly in the big Indian consultancies where it’s a billable line. I don’t mourn it. That’s the version of the role I never believed in.&lt;/p&gt;
&lt;p&gt;But there’s a tempting next step, and the frameworks I found have nearly all taken it: if nobody’s job is scrum master any more, make it an agent’s. I think that’s a mistake, and it’s the same mistake in a new place.&lt;/p&gt;
&lt;p&gt;An agent can do the parts of the job that were never the point. It can chase status, tidy the board, notice that something has been stuck for two days and say so. Let it. What it can’t do is the thing the role exists for. The impediments worth the name need someone with access, authority and standing: a credential only I hold, a decision only I can make, a conversation with a person who needs to hear it from a person. An agent scrum master can tell you the team is blocked. It can’t unblock it.&lt;/p&gt;
&lt;p&gt;That was always the critical piece. The scrum master was there to unblock the team and keep it productive, and everything else was housekeeping. Most of what the full-time version of the role fussed over was exactly that: housekeeping, done at length, to justify the position. It’s also, as far as I can tell, the plainest answer to why the humans are still needed. For now.&lt;/p&gt;
&lt;p&gt;So the role didn’t disappear. It moved to whoever is in the loop. Call it the agentic scrum master if it needs a name: a human on the team, doing real work, whose team happens to be agents. In my experience that’s the arrangement that gets results, and it’s the one the role should have been all along.&lt;/p&gt;
&lt;h2 id=&quot;why-content-and-not-fried&quot;&gt;Why content, and not fried&lt;/h2&gt;
&lt;p&gt;Not everyone doing this is happy. Boston Consulting Group surveyed about 1,500 workers this year and &lt;a href=&quot;https://builtin.com/articles/ai-brain-fry-software-developers&quot;&gt;found 14 per cent reporting a kind of mental hangover&lt;/a&gt; from working with AI tools, with productivity turning down once people took on a fourth agent. Flavio Copes, who has shipped a remarkable amount this way, &lt;a href=&quot;https://flaviocopes.com/ai-and-the-joy-of-programming/&quot;&gt;wrote last month&lt;/a&gt; that the small satisfactions of building things by hand have gone, and that the joy has moved to deciding what should exist.&lt;/p&gt;
&lt;p&gt;I recognise both. I usually have two or three sessions going at once, and that’s comfortable. I can push it to six, and at six I hit exactly what that survey describes: I’m switching context faster than I can rebuild it, and I can feel the quality of my decisions dropping. Their number and mine are close enough that I believe theirs.&lt;/p&gt;
&lt;p&gt;And the pressure is real. An agent sitting there waiting for me pulls harder than an unread email, harder than the red badge on Slack. It has stopped, it has said exactly what it needs, and the only thing between it and the next hour of work is me.&lt;/p&gt;
&lt;p&gt;But that feeling is borrowed from somewhere else. It comes from working with people, where waiting genuinely costs something. A person who’s blocked on you is losing their afternoon, and the clock on that is what makes so much of working life stressful. I learned to feel that clock, and now I feel it for something that doesn’t have one. A waiting agent costs nothing. It doesn’t get bored, it doesn’t lose the thread, and it will be in exactly the same state in an hour or tomorrow morning.&lt;/p&gt;
&lt;p&gt;So this is the adjustment, and I think it’s the one that decides whether you end up content or fried. Six sessions waiting is not six people waiting. When it’s too many, I let three of them sit, and nothing is lost. The urgency was never coming from the agents. We spent years building tools that turn every human interaction into something time-sensitive, and it would be a waste to do the same thing to ourselves with the first colleagues who genuinely don’t mind.&lt;/p&gt;
&lt;h2 id=&quot;the-retro&quot;&gt;The retro&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What went well.&lt;/strong&gt; Ideas get a hearing. I’ve shipped in languages I don’t write. The lessons get written down.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What didn’t.&lt;/strong&gt; I still brief from memory instead of from the system, and the agent still has to catch me at it. I’m the bottleneck more often than the agent is.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What I’d change.&lt;/strong&gt; Treat the unblocking as the job, not as an interruption to it. Write the lesson down the first time, not the third. Let a waiting agent wait.&lt;/p&gt;
&lt;p&gt;Fourteen years ago I was a scrum master because a project needed one and I was already on the team. It turns out I am again, for the same reason. The team just types faster.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Co-authored with Claude, because I don’t do anything alone any more.&lt;/em&gt;&lt;/p&gt;
</content:encoded><category>ai-agents-in-action</category><category>agile</category><category>scrum</category><category>software-development</category><category>claude-code</category><category>Writing</category></item><item><title>From Contact Form to Order Portal in Four Days: Building a 3D Printing Business’s Customer Portal on Netlify</title><link>https://geekconsulting.au/writing/order-portal-in-four-days/</link><guid isPermaLink="true">https://geekconsulting.au/writing/order-portal-in-four-days/</guid><description>The website could give you a ballpark price in three questions. The moment you pressed Send, it turned into an email to Tim, and everything after that lived in inboxes and memory.</description><pubDate>Sun, 27 Sep 2026 22:19:59 GMT</pubDate><content:encoded>&lt;h2 id=&quot;tldr&quot;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;G3DK Printing is a one-person 3D printing business in Sydney. Tim prints things from models customers send in. The website was friendly up to the point you asked for a quote; after that it was just an email, and the quoting, the back-and-forth, the photos, the invoice, the postage and the tracking number all happened by hand.&lt;/p&gt;
&lt;p&gt;Now every enquiry becomes an order with its own page. Customers sign in with a link emailed to them, with no password to forget. They can chat with Tim on the page or just reply to the emails, accept a quote with a tick-box, see photos of their finished print, pay an invoice that already includes the postage, and get a tracking link when it ships.&lt;/p&gt;
&lt;p&gt;It took four days to build, and it went live on 27 September 2026. The first real order went all the way through on launch day. It’s still labelled beta.&lt;/p&gt;
&lt;p&gt;If you’re here for the technical details, they’re at the bottom.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;At a glance&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Stack:&lt;/strong&gt; still a static site, now with Netlify Functions, Netlify Database (Postgres) and Netlify Blobs, plus Postmark for email, Square for payments and the Australia Post postage API. No framework, no extra host, no passwords.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Time:&lt;/strong&gt; four days, 24–27 September 2026, in eight milestones. Each was a pull request, tested on a deploy preview, then merged. Built pair-programming with Claude Code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;By the numbers:&lt;/strong&gt; about 4,800 lines of portal code, 1,900 lines of tests (91 tests against a real in-process Postgres), 23 serverless functions, 7 database migrations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Status:&lt;/strong&gt; live in production, in beta. First end-to-end order on launch day: enquiry, quote, chat by email, photos, live postage, card payment, shipped, arrived.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;everything-after-theemail&quot;&gt;Everything after the email&lt;/h2&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/order-portal-in-four-days/1.png&quot; alt=&quot;The front half of the job, which the site already did well.&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;The front half of the job, which the site already did well.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;G3DK Printing is a one-person workshop: two Bambu Lab P1S printers and a 16K resin printer. The website already did the front half of the job rather well. It had a price finder, where three plain-English questions get you a ballpark; an advanced calculator for people who want the dials; and a quote form, handled by Netlify Forms, that emailed Tim. It even had an AI agent endpoint, an A2A agent card, so another AI could get a quote and submit a request without a human in sight.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/order-portal-in-four-days/2.png&quot; alt=&quot;The price finder. Before the portal, the “Send this and my model” button was where the website’s job ended.&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;The price finder. Before the portal, the “Send this and my model” button was where the website’s job ended.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What it didn’t have was anything after the email. A customer who’d sent a model got a reply from a person, then another, then a price, then photos, then some way to pay, then a parcel. Every step of that was Tim remembering to do the next thing, in whichever inbox the last message had landed in.&lt;/p&gt;
&lt;p&gt;I built it pair-programming with Claude Code, on a simple split: &lt;strong&gt;I supply direction, the agent supplies execution.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&quot;eight-requirements-and-a-handful-ofrules&quot;&gt;Eight requirements and a handful of rules&lt;/h2&gt;
&lt;p&gt;The brief for phase 2 was eight requirements:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Run the order portal and back end on Netlify, integrated with Square.&lt;/li&gt;
&lt;li&gt;After an enquiry, customers log in with their email and a magic link that expires in a few hours.&lt;/li&gt;
&lt;li&gt;Each order has chat, and it works by email too.&lt;/li&gt;
&lt;li&gt;Quote generated, quote accepted by the customer.&lt;/li&gt;
&lt;li&gt;Printing, with photos uploaded.&lt;/li&gt;
&lt;li&gt;Invoice through Square, including postage.&lt;/li&gt;
&lt;li&gt;Ship once it’s paid.&lt;/li&gt;
&lt;li&gt;Optionally, Australia Post labels and pickups.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;And a few business rules that shaped everything: a $20 minimum, delivery from $15, no GST (the business isn’t registered), photos go out &lt;em&gt;before&lt;/em&gt; payment, payment comes &lt;em&gt;before&lt;/em&gt; shipping, and G3DK only prints supplied models, with no design work.&lt;/p&gt;
&lt;p&gt;Those last two orderings, photos before payment and payment before shipping, turned out to be the spine of the whole build. Most of the interesting code exists to make sure they can’t be skipped.&lt;/p&gt;
&lt;h2 id=&quot;keep-the-static-site-add-a-thin-layer-besideit&quot;&gt;Keep the static site, add a thin layer beside it&lt;/h2&gt;
&lt;p&gt;The site stayed static: plain HTML pages and vanilla JavaScript modules, no framework. Beside it went a thin server layer of Netlify Functions, one file per endpoint, with a Postgres database from Netlify, Netlify Blobs for files, and three outside services: Postmark for email, Square for payments and Australia Post for postage prices. No new host, nothing to patch at two in the morning.&lt;/p&gt;
&lt;p&gt;Every order moves through one status machine, from &lt;em&gt;enquiry&lt;/em&gt; through &lt;em&gt;quoted&lt;/em&gt;, &lt;em&gt;printing&lt;/em&gt;, &lt;em&gt;invoiced&lt;/em&gt; and &lt;em&gt;paid&lt;/em&gt; to &lt;em&gt;shipped&lt;/em&gt; and &lt;em&gt;completed&lt;/em&gt;. The server and both pages, the customer’s and Tim’s, share the same definition, so the wording can’t drift between what the customer reads and what Tim sees.&lt;/p&gt;
&lt;p&gt;And everything shipped &lt;strong&gt;dark&lt;/strong&gt;. Every portal function returns a 404 unless one environment variable, PORTAL_ENABLED, is true. Deploy previews had it on, with Square’s sandbox; production had it off until launch day. That one flag is also the rollback.&lt;/p&gt;
&lt;h2 id=&quot;four-days-eight-milestones&quot;&gt;Four days, eight milestones&lt;/h2&gt;
&lt;p&gt;The build went in as eight milestones, M0 to M7, each its own pull request, tested on a Netlify deploy preview and then merged:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;M0:&lt;/strong&gt; foundations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M1:&lt;/strong&gt; enquiries become orders, a “My orders” page, and Tim’s admin list.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M2:&lt;/strong&gt; a message thread on each order, the email bridge, and bounce handling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M3:&lt;/strong&gt; quotes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M4:&lt;/strong&gt; printing and photos.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M5:&lt;/strong&gt; delivery details, Square invoices and payment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M6:&lt;/strong&gt; shipping.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;M7:&lt;/strong&gt; go-live: the Privacy and Terms pages, enquiries going to the portal first, and “Your orders” links.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Because the whole thing was dark in production, each milestone could land on main the moment it was ready, with nothing visible to customers. The deploy previews were where it got exercised, against real Postmark and the Square sandbox.&lt;/p&gt;
&lt;h2 id=&quot;the-link-that-silently-didntsend&quot;&gt;The link that silently didn’t send&lt;/h2&gt;
&lt;p&gt;Sign-in is by magic link: no passwords, no accounts to create. To stop anyone using the “email me a link” button as a spam cannon, there’s a limit of five links per address per hour.&lt;/p&gt;
&lt;p&gt;Testing found the flaw. The limit was also counting the sign-in links inside our &lt;em&gt;own&lt;/em&gt; order emails. A fast-moving order sends a lot of those: quote, printing, printed, invoice. By the time the customer went looking for a fresh link, the order had spent the allowance for them, and their “email me a link” silently sent nothing. No error, no email, just a customer who can’t get in.&lt;/p&gt;
&lt;p&gt;The fix was to count only the links a &lt;em&gt;person&lt;/em&gt; asks for.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/order-portal-in-four-days/3.png&quot; alt=&quot;Sign-in: an email address, a link, no password. The Beta notice stays until the portal leaves beta.&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;Sign-in: an email address, a link, no password. The Beta notice stays until the portal leaves beta.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;A rate limit protects you from the people pressing the button. Make sure it’s counting them, and not your own system doing its job.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;what-the-unit-tests-couldntsee&quot;&gt;What the unit tests couldn’t see&lt;/h2&gt;
&lt;p&gt;npm test runs 91 tests against a real Postgres, running in-process, with every migration applied. The outside services are faked at the fetch level, so the tests can check the exact request bodies that would go to Square and Postmark. And for the rules that matter most (the webhook signature, customers only ever seeing their own orders, the payment gate, the guard against paying an invoice twice) we deliberately broke the code and confirmed a test failed. A test you’ve never seen fail is a test you’re taking on trust.&lt;/p&gt;
&lt;p&gt;Then came the browser runs. Headless Chromium drove whole orders through against a scratch server running the real function modules: the customer at phone width, Tim at desktop width. They caught what no unit test could:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;buttons touching each other,&lt;/li&gt;
&lt;li&gt;a pickup order ending on a “Delivered” badge,&lt;/li&gt;
&lt;li&gt;local drop-off customers being promised “tracking details” they were never going to get.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;None of those is a bug in the logic. The status was right, the rules held, every assertion passed. They’re bugs in what a person &lt;em&gt;reads&lt;/em&gt;, and you only find those by being the person.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;Unit tests prove the rules. Only walking an order through as the customer proves the story they’re told.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;the-unglamorous-half&quot;&gt;The unglamorous half&lt;/h2&gt;
&lt;p&gt;Getting email right took as long as some features.&lt;/p&gt;
&lt;p&gt;Postmark sends as &lt;a href=&quot;mailto:orders@g3dkprinting.au&quot;&gt;orders@g3dkprinting.au&lt;/a&gt;, with its own DKIM key and a custom return path, so SPF and DKIM both line up. Replies come back through a separate subdomain, so every order gets its own reply address and the customer can answer from their own inbox. The business mailbox had its own saga: Proton had no free domain slot, so the domain started on ImprovMX forwarding; later an old domain was retired from Proton, g3dkprinting.au moved in, and the old domain was parked on ImprovMX’s free plan. DMARC reports go to Postmark’s free weekly digest. The first Google report showed only Postmark sending, and everything passing.&lt;/p&gt;
&lt;p&gt;A few surprises along the way:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Square’s hosted payment page asked for a US ZIP code, because the sandbox test card is a US card.&lt;/li&gt;
&lt;li&gt;Proton kept “delivering” test emails to its own inbox, because an old linked Gmail address was still attached to the account.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;launch-day&quot;&gt;Launch day&lt;/h2&gt;
&lt;p&gt;Launch was a runbook, not an event. Production environment variables went in, the Postmark inbound webhook moved from the deploy preview to production, PORTAL_ENABLED went to true, and the site redeployed.&lt;/p&gt;
&lt;p&gt;Then a real order went all the way through: enquiry, quote, acceptance, a reply by email landing in the chat, printing and photos, live Australia Post prices, a production Square invoice paid by card, shipped, arrived.&lt;/p&gt;
&lt;p&gt;The “Payment received” email arrived &lt;strong&gt;one second&lt;/strong&gt; after the card payment.&lt;/p&gt;
&lt;h2 id=&quot;what-it-looks-likenow&quot;&gt;What it looks like now&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;An order page for every enquiry&lt;/strong&gt;, with the order’s status, its chat, its quote, photos of the print, the invoice and, once it ships, the tracking link.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Chat that nobody has to use.&lt;/strong&gt; Every message can be answered by replying to the email.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Square invoices with postage already in them&lt;/strong&gt;, priced live from Australia Post, payable by card, or by PayID for customers who’d rather do a bank transfer.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rules the server enforces, whatever the page shows:&lt;/strong&gt; no Printed without a photo, no shipping without payment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A “Move back a stage” control&lt;/strong&gt; for Tim, for when a print fails and needs re-quoting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;About 4,800 lines of portal code, 1,900 lines of tests, 23 functions and 7 migrations&lt;/strong&gt;, on the same Netlify site the static pages always lived on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Privacy and Terms pages&lt;/strong&gt; written from what the code actually does, including a promise the code keeps: files are deleted 12 months after an order closes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A “Beta” notice&lt;/strong&gt;, and a fallback: if the portal ever can’t take an enquiry, the quote form still posts to Netlify Forms, so nothing is lost.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-honestpart&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;The features were never the hard part. A quote builder is a form. A chat thread is a table. The four days went into the orderings: making “photos before payment, payment before shipping” true on the server, not just on the page; making sure going backwards can’t become a way around going forwards; checking that Square’s total matches ours to the cent before anything is sent. And they went into email, which is where a small business actually meets its customers, and where most of the surprises were.&lt;/p&gt;
&lt;p&gt;What made four days possible wasn’t typing speed. It was that the whole thing was dark until the last day. Each milestone landed small, got exercised on a preview against the real services, and merged, and nothing a customer could see changed until one flag flipped. The rollback was the same flag.&lt;/p&gt;
&lt;p&gt;And it’s not finished. It’s in beta. Australia Post labels and pickups, the optional eighth requirement, aren’t built. But the first real customer order went from “here’s my model” to “it’s arrived” without anyone having to remember what came next.&lt;/p&gt;
&lt;p&gt;Keep it dark until it’s done. Then flip one flag.&lt;/p&gt;
&lt;h2 id=&quot;under-the-hood-how-itsbuilt&quot;&gt;Under the hood: how it’s built&lt;/h2&gt;
&lt;h3 id=&quot;what-it-runson&quot;&gt;What it runs on&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.netlify.com/build/functions/overview/&quot;&gt;Netlify Functions&lt;/a&gt;: Node ESM, one file per endpoint, 23 in total, rate-limited at the edge.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.netlify.com/build/data-and-storage/netlify-database/&quot;&gt;Netlify Database&lt;/a&gt; (@netlify/database 2.0.1): Postgres, migrations applied on deploy, a database branch per deploy preview.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.netlify.com/build/data-and-storage/netlify-blobs/&quot;&gt;Netlify Blobs&lt;/a&gt; (@netlify/blobs 11.1.1): model files, photos and chat attachments.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://postmarkapp.com/&quot;&gt;Postmark&lt;/a&gt;: outbound email, plus &lt;a href=&quot;https://postmarkapp.com/developer/user-guide/inbound&quot;&gt;inbound&lt;/a&gt; for reply-by-email, and bounce and spam webhooks. &lt;a href=&quot;https://dmarc.postmarkapp.com/&quot;&gt;DMARC Digests&lt;/a&gt; for the weekly reports.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developer.squareup.com/docs/invoices-api/overview&quot;&gt;Square Invoices API&lt;/a&gt;: customers, orders, invoices, and &lt;a href=&quot;https://developer.squareup.com/docs/webhooks/step3validate&quot;&gt;signed webhooks&lt;/a&gt; for payment.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developers.auspost.com.au/apis/pac&quot;&gt;Australia Post Postage Assessment Calculator (PAC)&lt;/a&gt;: domestic parcel prices.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://pglite.dev/&quot;&gt;PGlite&lt;/a&gt; (@electric-sql/pglite 0.2.17): in-process Postgres for the tests.&lt;/li&gt;
&lt;li&gt;The existing site: static HTML with TailwindCSS and vanilla JavaScript modules, plus an &lt;a href=&quot;https://a2a-protocol.org/&quot;&gt;A2A&lt;/a&gt; agent endpoint.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;the-shape&quot;&gt;The shape&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Browser pages (static HTML + vanilla JS modules)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   │  fetch /api/…  (same-origin, JSON)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   ▼&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Netlify Functions (Node ESM, one file per endpoint, rate-limited at the edge)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   │&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   ├── server/*.mjs      business logic (auth, orders, messages, quotes, invoicing, dispatch…)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   ├── lib/*.mjs         shared with the browser: pricing maths, quote maths, order statuses&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   │&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   ├── Netlify Database  Postgres, migrations applied on deploy, a branch per deploy preview&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   ├── Netlify Blobs     model files, photos, chat attachments&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   │&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   └── Postmark · Square · Australia Post PAC&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The status machine, shared by the server and both pages:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;enquiry → quoting → quoted → accepted → printing → printed → invoiced → paid → shipped → completed&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                                        (plus declined · expired · cancelled)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;decisions&quot;&gt;Decisions&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Keep the static site; add functions beside it.&lt;/strong&gt; No framework, no extra host.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ship dark behind one flag.&lt;/strong&gt; PORTAL_ENABLED gates every portal function (404 otherwise), and it’s also the rollback.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No passwords.&lt;/strong&gt; Magic links only, stored as SHA-256 hashes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Staff emails never contain sign-in links.&lt;/strong&gt; An admin link sitting in an inbox would open every customer’s details.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The server re-prices every quote.&lt;/strong&gt; Totals from the browser are never trusted.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Square doesn’t email the invoice&lt;/strong&gt; (SHARE_MANUALLY). Our email carries the photos and the pay link, so the customer gets one clear message, not two.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;“Move back a stage” only goes backwards&lt;/strong&gt;, so the forward rules can’t be bypassed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A fallback for enquiries.&lt;/strong&gt; If the portal can’t take one, the form still posts to Netlify Forms.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;sign-in-without-passwords&quot;&gt;Sign-in without passwords&lt;/h3&gt;
&lt;p&gt;Customers never create an account. The quote form creates a customer and an order, and the confirmation email carries a one-time sign-in link.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Tokens are stored as SHA-256 hashes only.&lt;/strong&gt; A database leak gives nobody a working link or session.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The token travels in the URL fragment&lt;/strong&gt; (/portal-verify.html#t=…), which browsers never send to servers or put in Referer headers. The verify page also needs a &lt;strong&gt;click&lt;/strong&gt; before it uses the token, so email link scanners that “pre-open” links don’t burn it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Links expire after 3 hours and work once.&lt;/strong&gt; Customer sessions last 30 days and slide with use. Tim’s admin sessions last 12 hours.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Every POST goes through one wrapper&lt;/strong&gt; that checks the feature flag, the method, a same-origin Origin header and a JSON body.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Only links a person asks for count&lt;/strong&gt; towards the 5-per-address-per-hour limit (see the story above).&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;chat-that-works-byemail&quot;&gt;Chat that works by email&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Each order gets its own reply address&lt;/strong&gt;, g3dk-1001.&amp;lt;16 hex chars&amp;gt;@reply.g3dkprinting.au. The hex is an HMAC of the order number, so addresses can’t be guessed or forged.&lt;/li&gt;
&lt;li&gt;reply.g3dkprinting.au has an MX record pointing at &lt;strong&gt;Postmark inbound&lt;/strong&gt;, which posts every message to a webhook (basic auth in the URL). The function checks the signature, works out who sent it (the customer, or Tim from any of Tim’s known addresses), strips the quoted history, and adds it to the thread with any attachments.&lt;/li&gt;
&lt;li&gt;The other side is emailed &lt;strong&gt;unless they’re looking at the order right now&lt;/strong&gt;: a presence table records open pages, and a five-minute sweep catches the rest. Postmark bounce and spam webhooks mark a customer’s address as failing and warn Tim.&lt;/li&gt;
&lt;li&gt;Pages poll every 10 seconds. System lines in the chat (“Quote accepted”, “Payment received: $35.00”) double as the signal for the other side’s page to refresh itself.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;quotes&quot;&gt;Quotes&lt;/h3&gt;
&lt;p&gt;The admin quote builder suggests a price from the slicer weight (grams × material rate × 1.2) and allows extras and discounts. The server re-prices every quote. Quotes are versioned (a revision supersedes the last), expire after 14 days (a daily job), top up to the $20 minimum as a separate line, and are accepted with a tick-box. The time and IP of acceptance are recorded, and the tick-box links to the Terms.&lt;/p&gt;
&lt;h3 id=&quot;photos-you-cantrust&quot;&gt;Photos you can trust&lt;/h3&gt;
&lt;p&gt;Tim uploads photos from a phone. The admin page shrinks them to 2000 px JPEG in the browser and uploads them one at a time. The server then &lt;strong&gt;sniffs the magic bytes&lt;/strong&gt;: only real JPEG, PNG or WebP files are stored as photos. When one is shown inline, it’s served as that exact type with X-Content-Type-Options: nosniff and a CSP sandbox. Anything else is only ever a download. An “image” that’s really HTML can’t run on our domain.&lt;/p&gt;
&lt;h3 id=&quot;square-invoices-with-postageincluded&quot;&gt;Square invoices, with postage included&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Postage:&lt;/strong&gt; the packed parcel’s weight and size go to the Australia Post PAC API from the workshop’s postcode. Prices come back cheapest first, with the $15 delivery floor applied.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The invoice:&lt;/strong&gt; find or create the Square customer, build a Square &lt;strong&gt;order&lt;/strong&gt; (each quoted line, discounts as fixed-amount discounts, the minimum top-up and a delivery line), then &lt;strong&gt;check Square’s total against ours&lt;/strong&gt;. If they differ by a cent, nothing is sent. Then create and publish the invoice: card only, due in 7 days.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SHARE_MANUALLY:&lt;/strong&gt; Square doesn’t email the invoice; our email does.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Payment:&lt;/strong&gt; Square webhooks are signed with HMAC-SHA256 over the &lt;strong&gt;notification URL plus the body&lt;/strong&gt;, and deduplicated by event ID. invoice.payment_made moves the order to Paid, emails the customer, and tells Tim it’s ready to ship.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;PayID:&lt;/strong&gt; customers can pay by bank transfer instead. “Mark paid (PayID)” first checks Square hasn’t already been paid by card, then cancels the Square invoice so it can’t be paid twice.&lt;/li&gt;
&lt;li&gt;Reminders go out 3 days after sending and on the due date.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The webhook check, copied from the deployed source. Note the URL is part of what’s signed, so it must be exactly the one registered with Square:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/**&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; * Square signs each webhook with HMAC-SHA256(signature key, notification URL + raw body), base64.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; * The URL must be exactly the one registered in the Square Developer Dashboard.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; */&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export function validWebhookSignature({ signature, url, body }) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const key = env(&quot;SQUARE_WEBHOOK_SIGNATURE_KEY&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  if (!key || !signature) return false;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const expected = Buffer.from(createHmac(&quot;sha256&quot;, key).update(url + body).digest(&quot;base64&quot;));&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const given = Buffer.from(signature);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  return expected.length === given.length &amp;amp;&amp;amp; timingSafeEqual(expected, given);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;shipping-and-the-rules-that-cant-beskipped&quot;&gt;Shipping, and the rules that can’t be skipped&lt;/h3&gt;
&lt;p&gt;“Ship it” is enforced on the server: an order that isn’t paid can’t be shipped, whatever the page shows. Tim enters the tracking number (dashes and spaces from a pasted label are fine), the customer gets an email with an Australia Post tracking link, and the order completes when the customer taps “It’s arrived”, or 14 days after shipping (daily job). Pickup and local drop-off skip tracking and go straight from Paid to Collected or Delivered.&lt;/p&gt;
&lt;p&gt;“Move back a stage” never goes back past payment, needs a reason, and can optionally email the customer. It also tidies up: an unpaid invoice is withdrawn in Square, the old quote is set aside, and tracking details are cleared.&lt;/p&gt;
&lt;h3 id=&quot;testing-a-real-database-fake-everything-else&quot;&gt;Testing: a real database, fake everything else&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;91 tests against PGlite&lt;/strong&gt; with every migration applied, so the SQL is exactly what production runs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Postmark, Square and Australia Post are replaced by a fake&lt;/strong&gt; &lt;strong&gt;fetch&lt;/strong&gt; that behaves like the real APIs, so tests check the exact request bodies (for example that an invoice goes out card-only and SHARE_MANUALLY).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mutation checks&lt;/strong&gt; on the security-critical rules: webhook signature, customers only seeing their own orders, the payment gate, the paid-invoice guard.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Headless Chromium runs&lt;/strong&gt; against a scratch server running the real function modules, customer at phone width and Tim at desktop width.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deploy previews:&lt;/strong&gt; every milestone tested by hand against real Postmark and the Square sandbox.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;gotchas&quot;&gt;Gotchas&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Rate-limit only what people ask for.&lt;/strong&gt; Counting system-sent sign-in links let a busy order lock its own customer out, silently.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Put the token in the fragment and make the verify page need a click&lt;/strong&gt;, or link scanners will spend it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Square’s webhook signature covers the URL.&lt;/strong&gt; If the URL you check against differs from the registered one, every webhook fails.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Square’s sandbox test card is a US card&lt;/strong&gt;, so its payment page asks for a US ZIP code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Check what customers read, not just what the code does.&lt;/strong&gt; A pickup order ended on “Delivered”; local drop-off promised tracking.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;launch-runbook&quot;&gt;Launch runbook&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;Production env vars: staff emails, a new reply-signing secret, the Square production webhook key.&lt;/li&gt;
&lt;li&gt;The Postmark inbound webhook moved from the deploy preview to production.&lt;/li&gt;
&lt;li&gt;PORTAL_ENABLED=true and a redeploy.&lt;/li&gt;
&lt;li&gt;A smoke test with a real order: enquiry → quote → acceptance → a reply by email into the chat → printing and photos → live Australia Post prices → a production Square invoice paid by card → shipped → arrived.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;whats-next&quot;&gt;What’s next&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Australia Post labels and pickups from the admin page&lt;/strong&gt;, via an aggregator like Starshipit or Shippit, once there are enough parcels that copying tracking numbers gets tedious.&lt;/li&gt;
&lt;li&gt;Dropping the Beta label.&lt;/li&gt;
&lt;li&gt;Tightening DMARC to p=quarantine after a few clean weekly digests.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;G3DK Printing:&lt;/em&gt; &lt;a href=&quot;https://g3dkprinting.au&quot;&gt;&lt;em&gt;g3dkprinting.au&lt;/em&gt;&lt;/a&gt;&lt;em&gt;. Send a model link, get a price, follow your print all the way to your door.&lt;/em&gt;&lt;/p&gt;
</content:encoded><category>netlify</category><category>web-development</category><category>ai-agents-in-action</category><category>serverless</category><category>small-business</category><category>Writing</category></item><item><title>Project: Household karaoke</title><link>https://geekconsulting.au/projects/pikaraoke/</link><guid isPermaLink="true">https://geekconsulting.au/projects/pikaraoke/</guid><description>PiKaraoke in the homelab: guests queue K-pop, Mandopop and Cantopop from their phones, and local AI models keep the library tidy.</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate><category>Project</category><category>Homelab</category></item><item><title>A Spring Scorcher, a Powerwall, and a Spa That Only Heats on Spare Sunshine</title><link>https://geekconsulting.au/writing/spa-heats-on-spare-sunshine/</link><guid isPermaLink="true">https://geekconsulting.au/writing/spa-heats-on-spare-sunshine/</guid><description>I wrote a careful brief for this. Built as written, it would have quietly flattened my Powerwall any afternoon a cloud went over, and nothing in the house would have told me.</description><pubDate>Sat, 26 Sep 2026 10:34:06 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/spa-heats-on-spare-sunshine/1.jpg&quot; alt=&quot;&quot;&gt;&lt;/figure&gt;&lt;h2 id=&quot;tldr&quot;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;We have an inflatable spa in the backyard, solar panels on the roof and a Tesla Powerwall. The spa’s heater draws about what a kettle does, for hours at a time, and it used to run on a timer whether the sun was out or not.&lt;/p&gt;
&lt;p&gt;Now the house heats it only with sunshine it would otherwise send to the grid for very little, and only once the Powerwall is full, because the Powerwall is what keeps our evenings cheap. When the sun fades, the heater stops. If the water’s still cold at half past three, it gets topped up anyway, so it’s usable that evening.&lt;/p&gt;
&lt;p&gt;It went live on a hot spring afternoon, and so far it has proven one thing: it’s very good at turning the heater &lt;em&gt;off&lt;/em&gt;. Whether it turns on when it should is waiting on the next sunny day.&lt;/p&gt;
&lt;p&gt;If you’re here for the config, it’s at the bottom.&lt;/p&gt;
&lt;h2 id=&quot;the-spa-that-ran-on-atimer&quot;&gt;The spa that ran on a timer&lt;/h2&gt;
&lt;p&gt;The spa is a Bestway AirJet: the inflatable kind, a pump unit on the side, an app on the phone. The app has a timer, and the timer did what timers do: pump on at nine, off at four, every day, with the heater chasing its thermostat whenever the pump ran, regardless of whether the roof was making four kilowatts or none.&lt;/p&gt;
&lt;p&gt;Which is daft, because a spa is about the best thermal battery a house can own. A thousand-odd litres of water hold heat for hours. It doesn’t care &lt;em&gt;when&lt;/em&gt; it was heated, only that it’s warm by evening. It should soak up the solar the house can’t use and sit on it until someone gets in.&lt;/p&gt;
&lt;p&gt;I’d known that for as long as we’ve had it, and never done anything about it, for the usual reasons. The integration is cloud-only and works through a reverse-engineered API. The newer Bestway pumps use a different backend from the older ones. The setup instructions warn that the right server region is “often not the obvious one”. Every step was guesswork, and there was always something more useful to do with an afternoon.&lt;/p&gt;
&lt;p&gt;What finally moved it was the weather. Late September, a proper spring scorcher, the aircon going flat out, and the house pulling somewhere between nine and thirteen kilowatts, swinging minute to minute. On a day like that you start thinking hard about where your electrons go.&lt;/p&gt;
&lt;p&gt;I’ve written about this house before: &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;turning Home Assistant into a Grafana data platform&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-2/&quot;&gt;building the dashboard I’d abandoned three times&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-3/&quot;&gt;the plug-in alert that paid for the lot&lt;/a&gt;, and &lt;a href=&quot;https://geekconsulting.au/writing/proxmox-magi-panel/&quot;&gt;the MAGI panel on the cluster&lt;/a&gt;. Same collaborator, same principle: &lt;strong&gt;I supply direction, the agent supplies execution.&lt;/strong&gt; This time I thought I’d supply the direction in advance, in writing. That’s where it went wrong.&lt;/p&gt;
&lt;h2 id=&quot;the-brief-i-wrote-and-the-house-itforgot&quot;&gt;The brief I wrote, and the house it forgot&lt;/h2&gt;
&lt;p&gt;I did the research in the morning in Claude’s ordinary chat app, and came out with a tidy brief and a draft Home Assistant package. It’s the same Claude I build everything else with. What it didn’t have was the context. The homelab project lives in Claude Code: the repo of runbooks and notes about every device in the house, the memory of what we’ve built and what bit us, and live access to Home Assistant itself. The chat had my description of the house and nothing else. The rules it helped me write were simple:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Heat on&lt;/strong&gt; when the house has been exporting more than 2,400 W for five minutes, or when the price for exporting goes negative.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Heat off&lt;/strong&gt; when the house has been importing more than 300 W for five minutes, once the heater has run at least twenty minutes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A comfort floor&lt;/strong&gt; at 15:30: if the water’s under temperature, heat it regardless.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It had a task list, a section on known gaps, and a house rule at the bottom: &lt;em&gt;verify against the live system&lt;/em&gt;. I was quite pleased with it.&lt;/p&gt;
&lt;p&gt;Then I handed the brief to Claude Code, which has all that context and can see the house. The first task in the brief was “discover current state, read-only first”, and it did exactly that before touching anything. Among the integrations it listed was the one the brief never mentioned: &lt;strong&gt;Tessie, driving a Tesla Powerwall.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I haven’t forgotten we have a Tesla Powerwall; it had a starring role in the car alert piece. I just didn’t mention it in the chat, so the chat couldn’t know. And what I’d forgotten myself was what it &lt;em&gt;does&lt;/em&gt; to a rule that watches the grid.&lt;/p&gt;
&lt;p&gt;On a sunny afternoon spare solar goes into the battery first, so the house exports only once the battery’s full. So far, so harmless: the heat-on rule would simply wait. The problem is the other end. When a cloud comes over while the spa is heating, the house doesn’t start importing. &lt;strong&gt;The Powerwall covers the shortfall&lt;/strong&gt;, instantly, silently, which is the entire point of a Powerwall. Grid import stays near zero. The stop rule, watching for 300 W of import, never fires.&lt;/p&gt;
&lt;p&gt;The spa would keep heating, the battery would keep emptying, and every reading in the brief would say everything was fine. That evening the house would buy back at full price the power I’d poured into the spa at lunchtime.&lt;/p&gt;
&lt;p&gt;The redesign took one conversation. The spa doesn’t get “exported power”; it gets &lt;strong&gt;spare power&lt;/strong&gt;, defined so the battery comes first:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Export counts as spare.&lt;/li&gt;
&lt;li&gt;Battery charging counts as spare only once the battery is at &lt;strong&gt;95 %&lt;/strong&gt;. Below that, the battery gets it.&lt;/li&gt;
&lt;li&gt;Battery &lt;em&gt;discharging&lt;/em&gt; always counts as a shortfall. &lt;strong&gt;The spa never drains the Powerwall.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The part I find properly embarrassing: I’d already made this argument. In the plug-in alert piece I wrote that charging the car from solar while the battery sits at 60 % “is just moving the problem”. Then I sat down and specified a spa heater that did precisely that, because I’d researched it in a conversation that knew only what I remembered to tell it, and on that morning, I didn’t remember the battery.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;A brief written without access to the system is a hypothesis about the system. Its first task should be finding out what it got wrong, and this one’s did.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;two-meters-onelie&quot;&gt;Two meters, one lie&lt;/h2&gt;
&lt;p&gt;The house has two ways of measuring power. An Enphase Envoy on the switchboard, read locally, and the Powerwall’s own figures, read from Tesla’s cloud through Tessie. The first cut of the spare-power sensor used both: grid from the Envoy, battery from Tessie. It seemed sensible. The Envoy is local and fast; why go to the cloud for something you can read on the LAN?&lt;/p&gt;
&lt;p&gt;The first live reading came back at &lt;strong&gt;−8,721 W&lt;/strong&gt;. On a moment when the true figure was a few hundred.&lt;/p&gt;
&lt;p&gt;Nothing was broken. Both numbers were correct, just not at the same instant. They sample at different moments, and in a house whose load was lurching several kilowatts a minute that afternoon, one meter was describing a different second from the other. Subtract one from the other and you get a number that describes no moment at all.&lt;/p&gt;
&lt;p&gt;The fix was to take all three inputs, grid, battery and charge level, from the &lt;strong&gt;same Tessie poll&lt;/strong&gt;, where solar + battery + grid adds up to the house load exactly. Before trusting it, the agent checked that balance at nine points across the previous twenty-four hours, which is how the sign conventions got pinned down too: grid is positive when importing, battery is positive when &lt;em&gt;discharging&lt;/em&gt;. It didn’t assume either. They’re exactly the kind of thing that’s the other way round on someone else’s integration.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;Never compute a difference across two sources that sample at different moments. Take every term of a balance from the same reading, or the balance is fiction.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The same instinct settled a question the brief had left open: which way round does the Amber feed-in price go when exporting &lt;em&gt;costs&lt;/em&gt; you? Rather than guess, the agent read Home Assistant’s own Amber source code and found it multiplies feed-in by −1. So “less than zero” means you’re paying to export. The draft had it right. Now it was right on purpose.&lt;/p&gt;
&lt;h2 id=&quot;the-id-that-could-read-but-notwrite&quot;&gt;The ID that could read but not write&lt;/h2&gt;
&lt;p&gt;The Bestway integration went in cleanly, and the spa showed up in Home Assistant with its water temperature, its thermostat and its switches. Reads worked. So the agent tried a write: heater off, then on.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Error 10001.&lt;/strong&gt; Both times, at 15:27 and 15:30, from Bestway’s cloud, through both versions of its command API. And after the writes, the reads went bad too: the version sensors went to unknown and the temperatures to nothing.&lt;/p&gt;
&lt;p&gt;The setup had used the app’s own visitor ID. My phone’s app has never had an account: it’s a guest, and the integration will happily log in as that same guest. That turns out to be a known, open upstream issue, &lt;a href=&quot;https://github.com/cdpuk/ha-bestway/issues/121&quot;&gt;ha-bestway#121&lt;/a&gt;: when Home Assistant and the phone app share one identity, control breaks. The readings look fine until you try to &lt;em&gt;change&lt;/em&gt; something, which is the only thing we were there to do.&lt;/p&gt;
&lt;p&gt;The way out was a feature I’d assumed the new app had dropped. It still offers &lt;em&gt;Share the device&lt;/em&gt; with a QR code, the route the older setup instructions describe. I screenshotted it, the agent decoded it locally, never printing it (it grants control of the spa), and re-added the spa with it. Home Assistant got &lt;strong&gt;its own guest identity&lt;/strong&gt;. It shows up in the app as an extra “guest” user, which I’m now under strict instructions from myself never to tidy away.&lt;/p&gt;
&lt;p&gt;Then came the test that counts. Setpoint from 33 °C to 34 °C in Home Assistant, and I read 34 off the app on my phone. Heat off, heat on, both reported back from the spa itself. No 10001s.&lt;/p&gt;
&lt;p&gt;We couldn’t prove it from the power meter. A ~2 kW heater step is invisible when the house is swinging between nine and thirteen kilowatts because the aircon is fighting the afternoon. The app was the witness.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;An identity that can read isn’t one that can write. Prove control with a change you can see from the&lt;/em&gt; other &lt;em&gt;side, not with the tool’s report that it sent the command.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;there-is-no-heaterswitch&quot;&gt;There is no heater switch&lt;/h2&gt;
&lt;p&gt;The brief, the draft and the integration’s README all talked about a heater switch. On this pump, there isn’t one. The heater &lt;em&gt;is&lt;/em&gt; the thermostat’s mode: heat or off. The draft was rewritten around that before anything was deployed.&lt;/p&gt;
&lt;p&gt;At 15:55 the automations went in, with solar mode still off. At &lt;strong&gt;16:00:06&lt;/strong&gt; the heater and the filter pump both switched off.&lt;/p&gt;
&lt;p&gt;Nobody had touched them. Home Assistant’s logbook showed no automation, no user and no context, just the state changing. The timestamp was the clue: six seconds past the hour is what a schedule looks like, not a person. It was the Bestway app’s own filter timer, nine to four, still running. The heater can’t run without the pump, so when the timer stopped the pump at four, it stopped the heating as well, including, as it would turn out, any 15:30 comfort top-up still in progress.&lt;/p&gt;
&lt;p&gt;A quick test at 16:04 settled the rest: switching heat on from Home Assistant starts the pump by itself. So the app timer went (I disabled it on my phone), and filtration moved into Home Assistant: same nine-to-four window, but the pump is never switched off while the thermostat is in heat mode. A run that finishes after four stops the pump a minute later.&lt;/p&gt;
&lt;p&gt;One more thing from that minute: as the pump stopped, the reported water temperature jumped from 29 °C to 34 °C in sixty-five seconds. The water didn’t warm five degrees; the sensor lost its flow. Nothing now acts on the temperature straight after a pump change, and it’s on the trial list to watch.&lt;/p&gt;
&lt;h2 id=&quot;the-first-cycle-was-ano&quot;&gt;The first cycle was a no&lt;/h2&gt;
&lt;p&gt;At 16:14 I turned solar mode on.&lt;/p&gt;
&lt;p&gt;The spa was heating from the earlier test. The five-minute spare-power figure was about &lt;strong&gt;minus four kilowatts&lt;/strong&gt;: the aircon was eating everything the roof could make and then some. The rule’s view was clear. There was no spare power, so there should be no spa.&lt;/p&gt;
&lt;p&gt;It waited, because the heater had to have run for twenty minutes first. At &lt;strong&gt;16:25:00&lt;/strong&gt;, the first five-minute check after that, the heat-off rule fired and the spa went to idle. At &lt;strong&gt;16:26:00&lt;/strong&gt;, one minute later, the filter rule noticed it was after four with the heat off and stopped the pump.&lt;/p&gt;
&lt;p&gt;Nobody touched anything. It was the first thing the system did on its own, and it was to say no, correctly, on exactly the kind of day that started all this.&lt;/p&gt;
&lt;h2 id=&quot;smaller-traps-for-therecord&quot;&gt;Smaller traps, for the record&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The Envoy reports kilowatts, not watts.&lt;/strong&gt; The brief expected watts, so a 2,400 threshold would have needed 2,400 kW before it did anything. It’s moot now that the Envoy isn’t in the formula, but it’s the same units trap as the car alert, in the same house, again.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The server region was the non-obvious one&lt;/strong&gt;, as warned. The brief said try Europe first; it was the US.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A restart resets the heater’s “on since” time.&lt;/strong&gt; The brief worried that would cut a minimum run short. It can only &lt;em&gt;lengthen&lt;/em&gt; one: the twenty-minute clock starts again. Relay protection holds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A trigger that only fires on a crossing misses a value that’s already high.&lt;/strong&gt; If there’s already plenty of spare power when solar mode is switched on, nothing crosses anything. Every rule re-checks every five minutes as well.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hysteresis wider than the load.&lt;/strong&gt; Start at +2,400 W, stop at −300 W: a 2.7 kW gap, more than the heater’s ~2 kW, so switching the heater on can’t by itself trip the stop rule.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Helpers created without&lt;/strong&gt; &lt;strong&gt;initial:&lt;/strong&gt;, which would reset every threshold on restart. The brief flagged that one itself, which is the most useful thing a brief can do.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The dashboard took three passes.&lt;/strong&gt; A first cut; then a full-width two-by-two row of graphs; then three balanced columns, each graph sitting next to the thing it explains, so the whole story fits on one screen. Then I swapped two cards, because it’s my dashboard.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-it-looks-likenow&quot;&gt;What it looks like now&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Six automations and a script&lt;/strong&gt;: heat on, heat off, the 15:30 floor, target sync, the filter schedule, and an alert if the spa drops off the cloud for half an hour, because a cloud-only integration that fails silently means a cold spa.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One number to change.&lt;/strong&gt; The target temperature on the dashboard &lt;em&gt;is&lt;/em&gt; the spa’s setpoint; change it and it’s pushed to the spa. 33 °C, with a 31 °C floor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Every threshold on the dashboard&lt;/strong&gt;, not in code: 2,400 W to start, 300 W to stop, battery first to 95 %.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A spa dashboard&lt;/strong&gt; with the spa, the house’s power right now, the day’s graphs, and when each rule last fired.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;An eleven-point trial checklist&lt;/strong&gt;, each item with how to check it and what counts as a pass. Two items are ticked.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-honestpart&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;The other pieces in this series each ended on what actually unblocked the thing. Delegation, for the dashboard. The earlier projects already existing, for the car alert. For the MAGI panel, a fictional interface that forced every number on screen to justify itself. This one lands somewhere less flattering.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;I specified it wrong, and the most useful thing the agent did was not do what I said.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Not by overruling me: every call that mattered stayed mine. Battery first was mine. Heating regardless when the grid pays us to import was mine. Moving filtration off the app was mine. But the brief I wrote in the morning described a house without a battery, and it was confident, well-organised, and cleanly formatted. It’s worth being precise about why. The chat wasn’t a worse AI; it was the same one, with less context. Everything it knew about the house came from me in that one conversation. The Claude Code side had months of accumulated notes, the memory of every previous project in this house, and a live connection to the system. The pushback didn’t come from anything clever. It came from that context, and from reading the house before building on the brief, which is what the brief’s own first task said to do. I wrote “verify against the live system” at the bottom of a document that hadn’t.&lt;/p&gt;
&lt;p&gt;I think that’s the shift, and it’s a bit uncomfortable. The bottleneck used to be execution, and the series has been about that bottleneck going away. With execution cheap, the weak link moves upstream to the direction itself, to the brief written at a distance, from memory, about a system that’s changed since you last thought about it, and to wherever the context lives. The same model gave me a confident wrong design in one window and caught it in the other. An agent that does exactly what the brief says is dangerous in proportion to how good the brief looks.&lt;/p&gt;
&lt;p&gt;And the other honest thing: it hasn’t done its job yet. What’s proven is the stop side, the target sync, and control. The thing the whole project is &lt;em&gt;for&lt;/em&gt;, heating on a sunny afternoon from sunshine the battery didn’t need, hasn’t happened. The first time it ran, the right answer was no. I’ll know it works when a mild, sunny day comes along with the battery full by lunch, and the checklist gets its third tick.&lt;/p&gt;
&lt;p&gt;Until then I have a spa that’s very good at declining to heat, and a brief I’ve kept as a reminder.&lt;/p&gt;
&lt;p&gt;Write the brief. Then let something that can see the house tell you what’s wrong with it.&lt;/p&gt;
&lt;h2 id=&quot;under-the-hood-if-you-want-to-buildthis&quot;&gt;Under the hood: if you want to build this&lt;/h2&gt;
&lt;h3 id=&quot;what-it-runson&quot;&gt;What it runs on&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.home-assistant.io/&quot;&gt;Home Assistant&lt;/a&gt; 2026.9.3: runs everything; helpers and automations are UI/storage, not YAML packages.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/cdpuk/ha-bestway&quot;&gt;ha-bestway&lt;/a&gt; v1.11.1, via &lt;a href=&quot;https://hacs.xyz/&quot;&gt;HACS&lt;/a&gt;: the Bestway cloud integration; AirJet on the V02 (AWS IoT) backend. Cloud-only.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.home-assistant.io/integrations/tessie/&quot;&gt;Tessie&lt;/a&gt;: Powerwall grid, battery and charge readings, &lt;strong&gt;all three from one poll&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.home-assistant.io/integrations/amberelectric/&quot;&gt;Amber Electric&lt;/a&gt;: import and feed-in prices, for the negative-price rules.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.home-assistant.io/integrations/statistics/&quot;&gt;Statistics&lt;/a&gt; + &lt;a href=&quot;https://www.home-assistant.io/integrations/template/&quot;&gt;Template&lt;/a&gt;: the spare-power sensor and its 5-minute mean.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.home-assistant.io/integrations/ntfy/&quot;&gt;ntfy&lt;/a&gt;: the “spa’s gone offline” push.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.home-assistant.io/integrations/enphase_envoy/&quot;&gt;Enphase Envoy&lt;/a&gt;: &lt;strong&gt;deliberately not used&lt;/strong&gt; in the formula (see Gotchas).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Load-bearing upstream issue: &lt;a href=&quot;https://github.com/cdpuk/ha-bestway/issues/121&quot;&gt;ha-bestway#121&lt;/a&gt;. Sharing the app’s visitor ID breaks control with error 10001.&lt;/p&gt;
&lt;h3 id=&quot;decisions&quot;&gt;Decisions&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Battery first.&lt;/strong&gt; Charging counts as spare only at ≥ 95 %; discharging always counts as a shortfall. Otherwise the spa quietly drains the battery that covers your evening.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stop on “spare power”, not on grid import.&lt;/strong&gt; With a home battery, import barely moves when solar dips; the battery absorbs it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Negative &lt;em&gt;import&lt;/em&gt; price → heat regardless&lt;/strong&gt;, and heat-off is suppressed. You’re being paid to use power.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Negative &lt;em&gt;feed-in&lt;/em&gt; on its own only lowers the bar to start&lt;/strong&gt;, with the battery full and the sun up. Solar is probably being curtailed, so the real surplus doesn’t show in any sensor. Heat-off still catches a deficit after 20 minutes, and a failed attempt costs about 0.7 kWh from the battery, which then drops below 95 % and blocks the next try.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Min run 20 min, min off 10 min; start +2,400 W, stop −300 W.&lt;/strong&gt; The 2.7 kW gap is bigger than the ~2 kW heater, so switching the heater on can’t trip the stop rule.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fixed setpoint, toggle the heat mode&lt;/strong&gt;, and only write the setpoint when it differs. Fewer cloud writes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filtration in Home Assistant, not the app.&lt;/strong&gt; The app’s timer stops the pump on schedule, and the pump stopping stops the heating.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ha-bestway over&lt;/strong&gt; &lt;a href=&quot;https://github.com/M1R4G376/New_bestway_spa&quot;&gt;&lt;strong&gt;New_bestway_spa&lt;/strong&gt;&lt;/a&gt; (kept as the fallback) and over the ESP8266 local-control mod, which has no documented support for 2024+ pumps and voids the warranty. The price is cloud dependence, hence the offline alert.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;gotchas&quot;&gt;Gotchas&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Add the spa with the app’s &lt;strong&gt;Share the device QR&lt;/strong&gt;, not the app’s own user ID: reads work, every write fails with 10001. The QR gives HA its own guest user. Don’t remove it in the app.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;There’s no heater switch&lt;/strong&gt; on the AirJet V02. The heater is the climate entity’s hvac_mode.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Disable the app’s timers.&lt;/strong&gt; An on-the-hour off with no Home Assistant context in the logbook is an app schedule.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don’t mix a local meter with a cloud battery reading&lt;/strong&gt;: sampling skew gave −8,721 W.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Check signs and units live.&lt;/strong&gt; Tessie: grid + import, battery + discharging, in kW. HA’s Amber integration flips feed-in, so &amp;lt; 0 means exporting costs money.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The water temperature jumps when the pump stops&lt;/strong&gt; (29 → 34 °C in 65 s). Don’t act on it straight after a pump change.&lt;/li&gt;
&lt;li&gt;A house load swinging 9–13 kW hides a 2 kW step. &lt;strong&gt;Prove control in the app.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;the-spare-power-sensor&quot;&gt;The spare-power sensor&lt;/h3&gt;
&lt;p&gt;The core of it. Mine is a UI template helper; this is the YAML form, written from the same formula rather than exported, so treat it as reviewed rather than run:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;state: &amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {% set k = 1000 %}  {# 1000 if your sensors report kW, 1 for W #}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {% set grid = states(&apos;__GRID_POWER__&apos;) | float * k %}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {% set batt = states(&apos;__BATTERY_POWER__&apos;) | float * k %}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {% set charging = [0 - batt, 0] | max %}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {% set discharging = [batt, 0] | max %}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {% set full_enough = states(&apos;__BATTERY_SOC__&apos;) | float&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                       &amp;gt;= states(&apos;input_number.spa_battery_soc_min&apos;) | float(95) %}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {{ (0 - grid + (charging if full_enough else 0) - discharging) | round(0) }}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Positive is spare, negative is shortfall. A statistics helper takes its 5-minute mean, and the rules trigger on the mean, never the raw value.&lt;/p&gt;
&lt;h3 id=&quot;take-it-withyou&quot;&gt;Take it with you&lt;/h3&gt;
&lt;p&gt;The full set is on GitHub at &lt;a href=&quot;https://github.com/angusmaul/homelab-recipes/tree/main/spa-solar-heating&quot;&gt;angusmaul/homelab-recipes&lt;/a&gt;: the helpers, the six automations and the script, and a deploy script that refuses to run with anything unsubstituted or missing, then shows a diff before it writes. There’s a README written for people and for AI agents, and it’s MIT licensed.&lt;/p&gt;
&lt;p&gt;What to substitute: the spa’s climate entity and filter switch, your battery’s grid/battery/charge sensors, your Amber price sensors (or drop those two rules), and a notifier. Proven live so far: control, heat-off on a real deficit, the filter following it, and target sync. Not yet: heat-on from spare solar, the negative-price hold, the 15:30 floor, and a restart.&lt;/p&gt;
</content:encoded><category>ai-agents-in-action</category><category>solar-energy</category><category>homelab</category><category>home-automation</category><category>home-assistant</category><category>Writing</category></item><item><title>Project: The Cipher</title><link>https://geekconsulting.au/projects/the-cipher/</link><guid isPermaLink="true">https://geekconsulting.au/projects/the-cipher/</guid><description>A Gotham-esque noir escape room in the browser: five nights in the city of Greymoor, one case, rendered with WebGPU.</description><pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate><category>Project</category><category>Games</category></item><item><title>Project: Nine Lives &apos;84</title><link>https://geekconsulting.au/projects/nine-lives-84/</link><guid isPermaLink="true">https://geekconsulting.au/projects/nine-lives-84/</guid><description>A three.js remake inspired by Alley Cat (1983): 2D gameplay, a 3D toy-box alley.</description><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><category>Project</category><category>Games</category></item><item><title>Project: Sopwith 3D</title><link>https://geekconsulting.au/projects/sopwith-3d/</link><guid isPermaLink="true">https://geekconsulting.au/projects/sopwith-3d/</guid><description>A faithful port of the 1984 Sopwith simulation, rendered in three.js.</description><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><category>Project</category><category>Games</category></item><item><title>Project: EDGE//RUN Solitaire</title><link>https://geekconsulting.au/projects/edgerun-solitaire/</link><guid isPermaLink="true">https://geekconsulting.au/projects/edgerun-solitaire/</guid><description>Cyberpunk Klondike in three.js, with switchable themes.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate><category>Project</category><category>Games</category></item><item><title>The Salesforce Org Argues Back: Building a Data Seeder That Reads Your Validation Rules</title><link>https://geekconsulting.au/writing/salesforce-org-argues-back/</link><guid isPermaLink="true">https://geekconsulting.au/writing/salesforce-org-argues-back/</guid><description>I filed the AI integration as an epic on day two of this project. It took a year to land — and the part I’d actually been dreading arrived in the twenty-three minutes between two commits.</description><pubDate>Mon, 17 Aug 2026 11:13:39 GMT</pubDate><content:encoded>&lt;p&gt;Every Salesforce project starts in the same room: a fresh sandbox with nobody in it. No accounts, no contacts, no opportunities, no assets. You cannot demo a report with no rows in it. You cannot test a flow that has nothing to fire on. You cannot show a customer their own process working when the screen is a polite empty state saying &lt;em&gt;No records to display&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;So the first day of every engagement is spent making up people. A spreadsheet of invented companies, Data Loader, an insert that fails, a column deleted, another insert, a different failure, another column deleted, and eventually a few hundred rows of something that technically exists. Then the next project starts and you do it again from scratch, because the data you made was shaped for the &lt;em&gt;last&lt;/em&gt; org’s configuration and this one has different ideas.&lt;/p&gt;
&lt;p&gt;I have been doing that for years. It is not hard. It is not interesting. It is the definition of a someday problem: obviously worth automating, never quite worth the afternoon.&lt;/p&gt;
&lt;p&gt;I’ve written up projects like this before — &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;a Grafana data platform&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-2/&quot;&gt;the homelab dashboard I’d abandoned three times&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-3/&quot;&gt;an alert that paid for the whole lab&lt;/a&gt;, and &lt;a href=&quot;https://geekconsulting.au/writing/proxmox-magi-panel/&quot;&gt;a Proxmox cluster with a MAGI panel bolted to it&lt;/a&gt;. This one isn’t the house. It’s the day job, which makes the someday excuse considerably harder to justify — I have been doing something the slow way and invoicing for it.&lt;/p&gt;
&lt;p&gt;Same principle either way: &lt;strong&gt;I supply direction, the agent supplies execution.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The tool is on GitHub: &lt;a href=&quot;https://github.com/angusmaul/salesforce-sandbox-data-seeder&quot;&gt;&lt;strong&gt;github.com/angusmaul/salesforce-sandbox-data-seeder&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/salesforce-org-argues-back/1.jpg&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;h2 id=&quot;why-this-is-harder-than-generate-some-fake-companies&quot;&gt;Why this is harder than “generate some fake companies”&lt;/h2&gt;
&lt;p&gt;Generating a thousand plausible businesses is trivial. Faker does it in about four milliseconds. The difficulty is that a Salesforce org is not a schema — it is a schema that somebody has been &lt;em&gt;configuring&lt;/em&gt; for six years, and every one of those configuration decisions is a landmine your generated record has to walk past.&lt;/p&gt;
&lt;p&gt;A modest developer org hands back &lt;strong&gt;802 creatable objects&lt;/strong&gt;. Twelve of them are custom — the things somebody in this business actually built. The other 790 are Salesforce’s own plumbing, and none of them are labelled as such.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/salesforce-org-argues-back/2.jpg&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;p&gt;Underneath that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Picklist dependencies&lt;/strong&gt; are real and enforced. Country controls State, and the mapping is not published as a list — it’s a base64 validFor bitmap you have to decode per value to find out which states are legal under which country.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Restricted picklists&lt;/strong&gt; reject anything outside the active set, so a perfectly sensible value that used to be valid is now an error.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Precision and scale&lt;/strong&gt; are per field. 9477.9 is a fine number and it will not go into a Number(4,1), whose ceiling is 999.9.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lookups&lt;/strong&gt; need real IDs, which means the objects have to be loaded in dependency order, and each child needs the parent IDs the previous step actually produced.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compound address fields&lt;/strong&gt; are quietly derived — send both BillingState and BillingStateCode and the org tells you off for a mismatched integration value.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All of that is discoverable. Tedious, but discoverable: it’s in the metadata, and a sufficiently patient program can read it.&lt;/p&gt;
&lt;p&gt;And then there are validation rules, which are a different kind of problem entirely.&lt;/p&gt;
&lt;p&gt;A validation rule is a formula somebody wrote, in this org, for a business reason that isn’t recorded anywhere in the schema. The metadata will tell you that Active__c is a picklist of Yes and No. Nothing in the metadata tells you that in &lt;em&gt;this&lt;/em&gt; org, an Account whose Active__c is No cannot be created at all. That fact lives in a formula and a sentence of English, and the only way you learn it is by having your record thrown back at you.&lt;/p&gt;
&lt;p&gt;Everything else about seeding a Salesforce org is a data problem. Validation rules are a reading-comprehension problem.&lt;/p&gt;
&lt;h2 id=&quot;two-false-starts-six-monthsapart&quot;&gt;Two false starts, six months apart&lt;/h2&gt;
&lt;p&gt;The repository is honest about how this went. The first commit is August 2025: a CLI that discovers the data model and fills it with Faker output. It worked, in the sense that it ran.&lt;/p&gt;
&lt;p&gt;The very next commit, the same day, files an epic called &lt;strong&gt;AI-Integration&lt;/strong&gt;. I knew from day two what the tool needed. Then nothing happened for six months.&lt;/p&gt;
&lt;p&gt;In February 2026 the epic shipped: field classification via Claude Haiku, analysing the org’s field schemas once and mapping them into a semantic library of 8 categories and 66 subcategories, so the generated data is &lt;em&gt;correlated&lt;/em&gt; rather than merely plausible. A contact in Engineering gets an engineering job title. An Australian address gets an Australian phone format. A name and an email address belong to the same person. Two and a half thousand lines, 141 tests, one commit.&lt;/p&gt;
&lt;p&gt;That was a genuine improvement and it fixed the wrong problem. The data got much more realistic and it still didn’t go in. Because “realistic” and “acceptable to this org” are unrelated properties, and I had spent the whole effort on the first one.&lt;/p&gt;
&lt;p&gt;Then nothing happened for six more months.&lt;/p&gt;
&lt;h2 id=&quot;the-24run&quot;&gt;The 24% run&lt;/h2&gt;
&lt;p&gt;What broke the deadlock in August wasn’t an idea. It was running the thing against a real sandbox and reading what came back.&lt;/p&gt;
&lt;p&gt;A 275-record load succeeded at &lt;strong&gt;24%&lt;/strong&gt;. 67 records in, 208 rejected. Account limped in at 55 of 100 and Lead at 12 of 75, and Contact came in at &lt;strong&gt;0 out of 100&lt;/strong&gt; — not degraded, not partial, zero. Every single one refused.&lt;/p&gt;
&lt;p&gt;The logs sorted the wreckage into two deterministic classes almost immediately.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;166 of roughly 208 failures were state and country integrity.&lt;/strong&gt; And the reason is the detail I keep turning over, because it’s not a bug so much as an indictment. The app &lt;em&gt;already decoded the org’s country-to-state dependency map&lt;/em&gt; — pulled the validFor bitmaps during field analysis, decoded them, and stashed the result on the session. That work was being done on every run. It was then never read. Generation was reaching for a hardcoded four-country dictionary that had been written months earlier as a stopgap, and the elaborate, correct, org-specific answer was sitting untouched in memory the whole time.&lt;/p&gt;
&lt;p&gt;Sitting under that was the reason Contact scored zero: a function called extractAddressPrefix, which is meant to tell MailingCity from ShippingCity, returned Shipping for Mailing* fields. A copy-paste error, one word wrong, which guaranteed that every address lookup on Contact missed. Accounts have Billing and Shipping addresses and limped in at some percentage. Contacts have Mailing addresses and could never have worked at all.&lt;/p&gt;
&lt;p&gt;I would never have found that by reading the code. It looks right. You find it by asking why one object scored zero when a structurally identical one didn’t, and being unwilling to accept “probably something to do with addresses” as an answer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;46 further failures were numeric range&lt;/strong&gt;, which is the boring one and was fixed the boring way: field analysis now stores precision, scale and digit count, and values get clamped to the field’s maximum representable number before they’re sent.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; Before you add capability, check what the program already knows and isn’t asking itself. The most expensive bug in that run was a correct answer computed on every single run and read by nothing.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;remediation-you-can-writedown&quot;&gt;Remediation you can write down&lt;/h2&gt;
&lt;p&gt;That fixed the deterministic classes. The next increment dealt with the fact that a failed insert was, until that afternoon, simply logged and abandoned.&lt;/p&gt;
&lt;p&gt;Salesforce error codes are unusually well-behaved: most of them tell you precisely what to do. So each failure now gets up to two remediated resubmissions, with the remediation chosen by code:&lt;/p&gt;
&lt;p&gt;STRING_TOO_LONG truncates to the field’s real length. INVALID_OR_NULL_FOR_RESTRICTED_PICKLIST re-picks from the active values. NUMBER_OUTSIDE_VALID_RANGE clamps, then drops the field if clamping didn’t satisfy it. REQUIRED_FIELD_MISSING regenerates the named fields — unless it’s an unfillable required lookup, in which case the record is dropped rather than retried into the same wall. Duplicate-rule violations uniquify the identity fields, and do it email-aware, so a mangled address is still an address. Row locks get one unchanged resubmit, because a row lock is a timing complaint and not a data complaint.&lt;/p&gt;
&lt;p&gt;The piece I like most is the &lt;strong&gt;learned blocklist&lt;/strong&gt;. If removing a particular field is what rescued at least three records, and those recoveries account for at least 80% of everything recovered this pass, that field stops being generated for the rest of the session. The tool notices what it keeps getting wrong and stops doing it, which means object number seven doesn’t repeat the mistake object number two already paid for.&lt;/p&gt;
&lt;p&gt;One error code was deliberately left terminal: FIELD_CUSTOM_VALIDATION_EXCEPTION.&lt;/p&gt;
&lt;p&gt;Every other code is a mechanical instruction. That one is a human being telling you no. There is no generic remediation for it, because the correct action depends entirely on what the formula says — and the formula is not in the error. All you get back is somebody’s error message.&lt;/p&gt;
&lt;h2 id=&quot;the-hundred-percent-that-provednothing&quot;&gt;The hundred percent that proved nothing&lt;/h2&gt;
&lt;p&gt;With those two increments in, the same 275-record load went in clean. 100%. Account 100, Contact 100, Lead 75, in thirteen seconds.&lt;/p&gt;
&lt;p&gt;Which proved almost nothing, and I knew it while I was looking at it. That org’s Accounts had no awkward rules on them. A green run against a permissive org is not evidence that the tool handles a strict one; it’s evidence that you picked an easy org. Every consultant reading this has inherited the other kind — the org with eleven rules nobody documented, written by someone who left in 2019.&lt;/p&gt;
&lt;p&gt;So I went into Setup and wrote a validation rule specifically to break my own tool. Accounts must have Active__c set to Yes at creation, or they don’t get created.&lt;/p&gt;
&lt;p&gt;The next run came back at &lt;strong&gt;72%&lt;/strong&gt;. Contact still 100, Lead still 75, and Account collapsed to 23 of 100**–77 records**, failed terminally, every one of them with the same sentence attached:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Active is required to be Yes when creating accounts&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;A rule a competent human satisfies in about four seconds of reading, and the tool could not touch it.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; A passing test against convenient conditions is a measurement of the conditions, not the tool. If you can’t find something that breaks it, go and build the thing that breaks it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;the-rule-that-says-no-and-wont-saywhy&quot;&gt;The rule that says no and won’t say why&lt;/h2&gt;
&lt;p&gt;This was the blocker, and it had been the blocker long before I wrote that rule down. Not “difficult” — blocked. Every other failure class had a mechanical answer and this one required understanding a formula written by a stranger. It’s why the project had sat for six months twice: I could see the shape of the work and could not see the shape of the solution.&lt;/p&gt;
&lt;p&gt;The solution turned out to hinge on one inversion, and once you’ve seen it the whole thing collapses into something small.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A validation rule’s&lt;/strong&gt; &lt;strong&gt;errorConditionFormula evaluates to TRUE when the record should be rejected.&lt;/strong&gt; It isn’t a description of a valid record. It’s a description of an invalid one. So you don’t ask a model “what does this rule want?” — a vague question with an essay for an answer. You tell it: &lt;em&gt;this formula rejects when true; produce constraints that make it false.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;That’s the whole trick, and it changes the task from open-ended interpretation into a translation with a right answer.&lt;/p&gt;
&lt;p&gt;The second decision matters just as much, and it’s the one I’d defend hardest: &lt;strong&gt;the model is not allowed to answer freely.&lt;/strong&gt; It gets five constraint types and nothing else.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fixedValue      — field must equal value&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;picklistSubset  — field must be one of these values&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;range           — numeric bounds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;notNull         — field must be populated&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;maxLength       — string length cap&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Anything the model cannot express in that vocabulary — cross-field logic, record types, user context, anything requiring the model to be clever — it is instructed to name under unsupported rather than improvise. Those rules are then reported to the user, alongside a suggestion to use the disable-and-restore fallback for those specific rules if they want them out of the way.&lt;/p&gt;
&lt;p&gt;So the model does comprehension, in a domain where comprehension is genuinely required, and hands back a data structure. Ordinary deterministic code does the enforcement: applyConstraints() walks the record, applies each constraint, skips anything Salesforce won’t accept as a write anyway — formula fields, auto-numbers, non-createable fields — and returns human-readable notes about every change it made. You can hit /api/validation-rules/list/:sessionId and see every rule the org has, next to exactly what was derived from it.&lt;/p&gt;
&lt;p&gt;The retry loop then gets one carefully hedged upgrade. FIELD_CUSTOM_VALIDATION_EXCEPTION is promoted from terminal to adjust-and-resubmit — &lt;strong&gt;but only when constraints exist for that object and aren’t already satisfied.&lt;/strong&gt; If the rule was one the model couldn’t express, or the record already complies and failed anyway, it stays terminal. Without that guard, a rule nobody can interpret would burn every retry pass resubmitting identical records to an org that has already made its position clear.&lt;/p&gt;
&lt;p&gt;The test for all this is the failure I’d manufactured: Active__c=No fails, constraints are applied, the resubmit succeeds. 169 of 169 server tests green.&lt;/p&gt;
&lt;p&gt;Then the same 275-record load, against the same org, with my hostile rule still active and untouched: &lt;strong&gt;275 of 275&lt;/strong&gt;. The two runs are &lt;strong&gt;twelve minutes apart&lt;/strong&gt; in the logs — 72% either side of a change that reads, in the load log, as nothing at all. Same records requested, same rule enforcing, same org. The tool just read the rule this time.&lt;/p&gt;
&lt;p&gt;The whole interpreter is 127 lines. The constraint applier is 93. That is the entirety of the thing that had blocked the project for a year.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; When you put a model in a pipeline, shrink its job until the output is a data structure you can validate, and give it a way to say “I can’t express this.” A model with five allowed answers and an escape hatch is inspectable. A model asked to be helpful is a second bug you can’t step through in a debugger.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The commits either side of that work are &lt;strong&gt;twenty-three minutes apart&lt;/strong&gt;. I’ve stared at that gap a fair bit since. It is not that the code was hard — 127 lines is not hard. It’s that arriving at “invert the formula and constrain the vocabulary” required knowing that validation failures were a distinct, terminal, 77-record failure class, and knowing &lt;em&gt;that&lt;/em&gt; required the retry loop from an hour earlier, which required the failure taxonomy from an hour before that. The blocker was never the typing. It was two hundred failures nobody had read.&lt;/p&gt;
&lt;h2 id=&quot;the-slow-model-that-found-threebugs&quot;&gt;The slow model that found three bugs&lt;/h2&gt;
&lt;p&gt;One more thread worth pulling, because it’s the best bug of the day and it isn’t really an AI bug at all.&lt;/p&gt;
&lt;p&gt;Field classification had been hardcoded to Anthropic, via a server environment variable. That’s fine for me and useless for the actual use case, because this tool reads your org’s field names, picklist values, validation-rule formulas and error messages — which is to say, a fairly complete description of a client’s business logic. Plenty of engagements would never permit that leaving the building, and a tool that requires it is a tool that doesn’t get used.&lt;/p&gt;
&lt;p&gt;So it became bring-your-own: Anthropic, any OpenAI-compatible endpoint — OpenAI, Groq, OpenRouter, LM Studio, vLLM — or a native Ollama running on your own hardware, all configured in the UI rather than in a server env var. Failures return null and generation falls back to pattern-based rules, so a missing or broken provider degrades the output instead of breaking the run.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/salesforce-org-argues-back/3.jpg&quot; alt=&quot;The generation plan: 275 records across three objects in dependency order, and the banner stating exactly what goes to the model.&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;The generation plan: 275 records across three objects in dependency order, and the banner stating exactly what goes to the model.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The banner in that screenshot is the whole argument in one sentence: &lt;em&gt;field metadata (not your data) is sent to your local Ollama model for classification.&lt;/em&gt; The org’s schema is the thing being analysed, the records never are, and on that setup neither leaves the building. Note the amber &lt;strong&gt;Production&lt;/strong&gt; badge too — a Developer Edition org is, in Salesforce’s own taxonomy, a production org, which is exactly why the tool flags it rather than quietly proceeding.&lt;/p&gt;
&lt;p&gt;And then a local gpt-oss:20b took &lt;strong&gt;77 seconds&lt;/strong&gt; to classify a set of fields, and three separate bugs fell out of a latency regime nothing had ever been in before.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Next.js’s rewrite proxy severs connections at 30 seconds by default.&lt;/strong&gt; The browser reported socket hang up. The server, meanwhile, was completely fine and finished the job.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A session lost-update race.&lt;/strong&gt; The long-running analysis endpoints hold a reference to the session object across a multi-minute await, while the wizard’s ordinary PUTs replace that object in the store. Last writer wins. In the observed run, the finished AI plan was overwritten by a stale snapshot &lt;em&gt;seconds after being successfully cached&lt;/em&gt; — so the work completed, was saved, and then vanished.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The UI gave up when the request died&lt;/strong&gt;, despite the server completing the analysis regardless.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;All three diagnosed by reading a real run’s logs alongside the session store, rather than by reasoning about what ought to have happened. The fixes are unglamorous: raise the proxy timeout, have PUT mutate the stored object in place instead of replacing it, have long-running endpoints re-fetch the session before writing their results back, and have the UI poll the cached-plan endpoint instead of assuming a dead socket means dead work.&lt;/p&gt;
&lt;p&gt;Not one of those three is a bug about AI. They’re a read-modify-write race and two timeout assumptions, all of which had been sitting there since the beginning, all of them invisible while every call returned in two seconds.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; Supporting slow inference is a load test. A model that takes 77 seconds instead of 2 doesn’t just make things slower — it drags your code into a latency regime where every implicit timeout and every race you’d been getting away with becomes reproducible.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;smaller-traps-for-therecord&quot;&gt;Smaller traps, for the record&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;jsforce’s sObject Collections call fails wholesale past 200 records&lt;/strong&gt; unless you pass allowRecursive: true, at which point it chunks for you. Discovered the direct way.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Result-to-record pairing was using&lt;/strong&gt; &lt;strong&gt;indexOf&lt;/strong&gt;, which quietly does the wrong thing the moment two generated records are identical. The API returns results in input order; positional mapping is both correct and a prerequisite for retrying anything.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Never write&lt;/strong&gt; &lt;strong&gt;GeocodeAccuracy&lt;/strong&gt;, and suppress the text State/Country twins whenever the *Code fields exist — Salesforce derives the text itself and objects loudly if you send both.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A&lt;/strong&gt; &lt;strong&gt;StateCode must never be written without its&lt;/strong&gt; &lt;strong&gt;CountryCode.&lt;/strong&gt; There’s now a backstop for this in both record loops, because the rule was getting violated from two different code paths that each thought the other was handling it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Credentials are stored server-side at&lt;/strong&gt; &lt;strong&gt;0600 and never returned to the browser&lt;/strong&gt; — but they’re plaintext JSON on disk, which is written down in the README rather than glossed over. It’s a sandbox tool. It should still say what it is.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-it-looks-likenow&quot;&gt;What it looks like now&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A seven-step wizard&lt;/strong&gt; — connect, discover, select, configure, preview, execute, results — with live progress over WebSocket and a results dashboard that will hand you the whole run as a ZIP.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Saved org connections&lt;/strong&gt;, so authenticating once via an External Client App means one click on every subsequent session.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Any AI provider you like, including none.&lt;/strong&gt; Hosted, local, or pattern-based fallback. The org’s metadata never has to leave your network.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Validation-rule awareness&lt;/strong&gt;, with the rules it couldn’t interpret named explicitly rather than silently skipped.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;169 passing server tests&lt;/strong&gt;, including the exact 77-record failure that started the last increment, frozen as a regression case.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deployable&lt;/strong&gt; via Docker Compose, LXC with systemd units, or bare Node.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-honestpart&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;Every one of these write-ups ends with me working out what actually changed, and it’s never the thing I expected going in. With the dashboard it was that the tedious middle could be delegated. With the energy alert it was that the previous projects existing made the next one trivial. With the MAGI panel it was that an aesthetic constraint had done the engineering work.&lt;/p&gt;
&lt;p&gt;This one lands somewhere I didn’t like at first.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The AI’s job got smaller, and that’s what made it work.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The instinct — mine, certainly, and I think most people’s — is to hand the model more. When the generated data kept getting rejected, my first thought was that the model should generate better data. February’s version is that thought, fully committed: an AI plan for every field, semantic categories, correlation maps, 2,456 lines. It made the data lovely and it did not get one extra record into the org.&lt;/p&gt;
&lt;p&gt;What actually shipped does the opposite. The model never sees a record. It never generates a value. It reads a set of formulas and a field list and answers in five permitted shapes, with an explicit option to decline. Everything downstream — picking, clamping, truncating, uniquifying, retrying, blocklisting — is deterministic code you can step through at three in the morning.&lt;/p&gt;
&lt;p&gt;And the reason that design was even &lt;em&gt;available&lt;/em&gt; is that the logs existed. You can only shrink a model’s job to “translate this formula into a constraint” once you know that formula-shaped rejections are a distinct failure class of exactly 77 records, and that everything else in the pile is mechanical. In February I didn’t know that, so the AI had to be the whole feature. In August the AI is 127 lines in the middle of a loop whose actual engine is a log file.&lt;/p&gt;
&lt;p&gt;So when people say agentic engineering is moving quickly — and it is; three increments with tests in ninety-four minutes would not have happened in February — I don’t think the interesting part is that the models write better code. That’s true and it’s the least of it. The interesting part is that the cost of &lt;em&gt;going and looking&lt;/em&gt; collapsed. Running the thing against a real org, pulling two hundred failures apart, sorting them into classes, and letting each class dictate its own fix used to be an afternoon I would never spend. It’s now the cheapest step in the process, which means evidence has become cheaper than theorising.&lt;/p&gt;
&lt;p&gt;That’s the actual shift. Not that the agent is better at answering. That it’s finally cheap enough to stop guessing.&lt;/p&gt;
&lt;p&gt;Six years of hand-rolled CSVs, two abandoned attempts, and the thing that broke it open was a 24% success rate and somebody willing to read all of it.&lt;/p&gt;
&lt;p&gt;Point it at a real org. Read every failure. Then ask the model one small question.&lt;/p&gt;
</content:encoded><category>software-development</category><category>salesforce</category><category>developer-tools</category><category>anthropic-claude</category><category>ai-agent</category><category>Writing</category></item><item><title>I’m 47, I’ve Watched Evangelion Too Many Times, and Now My Proxmox Cluster Has a MAGI Panel</title><link>https://geekconsulting.au/writing/proxmox-magi-panel/</link><guid isPermaLink="true">https://geekconsulting.au/writing/proxmox-magi-panel/</guid><description>I’d imagined this screen for thirty years. The idea was never the hard part — execution was, and that’s the bit I finally had help with.</description><pubDate>Sun, 02 Aug 2026 07:15:28 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/proxmox-magi-panel/1.jpg&quot; alt=&quot;&quot;&gt;&lt;/figure&gt;&lt;p&gt;I have a rack in my study with a Raspberry Pi on top of it wired to a 1280×400 ultrawide bar LCD. For a while that screen did nothing, because the status of my homelab lived where it lives for everybody: in a browser tab I had to remember to open, which meant I opened it after something had already broken.&lt;/p&gt;
&lt;p&gt;I have also watched &lt;em&gt;Neon Genesis Evangelion&lt;/em&gt; more times than a man my age should admit to in public. It aired in 1995, when I was a teenager, and the thing that stuck wasn’t the giant robots. It was the MAGI: three supercomputers — Melchior, Balthasar and Casper — built by one scientist who split her own personality across them so that they would &lt;em&gt;argue&lt;/em&gt;. They deliberate. They vote. They deadlock, and when they do the screen says so in enormous orange letters and everyone in the room has to deal with it.&lt;/p&gt;
&lt;p&gt;Thirty years later I had three machines about to become a three-node Proxmox cluster with quorum. Three computers that vote, can disagree, and need a majority before anything is allowed to happen.&lt;/p&gt;
&lt;p&gt;That isn’t a metaphor. That is what quorum &lt;em&gt;is&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;So I did the obvious ridiculous thing: named the cluster magi, named the nodes after the three, and built them a status panel that looks like the one in the show. Then it got interesting, because an anime obsession is about &lt;em&gt;aesthetics&lt;/em&gt; and a monitoring panel is about &lt;em&gt;truth&lt;/em&gt;, and those two things fight — constantly, and over things I’d never have predicted. Every collision resolved in favour of truth, usually because the agent I was building it with went and measured something I’d have been happy to eyeball.&lt;/p&gt;
&lt;p&gt;I’ve written about this homelab before: &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;turning Home Assistant into a Grafana data platform&lt;/a&gt;, &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-2/&quot;&gt;building the dashboard I’d abandoned three times&lt;/a&gt;, and &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-3/&quot;&gt;the alert that made the whole thing pay for itself&lt;/a&gt;. Same collaborator, same principle: I supply direction, the agent supplies execution.&lt;/p&gt;
&lt;h2 id=&quot;the-setup&quot;&gt;The setup&lt;/h2&gt;
&lt;p&gt;Two panels, one collector, polling the Proxmox API with a read-only, privilege-separated token.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The rack bar — 1280×400, amber on black, three unit columns, live CPU traces. It runs on the Pi, named ibuki after Maya Ibuki, the bridge operator whose job is to read the MAGI’s status out loud. The name is doing work: that box observes the cluster through a token that can read status and &lt;em&gt;nothing else&lt;/em&gt;. It reads out loud; it has no authority.&lt;/li&gt;
&lt;li&gt;The desk console — 1080×1920, portrait, fullscreen, under a wall of security cameras. A much closer reading of the show’s visual language, and it absorbed the Uptime Kuma board that used to have its own window.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/proxmox-magi-panel/2.png&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;p&gt;One structural decision shaped everything downstream: both panels run their own collector and their own web server, locally. A panel whose job is to tell you the cluster is unhealthy must not be served &lt;em&gt;by&lt;/em&gt; the cluster. Hosted on the node, a node outage blanks the screen that would have told you about it. Hosted locally, that same outage renders as MELCHIOR 停止 in red on a panel that is still up.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Lesson: The thing that reports a failure must not depend on the thing that can fail.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The whole cost is about 1.4 small authenticated JSON GETs per second across the lab, against a node running 37 guests. Both panels, the collector and the systemd units are on GitHub: &lt;a href=&quot;https://github.com/angusmaul/magi-display/&quot;&gt;github.com/angusmaul/magi-display&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;the-panel-that-can-look-perfect-and-knownothing&quot;&gt;The panel that can look perfect and know nothing&lt;/h2&gt;
&lt;p&gt;The architecture is simple. A collector polls the Proxmox API every five seconds and writes data.json to disk. A local web server serves that directory. The page fetches data.json every five seconds and re-renders. And the clock in the header ticks in the &lt;em&gt;browser&lt;/em&gt;, once a second, on a setInterval that knows nothing about any of the above.&lt;/p&gt;
&lt;p&gt;So what does the screen look like if the collector dies?&lt;/p&gt;
&lt;p&gt;Perfect. The web server is still up, so the fetch still succeeds. The last data.json is still on disk, so the response is HTTP 200 with valid JSON in it. Every node reads online, every gauge holds its last value, and the seconds keep counting.&lt;/p&gt;
&lt;p&gt;A frozen panel and a healthy panel are, visually, the same panel.&lt;/p&gt;
&lt;p&gt;Failing to a red LINK LOST when the fetch throws catches nothing here, because nothing throws. The request is fine. The JSON is fine. The data is three hours old. So the collector stamps every payload with the time it was written, and the page checks that stamp before it believes a single number:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;if(d.updated &amp;amp;&amp;amp; (Date.now()/1000 - d.updated) &amp;gt; STALE_AFTER) d.link = false;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;STALE_AFTER is 30 seconds — six poll intervals, so one missed tick never trips it.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Lesson: “The request succeeded” is not the same claim as “this data is current.” If a payload can’t tell you when it was written, you can’t distinguish a live system from a corpse — and every animated thing on the page will vote for&lt;/em&gt; live*.*&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The two red states are deliberately different. EMERGENCY asserts the cluster is broken; LINK LOST asserts only that we can’t &lt;em&gt;see&lt;/em&gt; it. Under LINK LOST the units render 不明 (unknown) rather than 停止 (stopped), because a dead collector is no evidence about any node. At 2am that’s the difference between driving to the rack and going back to bed.&lt;/p&gt;
&lt;h2 id=&quot;every-field-isreal&quot;&gt;Every field is real&lt;/h2&gt;
&lt;p&gt;The reference frames are dense with gorgeous nonsense. CODE : 666. EXTENTION : 0256 — misspelled, in the original, in a show that ran on national television. PRIORITY : S+. It’s fantastic, and it’s a lie generator: a display covered in numbers that mean nothing trains you to stop reading the numbers that do.&lt;/p&gt;
&lt;p&gt;So the rule is nothing on screen is set dressing. The only non-data text is the frame words CURRENT and OPERATIONS, and the unit names. Everything else is a live reading:&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/proxmox-magi-panel/3.png&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;p&gt;The misspelled EXTENTION is the reference’s own error and it stays — it’s the most recognisable detail on the screen and the number under it is honest. One field didn’t survive: TIME REMAINING TO COLLAPSE, the best thing in the alert frame, which has nothing to count down to. The fourth corner carries TIME SINCE RAISED instead, a real elapsed timer on a real fault.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/proxmox-magi-panel/4.png&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;p&gt;The rule ran back through the rack panel too. An earlier layout had CODE in the data block showing failing checks &lt;em&gt;and&lt;/em&gt; a second CODE: in the footer showing something else — two fields, same name, different values, one screen. The rack panel’s footer had also carried a hardcoded CODE:601 since day one. Both gone.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Lesson: Decorative data isn’t neutral. It teaches the reader that the numbers on this screen are vibes, and that lesson generalises to the numbers that matter.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;the-permanent-emergency-i-nearlyshipped&quot;&gt;The permanent emergency I nearly shipped&lt;/h2&gt;
&lt;p&gt;Every five seconds the panel picks one of four verdicts, and that rule is fixed in code:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const crit = lost || !d.quorate || d.priority === &quot;S+&quot;;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Quorum is tested first, before anything is painted, because the colours everything else is drawn in depend on the outcome. Losing quorum can therefore never render as PROTECTED or even as amber DILEMMA — not even when every other signal is healthy. All five states were checked against synthetic payloads, including “quorum lost and nothing else wrong at all.”&lt;/p&gt;
&lt;p&gt;A missing quorate field is falsy, and therefore also critical — unknown quorum fails safe rather than nominal. That choice nearly shipped a disaster of its own: a collector that doesn’t emit the field puts the console into a &lt;em&gt;permanent&lt;/em&gt; EMERGENCY the moment it sees real data, which is exactly what the collector did. The quorum call was gated behind an environment variable the console hadn’t been given. Caught before first deploy, by testing the collector’s actual output rather than assuming it agreed with the panel about the payload shape.&lt;/p&gt;
&lt;p&gt;Same trap in the service list: if Kuma is unreachable the collector emits an explicit null, not an empty array, because services: [] renders as seven silent green rows — seven services, all fine, none of them checked.&lt;/p&gt;
&lt;p&gt;When you can’t distinguish “fine” from “unknown”, make the code land on &lt;em&gt;unknown&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;drama-versus-diagnosis&quot;&gt;Drama versus diagnosis&lt;/h2&gt;
&lt;p&gt;The show’s alert frame is a full-screen red takeover: radar rings, rotating wireframe, an enormous centre word. Best-looking thing in the reference, and nearly useless, because it shows four fields. At the moment things are worst you lose the schematic, the vitals and the per-unit detail. Nobody sees the drama at 03:00; the person who walks up at 08:00 needs to know &lt;em&gt;what&lt;/em&gt; broke.&lt;/p&gt;
&lt;p&gt;So it’s split. The critical &lt;em&gt;palette&lt;/em&gt; lasts the whole episode; the full-screen &lt;em&gt;takeover&lt;/em&gt; runs 30 seconds and then retires to the normal diagnostic panel, still in red.&lt;/p&gt;
&lt;p&gt;It re-fires on a change in the kind of fault, never on the failing-check count. Key it on the count and a number flickering 11 → 12 → 11 restarts the window every poll and the takeover never retires — it just sits there forever, four fields deep, being dramatic. A retry loop with better art direction.&lt;/p&gt;
&lt;p&gt;One consequence I didn’t see coming: because the panel now &lt;em&gt;stays&lt;/em&gt; in red, anything drawn with a hardcoded colour ends up blue on a red display. Canvas can’t read a CSS variable, so every colour the JavaScript writes had to be declared as a variable and read back once per render. This shipped wrong for half a day — the panel went correctly red and left three cheerful blue CPU traces on top of it.&lt;/p&gt;
&lt;h2 id=&quot;reading-the-sourceproperly&quot;&gt;Reading the source properly&lt;/h2&gt;
&lt;p&gt;The trilobe schematic — Balthasar as a pentagon at the top, Casper and Melchior as slabs below, joined into a Y by thick orange bands — isn’t eyeballed from memory. Four frames of the reference were read directly and the geometry traced off them.&lt;/p&gt;
&lt;p&gt;What came out of that: the connecting bands branch from the exact midpoint of each chamfered corner, and the chamfer carries on past the branch. Balthasar’s chamfer midpoints sit at (527,400) and (751,400); the bands leave at (525,397) and (753,397). Two and three pixels off the midpoint — which is to say, the midpoint.&lt;/p&gt;
&lt;p&gt;An earlier version joined the shapes corner to corner, which is what you’d do from memory. It’s wrong in a way that’s hard to name until you see them side by side: the junctions read as &lt;em&gt;folds&lt;/em&gt; rather than as a branching bus.&lt;/p&gt;
&lt;p&gt;I would have shipped that version, looked at it, thought “yeah, that’s the MAGI”, and moved on. It’s right because the agent went back to the actual frames and measured — the same discipline that read the 404 instead of assuming the token was fine on the dashboard build. Just pointed at a cartoon instead of an API.&lt;/p&gt;
&lt;h2 id=&quot;the-typography-rabbithole&quot;&gt;The typography rabbit hole&lt;/h2&gt;
&lt;p&gt;Evangelion’s real faces aren’t a guess: Matisse EB for the NERV interfaces, Eurostile for signage and UI. Matisse is the recognisable one — an ultra-bold Mincho, thin horizontals, thick verticals, triangular feet, chosen by Anno against the era’s fashion for sans. Both are commercial, so: stand-ins.&lt;/p&gt;
&lt;p&gt;The Japanese one resolved itself embarrassingly easily. A heavy Mincho was already installed — Noto Serif CJK JP Black, sitting in a font package, zero downloads. Found by looking at what the machine already had rather than by shopping.&lt;/p&gt;
&lt;p&gt;The Latin one is the story. A “free Eurostile” from a font-download site turned out to be a 1991 Digital Typeface Corporation clone, distributed under a name Monotype holds the trademark to, with no licence fields in the file at all. Not ambiguous — &lt;em&gt;empty&lt;/em&gt;. Rejected, and Michroma went in instead: OFL, 64 KB, close enough in spirit.&lt;/p&gt;
&lt;p&gt;That cost a full round of resizing, because Michroma runs ~35% wider than the fallback and has exactly one weight — a synthetic bold on a face that wide smears the verticals. CURRENT / OPERATIONS went from 44px bold to 34px regular, the unit labels 30 to 23, the verdict badge 26 to 19, each checked against a screenshot of the real screen rather than a browser preview.&lt;/p&gt;
&lt;p&gt;Eurostile is still named first in every font stack, so a licensed copy would take over with no code change.&lt;/p&gt;
&lt;h2 id=&quot;names-that-dowork&quot;&gt;Names that do work&lt;/h2&gt;
&lt;p&gt;Both panels are keyed on the Proxmox node name. A unit with no matching node renders dark and greyed, reading 未配備 — “not deployed”. So casper, currently a Windows box that hasn’t joined anything, sits on both screens as a dark third lobe. The day it joins the cluster it lights up by itself. No code change, no redeploy.&lt;/p&gt;
&lt;p&gt;And the desk console is the funny one: that machine is also the cluster’s QDevice, the external tiebreaker vote that lets a two-node cluster survive losing a node. So the console renders a quorum state that it is itself one third of. A dark panel is also a lost vote. That’s not a bug — the box that shows you quorum is the box whose absence changes it — but you want to know it before you read the header.&lt;/p&gt;
&lt;p&gt;The header says QUORUM HELD, not 2/3, because /cluster/status returns quorate as a bare flag with no vote counts and no mention of the QDevice. The counts exist in pvecm status, which is a shell command, and reaching it would mean giving a wall display SSH into the cluster. The panel says exactly what the API can support.&lt;/p&gt;
&lt;h2 id=&quot;smaller-traps-for-therecord&quot;&gt;Smaller traps, for the record&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The page must be served over http, not file://. A file:// origin is opaque, so fetch(“data.json”) is blocked as cross-origin and silently never returns — and there’s no console anybody is reading on a wall display.&lt;/li&gt;
&lt;li&gt;Copying the page to the host changes nothing on screen. The poll fetches data.json, not the HTML, so the browser keeps rendering the markup it loaded at session start. A page change needs the browser restarted. An earlier version of my own README confidently stated the opposite.&lt;/li&gt;
&lt;li&gt;A hostname rename left a stale browser profile lock symlinked to the literal old hostname. Chromium compared it to the current one, decided another machine held the profile, and refused to start — but only on a &lt;em&gt;manual&lt;/em&gt; relaunch, so it read as a broken command rather than a stale lock.&lt;/li&gt;
&lt;li&gt;Screenshotting the console is its own saga. X11, so grim doesn’t apply; xwd -root fails with BadColor on that visual, so you capture by window id; no PIL on the box so you decode locally; and the XWD header’s bytes_per_line is field 12, not 14 — an off-by-two gives “not enough image data” and no hint why.&lt;/li&gt;
&lt;li&gt;Pixel-scanning a screenshot to check alignment works, but bound the scan. One unbounded scan caught the SVG frame’s right edge and reported OPERATIONS as overflowing when it had 60px of clearance. A font size got cut for nothing.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-it-looks-likenow&quot;&gt;What it looks like now&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Two live panels on separate hardware, each with its own collector, web server and separate read-only token — so rotating one can’t silently blank the other.&lt;/li&gt;
&lt;li&gt;One collector, not a fork. Console-only extras are gated behind an environment variable, so the rack Pi costs the cluster exactly what it always did.&lt;/li&gt;
&lt;li&gt;Every field sourced. Two frame words and three unit names are the only text on either screen that isn’t live data.&lt;/li&gt;
&lt;li&gt;Four verdict states, verified against synthetic payloads, previewable with keys 1–4 without inducing a real outage.&lt;/li&gt;
&lt;li&gt;A dark third lobe waiting for a node that doesn’t exist yet.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-honestpart&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;Every time I write one of these up I end on what actually unblocked it, and it’s never what I expect. With the dashboard it was that the tedious middle could be delegated. With the energy alert it was that the two projects before it already existed. This one lands somewhere else.&lt;/p&gt;
&lt;p&gt;The aesthetic constraint did the engineering work.&lt;/p&gt;
&lt;p&gt;I did not set out to build an honest monitoring panel. I set out to build a screen that looks like the MAGI. And because I was copying a fictional interface field by field, I had to ask of every element, “what is this actually saying?” — a question I have never once asked of a Grafana panel, where the fields arrive pre-justified. The tool emitted it, so it must mean something, so onto the screen it goes. Nothing forces the audit. A reference frame that is &lt;em&gt;fiction&lt;/em&gt; forces it on every single field.&lt;/p&gt;
&lt;p&gt;The clock is the sharpest version of it. It’s the most convincing thing on either screen — it moves, it’s precise, it’s the first thing your eye confirms — and it’s the only element guaranteed to keep working no matter how comprehensively everything behind it has failed. It costs one line of JavaScript and it’s worth nothing. A status screen borrows its credibility, by default, from the one component that knows nothing at all.&lt;/p&gt;
&lt;p&gt;The agent’s contribution wasn’t the code, and it certainly wasn’t the idea. It was a willingness to do the part that has always sat between the two: tracing chamfer midpoints off video frames, reading licence fields inside a font file, testing five verdict states against synthetic payloads, checking what the collector actually emits rather than what it ought to. None of that is creative work. All of it is the reason most good ideas stay ideas.&lt;/p&gt;
&lt;p&gt;And it did something I hadn’t expected delegation to do — it made the idea &lt;em&gt;better&lt;/em&gt;. Left to myself I’d have shipped corner-to-corner geometry, a faked countdown and a font clone with no licence: a screen that looked about right and was wrong in three separate ways I’d never have noticed. What exists is more faithful to the thing I’d been picturing for thirty years than the version I would have built, because it went and checked what I was actually picturing.&lt;/p&gt;
&lt;p&gt;The teenager who watched this show in 1995 would be delighted by the screen and wouldn’t have asked a single question about where the numbers came from. The 47-year-old needed it to be true, because there’s a real cluster behind it running things my family uses, and a beautiful screen that quietly stops knowing anything is worse than no screen at all. Both of them got what they wanted, which isn’t how these things usually go.&lt;/p&gt;
&lt;p&gt;So: have the stupid idea. Then make it prove every number on it. The distance between imagining a thing and running it on your wall is the shortest it has ever been — and the idea was never the part you were short of.&lt;/p&gt;
</content:encoded><category>ai-agent</category><category>homelab</category><category>proxmox</category><category>neon-genesis-evangelion</category><category>self-hosting</category><category>Writing</category></item><item><title>Project: MAGI Display</title><link>https://geekconsulting.au/projects/magi-display/</link><guid isPermaLink="true">https://geekconsulting.au/projects/magi-display/</guid><description>Evangelion-style status panels for a Proxmox cluster. Every field on screen is live data.</description><pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate><category>Project</category><category>Homelab</category></item><item><title>Project: mcp-uptime-kuma</title><link>https://geekconsulting.au/projects/mcp-uptime-kuma/</link><guid isPermaLink="true">https://geekconsulting.au/projects/mcp-uptime-kuma/</guid><description>Seven fixes merged into the Uptime Kuma MCP server, so an AI assistant&apos;s monitor changes actually stick.</description><pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate><category>Project</category><category>Open source</category></item><item><title>Project: MetaKavita</title><link>https://geekconsulting.au/projects/metakavita/</link><guid isPermaLink="true">https://geekconsulting.au/projects/metakavita/</guid><description>Security and operations work merged into a metadata enrichment tool for the Kavita reading server: authentication, a hardened container and dependency fixes.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Open source</category></item><item><title>Project: Proxmox VE Helper-Scripts</title><link>https://geekconsulting.au/projects/proxmox-helper-scripts/</link><guid isPermaLink="true">https://geekconsulting.au/projects/proxmox-helper-scripts/</guid><description>Eleven fixes and new scripts merged into the community&apos;s one-command installers for Proxmox, a project with nearly 30,000 GitHub stars.</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Open source</category></item><item><title>Crossfadarr — How Dropping Spotify Turned Into a Two-Day Software Project</title><link>https://geekconsulting.au/writing/crossfadarr-dropping-spotify/</link><guid isPermaLink="true">https://geekconsulting.au/writing/crossfadarr-dropping-spotify/</guid><description>Or: what happens when a homelabber decides his music taste shouldn’t live in someone else’s cloud</description><pubDate>Sun, 19 Jul 2026 03:35:20 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/crossfadarr-dropping-spotify/1.png&quot; alt=&quot;&quot;&gt;&lt;/figure&gt;&lt;p&gt;It started, like a lot of these things do, with a car.&lt;/p&gt;
&lt;p&gt;When Tesla added native YouTube Music support, I did the math and cancelled Spotify. I was already paying for the frankly-not-cheap YouTube family plan, and that plan &lt;em&gt;includes&lt;/em&gt; YouTube Music. Keeping a separate Spotify subscription on top of it was paying twice for the same thing. So: one fewer subscription, no change to how I actually listen. Easy win.&lt;/p&gt;
&lt;p&gt;Except it wasn’t quite that clean.&lt;/p&gt;
&lt;p&gt;YouTube Music’s discovery is genuinely decent — the radios and “because you liked” mixes surface stuff I’d never have found. But everything it learns about me &lt;em&gt;stays&lt;/em&gt; in YouTube Music. My liked songs, the artists I’ve saved, the channels I subscribe to: all of it lives behind Google’s login, useful only inside Google’s app. And I’m a homelabber. My whole hobby, lately, has been the slow, satisfying project of pulling my digital life &lt;em&gt;out&lt;/em&gt; of other people’s clouds and onto hardware I own. Photos, notes, home automation, monitoring — one by one, off the rent-forever platforms and onto the rack in my office.&lt;/p&gt;
&lt;p&gt;My music taste was still a hostage.&lt;/p&gt;
&lt;h2 id=&quot;the-gap&quot;&gt;The gap&lt;/h2&gt;
&lt;p&gt;I already run &lt;a href=&quot;https://lidarr.audio/&quot;&gt;Lidarr&lt;/a&gt; — the music member of the &lt;em&gt;arr&lt;/em&gt; family, the self-hosted apps that manage your media library. Lidarr is great at organising and monitoring artists you tell it about. It is not great at &lt;em&gt;knowing what you like&lt;/em&gt;, because it has no idea what you’ve been thumbs-upping on some streaming service all year.&lt;/p&gt;
&lt;p&gt;So the shape of what I wanted was obvious: take the three signals YouTube Music already has about me — the artists whose music I’ve saved, the channels I subscribe to, and the artists behind my liked songs — and turn them into a clean list I could review and hand to Lidarr. A bridge. Read from one side, review, write to the other.&lt;/p&gt;
&lt;p&gt;I went looking for that bridge. It doesn’t really exist. There’s an old Lidarr pull request that only handles &lt;em&gt;public&lt;/em&gt; playlists. There’s a community plugin that won’t build against the one library that can actually read a private YouTube Music account. And that library — ytmusicapi — is unofficial, because Google offers no official API for reading your own YouTube Music library. That’s the crux of the whole problem: the data is &lt;em&gt;yours&lt;/em&gt;, but the only door to it is a side door.&lt;/p&gt;
&lt;p&gt;Fine. I’d build the bridge myself. I gave myself a weekend.&lt;/p&gt;
&lt;h2 id=&quot;building-it-with-an-agent-not-byhand&quot;&gt;Building it with an agent, not by hand&lt;/h2&gt;
&lt;p&gt;I didn’t write most of this by hand. I built it in a tight loop with an AI coding agent (Anthropic’s Claude, running in my terminal), the same way I’ve built the last few homelab projects — I describe the outcome and the constraints, it writes and runs the code, I steer, we verify against the real thing. Every commit in the repo carries co-author attribution to make that honest and visible.&lt;/p&gt;
&lt;p&gt;The point I want to make here isn’t “AI wrote my app.” It’s that the agent let me spend two evenings on the &lt;em&gt;interesting&lt;/em&gt; problems — the matching, the auth, the judgment calls — instead of two weekends on boilerplate. The project went from empty folder to a public v1.0, dockerised and published, inside about two days. Here’s what actually turned out to be hard.&lt;/p&gt;
&lt;h2 id=&quot;hard-problem-1-the-front-door-is-held-shut-withtape&quot;&gt;Hard problem #1: the front door is held shut with tape&lt;/h2&gt;
&lt;p&gt;ytmusicapi authenticates by borrowing your browser’s session — you copy the request headers out of a logged-in YouTube Music tab and hand them over. It works. It is also fragile in a specific, maddening way: Google rotates the cookies of any session that stays active, so if you grab headers from your everyday browser, that same browser keeps &lt;em&gt;using&lt;/em&gt; the session, rotates the cookie out from under you, and your saved credentials can die within hours.&lt;/p&gt;
&lt;p&gt;The fix is almost silly, and it took real testing to land on: do it in an incognito window. Log in, copy the headers, then &lt;em&gt;close the window without logging out&lt;/em&gt;. A closed private session never rotates its cookies again, so the snapshot you took stays valid for weeks instead of hours. Crossfadarr walks you through exactly that, and validates the paste against a live call to your library &lt;em&gt;before&lt;/em&gt; it stores anything, so you find out immediately if it worked.&lt;/p&gt;
&lt;p&gt;I also went down the “proper” road — I built the full OAuth device-code flow, the durable, grown-up way to authenticate. It works perfectly, right up until you use the token: YouTube Music’s internal API currently rejects valid OAuth tokens outright, HTTP 400 on every endpoint, even though the exact same token is happily accepted by Google’s &lt;em&gt;official&lt;/em&gt; YouTube Data API. It’s a known, Google-side limitation that hits every tool built on this library. So the OAuth code is still in the app, built and working and marked “unavailable,” waiting for Google to flip a switch it may never flip. Sometimes the honest engineering outcome is a feature that sits there labelled &lt;em&gt;doesn’t work yet, not our fault.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;hard-problem-2-is-澤野弘之-the-same-artist-as-hiroyuki-sawano&quot;&gt;Hard problem #2: is “澤野弘之” the same artist as “Hiroyuki Sawano”?&lt;/h2&gt;
&lt;p&gt;To hand an artist to Lidarr, you need their MusicBrainz ID — MusicBrainz being the open, crowd-sourced encyclopedia of music that the entire &lt;em&gt;arr&lt;/em&gt; ecosystem uses as its source of truth. So Crossfadarr searches MusicBrainz for every artist it pulls out of YouTube Music.&lt;/p&gt;
&lt;p&gt;This is where a music library like mine gets spicy. A big chunk of what I listen to is Korean and Japanese, and those artists show up under native-script names, romanised names, stage names, and aliases — often all at once. A naïve string match drops half of them on the floor. Getting this right meant Unicode normalisation and matching a search against an artist’s &lt;em&gt;primary name plus every alias&lt;/em&gt; MusicBrainz knows, so that 澤野弘之 and “Hiroyuki Sawano” resolve to the same person. Each match comes back tagged with a confidence tier — a green/amber/grey dot in the UI — and an alternates dropdown so I can fix the occasional wrong guess by hand instead of trusting the machine blindly.&lt;/p&gt;
&lt;p&gt;There was an unexpected bonus here. I added a flag for artists that MusicBrainz lists no releases for — normally a sign you’d be adding an empty shell to Lidarr. It turned out to &lt;em&gt;also&lt;/em&gt; be a great bad-match detector: a couple of famous artists were quietly matching to obscure bootleg or placeholder entries, and the “no releases” badge is what surfaced them. A feature I built to prevent one problem caught a different one for free.&lt;/p&gt;
&lt;h2 id=&quot;hard-problem-3-making-it-feelfinished&quot;&gt;Hard problem #3: making it feel finished&lt;/h2&gt;
&lt;p&gt;The bones were done in a day. The second day was almost entirely the difference between “a script that works” and “a thing a stranger could run”:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Artwork. Circular artist portraits, because a wall of grey initials is depressing. It fetches from TheAudioDB where it can and falls back to the thumbnails YouTube Music already provides, so the review grid is fully illustrated either way.&lt;/li&gt;
&lt;li&gt;One-click scan. The whole pipeline — ingest, match, artwork, genres — runs inside the app behind a single “⟳ Refresh from YouTube Music” button with a live progress bar, instead of five scripts you run in order. (It’s slow &lt;em&gt;once&lt;/em&gt;, thanks to MusicBrainz’s polite ~1-request-per-second rate limit, then cached and fast forever after.)&lt;/li&gt;
&lt;li&gt;Review, always. Search, filters by confidence/source/type/genre, card and list views — and crucially, nothing is sent to Lidarr until you tick artists and click &lt;em&gt;Add selected.&lt;/em&gt; It dedupes against what’s already in your library and keeps a history of every add.&lt;/li&gt;
&lt;li&gt;Actually shippable. MIT licence, an optional arr-style login for the paranoid, a multi-arch Docker image published to GHCR, and a README that tells the truth about the fragile auth instead of hiding it.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-line-i-deliberately-didntcross&quot;&gt;The line I deliberately didn’t cross&lt;/h2&gt;
&lt;p&gt;Here’s the part I care about most. Crossfadarr downloads nothing. It talks to no torrent indexers, touches no piracy infrastructure, and circumvents no copy protection. It reads metadata from an account you own and writes metadata to software you run. What Lidarr does &lt;em&gt;after&lt;/em&gt; an artist is added is Lidarr’s business, configured by you, in Lidarr.&lt;/p&gt;
&lt;p&gt;That’s not a legal disclaimer I bolted on at the end — it’s a design constraint I held from the first commit, and it’s the line that keeps a project like this squarely in “managing my own library” territory rather than anywhere near the piracy conversation. The bridge only ever carries names, not music.&lt;/p&gt;
&lt;h2 id=&quot;where-itlanded&quot;&gt;Where it landed&lt;/h2&gt;
&lt;p&gt;Crossfadarr is public: &lt;a href=&quot;https://github.com/crossfadarr/crossfadarr&quot;&gt;github.com/crossfadarr/crossfadarr&lt;/a&gt;, MIT-licensed, with a docker compose up and a ready-made image. Point it at your Lidarr, paste your YouTube Music headers (from an incognito window — you’ve read this far, you know why), hit scan, review the grid, add what you want.&lt;/p&gt;
&lt;p&gt;The subscription I cancelled saved me a few dollars a month. The two days I spent making my own listening data portable saved me something I value more: another corner of my digital life that answers to me instead of to a login screen. That’s the whole homelab hobby in miniature, really — not saving money, exactly, but slowly taking the keys back.&lt;/p&gt;
&lt;p&gt;One fewer subscription. One more thing I own.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Crossfadarr is an independent project, not affiliated with or endorsed by Google, YouTube, Lidarr, MusicBrainz, or TheAudioDB — those names are used only to describe what it connects to. It accesses your own account’s data through an unofficial API, which may conflict with YouTube’s Terms of Service; it’s for personal use, provided as-is. It manages metadata only and downloads nothing.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Built in a two-day sprint with substantial AI assistance (Anthropic’s Claude); every commit carries co-author attribution. Source, warts and honest auth story included, at&lt;/em&gt; &lt;a href=&quot;https://github.com/crossfadarr/crossfadarr&quot;&gt;&lt;em&gt;github.com/crossfadarr/crossfadarr&lt;/em&gt;&lt;/a&gt;&lt;em&gt;.&lt;/em&gt;&lt;/p&gt;
</content:encoded><category>homelab</category><category>lidarr</category><category>ai-agent</category><category>self-hosting</category><category>youtube-music</category><category>Writing</category></item><item><title>Project: Crossfadarr</title><link>https://geekconsulting.au/projects/crossfadarr/</link><guid isPermaLink="true">https://geekconsulting.au/projects/crossfadarr/</guid><description>Reads your YouTube Music library and adds those artists to Lidarr. Built in two days after leaving Spotify.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Open source</category></item><item><title>Project: vikunja-mcp-ng</title><link>https://geekconsulting.au/projects/vikunja-mcp/</link><guid isPermaLink="true">https://geekconsulting.au/projects/vikunja-mcp/</guid><description>Three fixes merged into the Vikunja MCP server, so an AI assistant can bulk-edit a task board without losing writes.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Open source</category></item><item><title>Project: feedme</title><link>https://geekconsulting.au/projects/feedme/</link><guid isPermaLink="true">https://geekconsulting.au/projects/feedme/</guid><description>A dinner-deciding app for couples: ask for somewhere new in plain English and get places you haven&apos;t been, ranked by your own dining history.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Apps</category></item><item><title>The Someday Project, Part III: The Homelab Paid For Itself In One Push Notification</title><link>https://geekconsulting.au/writing/someday-project-part-3/</link><guid isPermaLink="true">https://geekconsulting.au/writing/someday-project-part-3/</guid><description>I built the foundation over a weekend. The thing I actually wanted took an hour.</description><pubDate>Fri, 17 Jul 2026 01:11:46 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-3/1.jpg&quot; alt=&quot;&quot;&gt;&lt;/figure&gt;&lt;p&gt;The Model Y has a bad habit. It sits in the garage, unplugged, while the roof throws 4 kW at an already-full Powerwall and the grid price briefly drops under ten cents. Free energy, one door away, and nobody in the house knows — because the car is in the &lt;em&gt;garage&lt;/em&gt;. You don’t walk past it. You don’t see it. There’s no moment in the day where the thought “I should plug that in” is prompted by anything at all.&lt;/p&gt;
&lt;p&gt;The obvious fix is a phone reminder at 11am. The obvious fix is wrong, because the answer to “should I plug in right now?” changes minute to minute and depends on four things that live in four different systems. A fixed daily reminder is either noise or it’s silent at exactly the wrong time. Usually both, in that order, until you swipe it away forever.&lt;/p&gt;
&lt;p&gt;This is the third part of the series. &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;Part I&lt;/a&gt; turned my Home Assistant instance into a proper Grafana data platform. &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-2/&quot;&gt;Part II&lt;/a&gt; built the Homarr control plane I’d abandoned three times. Same collaborator, same house, same principle: I supply direction, the agent supplies execution.&lt;/p&gt;
&lt;p&gt;Here’s the part I didn’t expect. Parts I and II were the point. This one is the &lt;em&gt;payoff&lt;/em&gt; — and it’s the first time the homelab stopped being a hobby and started being useful in a way I can put a dollar figure on.&lt;/p&gt;
&lt;h2 id=&quot;the-question-is-harder-than-itlooks&quot;&gt;The question is harder than it looks&lt;/h2&gt;
&lt;p&gt;Here’s what has to be true, simultaneously, for a plug-in alert to be worth sending:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I’m home. Alerting me about free electricity while I’m at the shops teaches me to ignore the alert. One bad push and the channel is dead.&lt;/li&gt;
&lt;li&gt;The car is unplugged. Obvious, but it needs a live cable sensor, not an assumption about what I did last night.&lt;/li&gt;
&lt;li&gt;There’s genuinely surplus energy. And this is the subtle one — either the solar is producing more than the house is eating &lt;em&gt;and&lt;/em&gt; the Powerwall is already full (otherwise the battery should get those electrons first; charging the car from solar while the battery is at 60% is just moving the problem), or the grid price has dropped low enough that it doesn’t matter where the electrons come from.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Four data sources. Tessie for the car’s cable state and battery percentage. Tessie again for the Powerwall charge and live solar output. Amber Electric for the live wholesale price. Home Assistant for presence. Each with its own auth, its own units, its own refresh cadence.&lt;/p&gt;
&lt;p&gt;Doing this manually means opening four apps and doing mental arithmetic about whether 2.3 kW of surplus into a 96%-full battery justifies a charge session — and doing it &lt;em&gt;repeatedly, all day&lt;/em&gt;, on the off-chance the numbers happen to line up. Nobody does that. That’s the whole reason the car sits unplugged. The task isn’t hard, it’s just impossible to remember to do at the exact moment it matters.&lt;/p&gt;
&lt;h2 id=&quot;the-weekend-that-made-it-a-one-hourjob&quot;&gt;The weekend that made it a one-hour job&lt;/h2&gt;
&lt;p&gt;Here’s the thing I want to be precise about, because it’s the actual argument of this whole series.&lt;/p&gt;
&lt;p&gt;None of the four integrations were built for this.&lt;/p&gt;
&lt;p&gt;Tessie was already in Home Assistant because I wanted the car on a dashboard. Amber Electric went in for the Grafana energy graphs — I wanted to &lt;em&gt;look&lt;/em&gt; at what power costs, that’s all. ntfy — a self-hosted push server behind Caddy, with a Cloudflare tunnel so it reaches my phone off-LAN — exists because Uptime Kuma and Beszel needed somewhere to shout when a container fell over. Home Assistant was there because of course it was.&lt;/p&gt;
&lt;p&gt;Every one of those went in for its own unrelated reason, over a weekend, as part of Parts I and II. And because they did, the plug-in alert cost almost nothing. It was assembly, not construction.&lt;/p&gt;
&lt;p&gt;I’d have told you, before this, that the return on a homelab is the services. It isn’t. The return is that the marginal cost of the next idea keeps falling. Once you have a notification bus, a live energy price feed, and a car API all in the same place, speaking the same language, “tell me when charging is free” stops being a project and becomes a config change. The alert itself is one template expression and one push call.&lt;/p&gt;
&lt;p&gt;That’s what a foundation &lt;em&gt;is&lt;/em&gt;. Not two years of accumulated cruft — a weekend of boring glue, built once, that turns every subsequent someday-idea into an afternoon. I didn’t know that until the third one was trivial.&lt;/p&gt;
&lt;h2 id=&quot;technical-insights-and-the-things-thatbit&quot;&gt;Technical insights, and the things that bit&lt;/h2&gt;
&lt;p&gt;Units will lie to you, silently. The Powerwall’s …_solar_power sensor reports in kW, not W. The first threshold written was &amp;gt; 2000 — a perfectly reasonable number that would never, ever fire. It doesn’t error. It doesn’t warn. It just quietly never triggers, and six weeks later you conclude the sun isn’t shining hard enough in Australia. The agent caught it by evaluating the template against live state before trusting it — the same “ask the running system, don’t assume” discipline that ran through Part II.&lt;/p&gt;
&lt;p&gt;Fail inert, not loud. The cheap-grid branch reads sensor.amber_general_price | float(999). If Amber’s API is down or the entity goes unavailable, that 999 default makes 999 &amp;lt; 0.10 false and the branch quietly disarms — while the solar branch keeps working. A default of 0 would have turned every Amber outage into a 3am push insisting electricity is free. When you pick a fallback value, pick the one that fails towards silence.&lt;/p&gt;
&lt;p&gt;Names are load-bearing. The Amber integration’s site is named “Amber” specifically so its entities land on sensor.amber_general_price. Rename that device in the UI and the automation doesn’t break — it &lt;em&gt;keeps running with the cheap-grid branch permanently disabled&lt;/em&gt;, courtesy of that same float(999). Graceful degradation and silent failure turn out to be the same mechanism wearing a different hat. Worth writing down somewhere you’ll find it.&lt;/p&gt;
&lt;p&gt;Home Assistant is fully drivable headlessly. All of this went in over REST — the automation via /api/config/automation/config/&lt;id&gt;, and the template helper via /api/config/config_entries/flow (handler template, then a menu step, then a form). No clicking through the UI, which means the whole thing is diffable, reproducible, and re-deployable from the repo.&lt;/p&gt;
&lt;h2 id=&quot;revision-1-the-version-that-tested-perfectly&quot;&gt;Revision 1: the version that tested perfectly&lt;/h2&gt;
&lt;p&gt;The first cut put the whole opportunity expression inline in the automation’s conditions, with two triggers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;sustained — a state trigger with for: 00:10:00, so a passing cloud doesn’t fire an alert&lt;/li&gt;
&lt;li&gt;reminder — a time_pattern: /15 poll that re-pushes every 2 hours while the opportunity holds, throttled against the automation’s own last_triggered&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The sustain trigger deliberately &lt;em&gt;bypasses&lt;/em&gt; the throttle. That’s intentional: if I plug in, then unplug an hour later while it’s still sunny, that off→on transition should be able to alert me again immediately, even inside the 2-hour reminder window. A feature, not an oversight.&lt;/p&gt;
&lt;p&gt;It tested clean. Every gate evaluated correctly against live state. A forced fire landed on the phone with the right numbers. The throttle template correctly blocked the reminder path.&lt;/p&gt;
&lt;p&gt;Then it met the real world.&lt;/p&gt;
&lt;h2 id=&quot;revision-2-the-double-push-race&quot;&gt;Revision 2: the double-push race&lt;/h2&gt;
&lt;p&gt;On 15 July, around midday, the conditions genuinely lined up for the first time — the first &lt;em&gt;organic&lt;/em&gt; fire, as opposed to the ones we’d triggered by hand.&lt;/p&gt;
&lt;p&gt;It pushed twice, 27 seconds apart.&lt;/p&gt;
&lt;p&gt;The mechanism, once we went and read the timestamps rather than theorising: the reminder poll fired at 12:00:01, with the opportunity having been true for about nine and a half minutes and last_triggered stale enough to sail straight through the 2-hour throttle. So it pushed. Then 27 seconds later the opportunity crossed its ten-minute mark, the sustained trigger fired — and sustained bypasses the throttle &lt;em&gt;by design&lt;/em&gt; — so it pushed again.&lt;/p&gt;
&lt;p&gt;Two independently correct rules, racing each other. The reminder path had no idea the initial sustain window was still counting down, because it had no way to know.&lt;/p&gt;
&lt;p&gt;The fix was two changes, and only one of them is interesting:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Extract the state. The opportunity expression moved out of the automation and into a template binary sensor helper, binary_sensor.tesla_plug_in_opportunity. “Is this an opportunity?” is now a first-class entity with its own last_changed timestamp, rather than an expression re-evaluated from scratch in two places that can’t see each other. As a bonus, all three thresholds — 2 kW solar, 90% Powerwall, $0.10/kWh — now live in exactly one template instead of being smeared across the conditions.&lt;/li&gt;
&lt;li&gt;Give the reminder path a floor. It now additionally requires the sensor’s last_changed to be more than 600 seconds ago. The reminder now &lt;em&gt;physically cannot&lt;/em&gt; pre-empt the initial sustain, because both paths are finally reading the same clock.&lt;/li&gt;
&lt;/ol&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Lesson: The bug wasn’t the throttle. It was that two triggers were reasoning about the same condition without sharing a source of truth for&lt;/em&gt; when that condition started*. The 600-second floor is the fix you’d write in five seconds; making the opportunity a sensor is the fix that made it possible to write, because there was finally something to ask.*&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That’s a lesson that generalises well past Home Assistant. When two rules disagree, the first question isn’t “which rule is wrong” — it’s “what do they both need to know that neither can see.”&lt;/p&gt;
&lt;h2 id=&quot;where-itstands&quot;&gt;Where it stands&lt;/h2&gt;
&lt;p&gt;Live, verified, waiting on the next organic fire to confirm exactly one push. When it goes off, the message says &lt;em&gt;why&lt;/em&gt;: solar 0.9 kW · Powerwall 96% · car at 80%, priority 4, zap emoji. A notification that just says “plug in the car” is one you’ll eventually stop believing; a notification that shows its working is one you’ll act on.&lt;/p&gt;
&lt;p&gt;The Amber price is also plotted on the Grafana energy dashboard from Part I, with a red threshold line drawn at 10c — the exact number the automation triggers on. If the alert ever behaves oddly, the graph and the trigger tell the same story, because they’re the same number.&lt;/p&gt;
&lt;p&gt;Total new infrastructure required: one template helper.&lt;/p&gt;
&lt;h2 id=&quot;the-honestpart&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;Part II ended with me admitting that what unblocked the dashboard wasn’t a tool — it was that the tedious work between “I want this” and “this exists” could be delegated to something that doesn’t get bored.&lt;/p&gt;
&lt;p&gt;This one’s different, and I think it’s the more interesting admission. The plug-in alert wasn’t unblocked by delegation. It was unblocked by the fact that the last two projects existed. The agent’s contribution here was the weekend that came before, not the hour that produced the thing I actually wanted.&lt;/p&gt;
&lt;p&gt;I’ve been thinking about homelabs wrong for years. I treated each service as the deliverable — Plex is for watching things, Grafana is for looking at graphs, ntfy is for container alerts. But the value compounded somewhere I wasn’t looking: in the &lt;em&gt;adjacency&lt;/em&gt;. Four integrations that had nothing to do with each other, sitting in one system, turned out to already contain the answer to a question I’d never asked them.&lt;/p&gt;
&lt;p&gt;The car gets plugged in now. Not because I’m more disciplined — because the house finally has enough context to notice something I never could, standing in the kitchen, with the garage door shut.&lt;/p&gt;
&lt;p&gt;Build the boring glue once. The third idea is where it pays.&lt;/p&gt;
</content:encoded><category>home-automation</category><category>homelab</category><category>tesla</category><category>home-assistant</category><category>claude-code</category><category>Writing</category></item><item><title>I’m a Salesforce Architect, Not a Rust Developer. I Just Shipped Four PRs to a Rust Project Anyway</title><link>https://geekconsulting.au/writing/salesforce-architect-rust-prs/</link><guid isPermaLink="true">https://geekconsulting.au/writing/salesforce-architect-rust-prs/</guid><description>How domain expertise plus an AI coding agent added a full Salesforce integration to an open-source ETL studio in five days — without me writing a line of Rust.</description><pubDate>Wed, 15 Jul 2026 09:28:37 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/salesforce-architect-rust-prs/1.jpg&quot; alt=&quot;&quot;&gt;&lt;/figure&gt;&lt;p&gt;I’ve spent my career in the Salesforce ecosystem: data models, integration patterns, org governance, the sharp edges of the platform’s APIs. What I have &lt;em&gt;not&lt;/em&gt; spent my career doing is writing Rust, or TypeScript build tooling, or AES-GCM credential encryption.&lt;/p&gt;
&lt;p&gt;And yet, over five days in July, I contributed a complete Salesforce sink connector to &lt;a href=&quot;https://github.com/slothflowlabs/duckle&quot;&gt;Duckle&lt;/a&gt; — an open-source, local-first ETL studio written in Rust on top of DuckDB, with a TypeScript/React frontend. Four pull requests merged, one genuine engine bug discovered and filed along the way.&lt;/p&gt;
&lt;p&gt;I didn’t write the code. Claude (Anthropic’s coding agent, running in Claude Code) wrote the code. My job was everything around the code — and it turns out that’s where most of the leverage is.&lt;/p&gt;
&lt;h2 id=&quot;the-setup&quot;&gt;The setup&lt;/h2&gt;
&lt;p&gt;Duckle is exactly the kind of tool that appeals to a data architect: visual pipelines, native DuckDB speed, no cloud dependency, no lock-in. It had 360+ components. It did not have a Salesforce connector worth the name — and Salesforce is where an enormous amount of enterprise data lives.&lt;/p&gt;
&lt;p&gt;I knew precisely what a good Salesforce sink needed to do, because I’ve lived with the alternatives (Data Loader, MuleSoft, half a dozen middleware products) for years:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Insert, update, upsert, delete via the sObject Collections REST API, with upsert keyed on an external ID field — the pattern every real migration uses&lt;/li&gt;
&lt;li&gt;OAuth 2.0 Client Credentials flow, not username/password or copy-pasted session tokens&lt;/li&gt;
&lt;li&gt;Per-record error handling — Salesforce fails records individually, and a tool that fails the whole batch on one DUPLICATE_VALUE error is useless&lt;/li&gt;
&lt;li&gt;Data Loader-style success/error CSV files, because that’s the artifact every Salesforce data person expects at the end of a load&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What I didn’t know was how Duckle’s Rust plan/execution engine worked, how its frontend palette and field-manifest system generated UI forms, or how its desktop app encrypted saved credentials. That’s a real stack: a RuntimeSpec enum dispatched through a DuckDB-backed engine, a code-generated component catalog, a Tauri-style desktop shell.&lt;/p&gt;
&lt;p&gt;The bet was that the division of labour could be clean: I supply the &lt;em&gt;what&lt;/em&gt; and the &lt;em&gt;why&lt;/em&gt;, the agent supplies the &lt;em&gt;how&lt;/em&gt;, and I verify the result against a real Salesforce org.&lt;/p&gt;
&lt;h2 id=&quot;what-actuallyshipped&quot;&gt;What actually shipped&lt;/h2&gt;
&lt;p&gt;PR &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/165&quot;&gt;#165&lt;/a&gt; — the Tier 1 sink (957 lines across 12 files, merged the next morning). A new snk.salesforce component: engine spec, REST envelope building, per-record result parsing, three mock-server integration tests, frontend palette entry and configuration form, and documentation. My input was a spec: which API, which operations, how batching should chunk (200 records per sObject Collections call), what an upsert on External_ID__c must look like on the wire, and what error text a Salesforce admin would actually understand.&lt;/p&gt;
&lt;p&gt;The design conversation that mattered more than code. After the sink merged, I opened a discussion about authentication (#166): the sink shouldn’t need a hand-minted bearer token; it should mint its own via Client Credentials at run time. I asked the maintainer two design questions — what the config payload should look like, and &lt;em&gt;where&lt;/em&gt; in the architecture token resolution should happen (in the engine, or in the host process that launches runs). He answered both, then implemented that first stage himself. That’s open source working as intended: the contribution was the design pressure and the domain requirements, and the maintainer built it his way, faster than a PR round-trip would have.&lt;/p&gt;
&lt;p&gt;PR &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/171&quot;&gt;#171&lt;/a&gt; — saved encrypted connections (1,434 additions across 39 files, merged within two hours). This was the deep one: a new shared Rust crate (duckle-secrets) extracted from the desktop app’s credential store, AES-GCM encryption for client secrets, run-time resolution of connection references across four different host binaries (desktop, scheduler, runner, headless server), and a frontend connection picker that only ever passes a &lt;em&gt;reference&lt;/em&gt; — so a pipeline file never contains a secret. I could not have written this. But I could specify the security property that mattered: credentials must never land in a pipeline JSON file that someone commits to git, because I have watched that exact incident happen in real Salesforce projects.&lt;/p&gt;
&lt;p&gt;PR &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/176&quot;&gt;#176&lt;/a&gt; — visibleWhen conditional form fields (merged). A pure frontend feature with no Salesforce code in it at all — configuration fields that show or hide based on other fields’ values. It exists because the Salesforce connector needed it: once you support both bearer-token and Client Credentials auth, showing both sets of fields at once is confusing and wrong. The Salesforce use case justified a general capability the whole component catalog can now use. Contributions compound like that.&lt;/p&gt;
&lt;p&gt;Issue &lt;a href=&quot;https://github.com/slothflowlabs/duckle/issues/170&quot;&gt;#170&lt;/a&gt; — a real bug, found by testing like an architect. While building live test pipelines, an edge case I always test — &lt;em&gt;what happens when the source query returns zero records?&lt;/em&gt; — broke Duckle’s REST source entirely: it materialized a single raw json column and every downstream SQL step failed with a binder error. Pre-existing bug, nothing to do with our code. I filed it with a reproduction; the maintainer fixed it within two days. Knowing &lt;em&gt;which edge cases matter&lt;/em&gt; is domain knowledge too.&lt;/p&gt;
&lt;p&gt;PR &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/181&quot;&gt;#181&lt;/a&gt; — resultsPath result files (615 lines, 8 unit tests + 3 integration tests, merged in under two hours). Data Loader-style stamped success/error CSVs — Account_upsert_20260715T031500Z_success.csv — written on every exit path, accumulating across runs rather than overwriting. The accumulate-don’t-overwrite decision was mine, made while testing: a scheduled pipeline that silently overwrites last night’s error file is a tool that loses audit history.&lt;/p&gt;
&lt;p&gt;Alongside the PRs, we built a 17-scenario live test suite that runs the full operation-by-auth-mode matrix against a real Salesforce org — with a credential-free design (${ENV:…} placeholders) so the suite itself can live in the repo without a secret in sight.&lt;/p&gt;
&lt;h2 id=&quot;how-the-collaboration-actuallyworked&quot;&gt;How the collaboration actually worked&lt;/h2&gt;
&lt;p&gt;The honest version, not the demo-reel version:&lt;/p&gt;
&lt;p&gt;I was the product manager, architect, and QA. The agent was the engineering team. Every session started with me describing behaviour in Salesforce terms: “upsert must use the external ID in the URL path, not the body”, “a 200 response can still contain per-record failures”, “field truncation errors need to name the field”. Claude translated that into idiomatic Rust and TypeScript that matched the existing codebase’s conventions — which it read and learned first.&lt;/p&gt;
&lt;p&gt;Testing was my half of the loop, and it was not optional. Every feature got exercised three ways: the agent’s own unit and mock-server tests, the live suite against my test org, and me personally clicking through the desktop app — configuring a connection, running a load, forcing real DUPLICATE_VALUE errors, checking the error CSV said something useful. The agent is good; it is not a substitute for a domain expert watching the actual product do the actual thing. Several of my best inputs (the accumulate-vs-overwrite call, the zero-records bug) came from testing, not from specifying.&lt;/p&gt;
&lt;p&gt;Small PRs, design questions first, maintainer’s conventions always. We asked before building, kept each PR to one concern, matched the repo’s formatting rather than running our own formatter over it (a repo-wide cargo fmt would have touched 11,000 unrelated lines — the agent flagged this before I would have known to care), and flagged known trade-offs in the PR descriptions rather than hoping reviewers wouldn’t notice. All four PRs merged in under 24 hours each — the last in 95 minutes. Maintainers respond to contributors who respect their time; that’s true whether the code came from a human or an agent.&lt;/p&gt;
&lt;p&gt;The friction points were real but manageable. The agent occasionally needed steering back to the maintainer’s stated architecture. The Windows/Rust toolchain setup (MSVC, vendored DuckDB CLI, PATH quirks) consumed a real afternoon. And AI attribution in open source is still an unsettled question — norms vary project to project, and it’s worth having that conversation with a maintainer early rather than assuming.&lt;/p&gt;
&lt;h2 id=&quot;what-this-meansmaybe&quot;&gt;What this means, maybe&lt;/h2&gt;
&lt;p&gt;The conventional wisdom is that open source has a contribution funnel problem: the people with the deepest domain knowledge about what a tool &lt;em&gt;should&lt;/em&gt; do are rarely the people fluent in the tool’s implementation language. A Salesforce architect knows exactly what a Salesforce connector needs; a Rust systems programmer knows exactly how to build it; those are almost never the same person, and historically the connector didn’t get built.&lt;/p&gt;
&lt;p&gt;That wall just got a lot shorter. Not gone — I still needed to understand APIs deeply, make architectural judgment calls, test rigorously, and communicate with a maintainer like a professional. The five days were full days. But the specific skill of &lt;em&gt;writing Rust&lt;/em&gt; stopped being the gate.&lt;/p&gt;
&lt;p&gt;If you’re a domain expert who’s been watching an open-source project miss the feature you need: the excuse inventory is shrinking. Pick the project. Write the spec you already have in your head. Direct the agent. Test like your production org depends on it — because eventually, someone’s will.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The work described: Duckle PRs&lt;/em&gt; &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/165&quot;&gt;&lt;em&gt;#165&lt;/em&gt;&lt;/a&gt;&lt;em&gt;,&lt;/em&gt; &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/171&quot;&gt;&lt;em&gt;#171&lt;/em&gt;&lt;/a&gt;&lt;em&gt;,&lt;/em&gt; &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/176&quot;&gt;&lt;em&gt;#176&lt;/em&gt;&lt;/a&gt;&lt;em&gt;,&lt;/em&gt; &lt;a href=&quot;https://github.com/slothflowlabs/duckle/pull/181&quot;&gt;&lt;em&gt;#181&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, and issue&lt;/em&gt; &lt;a href=&quot;https://github.com/slothflowlabs/duckle/issues/170&quot;&gt;&lt;em&gt;#170&lt;/em&gt;&lt;/a&gt;&lt;em&gt;. Built with Claude Code. Tested against a real Salesforce org with an External Client App on the Client Credentials flow.&lt;/em&gt;&lt;/p&gt;
</content:encoded><category>claude-code</category><category>duckdb</category><category>salesforce</category><category>data</category><category>ai-coding</category><category>Writing</category></item><item><title>Project: Duckle: Salesforce connectors</title><link>https://geekconsulting.au/projects/duckle-salesforce/</link><guid isPermaLink="true">https://geekconsulting.au/projects/duckle-salesforce/</guid><description>Salesforce source and sink connectors for Duckle, a local-first ETL studio written in Rust. Seven PRs merged.</description><pubDate>Sun, 12 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Salesforce</category></item><item><title>The Someday Project, Part II: How an AI Agent Built the Homelab Dashboard I’d Given Up On</title><link>https://geekconsulting.au/writing/someday-project-part-2/</link><guid isPermaLink="true">https://geekconsulting.au/writing/someday-project-part-2/</guid><description>I’d installed Homarr three times. I’d never finished it once.</description><pubDate>Fri, 10 Jul 2026 10:20:50 GMT</pubDate><content:encoded>&lt;p&gt;My media stack worked fine. Two Sonarrs, a Radarr, two Prowlarrs, qBittorrent, Plex, Jellyfin, Seerr (my request manager) — all humming along on three machines. The problem was that “working fine” lived in a dozen browser tabs and a mental map only I could read. Homarr — the self-hosted homelab dashboard — promised to collapse all of that into one pane of glass. I’d tried. Every time I’d wire up two integrations, hit the fiddly bits — an API key here, a reverse-proxy hostname there, a widget that refused to authenticate — lose the thread, and the tab would gather dust until I forgot the admin password.&lt;/p&gt;
&lt;p&gt;The payoff never justified the faffing about. It’s the textbook &lt;em&gt;someday&lt;/em&gt; project: obviously nice to have, never quite worth the afternoon.&lt;/p&gt;
&lt;p&gt;What changed this time isn’t Homarr. It’s that I could hand the faffing about to an agent.&lt;/p&gt;
&lt;p&gt;This is the follow-up to &lt;a href=&quot;https://geekconsulting.au/writing/someday-project-part-1/&quot;&gt;last week’s write-up&lt;/a&gt;, where Claude Code turned my Home Assistant instance into a proper Grafana data platform. Same collaborator, same house, same principle: &lt;strong&gt;I supply direction, the agent supplies execution.&lt;/strong&gt; This time the target was the media stack and the monitoring layer that had been on my “someday” list for a year.&lt;/p&gt;
&lt;h2 id=&quot;three-goals&quot;&gt;Three goals&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;One control plane.&lt;/strong&gt; Every *arr app, both media servers, downloads, requests, and Home Assistant behind a single dashboard I’d actually open — with the ability to start/stop containers from it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real metrics, not green dots.&lt;/strong&gt; Throughput, queue depth, container CPU/RAM, and internet speed flowing into the same VictoriaMetrics + Grafana I stood up yesterday, so the media stack lives next to the smart-home dashboards.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Do it without me doing it.&lt;/strong&gt; My job was to decide &lt;em&gt;what&lt;/em&gt; and approve the shape. Everything tedious — the keys, the widgets, the Docker plumbing — was the agent’s.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;the-stack&quot;&gt;The stack&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Homarr&lt;/strong&gt; for the control plane. I picked it over Homepage, Dashy and friends for one reason: it has the best machine interface of any dashboard I’ve seen — a full OpenAPI + tRPC API, &lt;em&gt;and&lt;/em&gt; its own built-in MCP server. That last part matters more than it sounds. Once Homarr was running, Claude didn’t have to click a UI or hand-edit YAML; it could talk to the dashboard through the dashboard’s own agent interface — list boards, create integrations, add widgets, all as first-class tool calls. The dashboard configured itself, supervised.&lt;/p&gt;
&lt;p&gt;For metrics I reused yesterday’s platform rather than building a new one:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Scraparr&lt;/strong&gt; — one container that scrapes every *arr instance (plus Kavita and Seerr) and exposes Prometheus metrics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;cAdvisor + node-exporter&lt;/strong&gt; — Docker container and host metrics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;vmagent&lt;/strong&gt; → &lt;strong&gt;VictoriaMetrics&lt;/strong&gt; → &lt;strong&gt;Grafana&lt;/strong&gt; — the same pipeline the smart-home data already flows through. Ten-year retention, LAN-only, behind Caddy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Homarr itself sits behind Caddy on my Linux Docker host, reaching Docker through a &lt;strong&gt;socket-proxy&lt;/strong&gt; — a small container that exposes a scoped, read-mostly slice of the Docker socket (start/stop/restart, but no arbitrary exec) so the dashboard can manage containers without being handed root over the whole host. Direction from me: “front the socket, don’t bind-mount it raw.” Execution from the agent.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-2/1.png&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;h2 id=&quot;the-build&quot;&gt;The build&lt;/h2&gt;
&lt;h2 id=&quot;thirteen-integrations-and-a-recurring-lie&quot;&gt;Thirteen integrations, and a recurring lie&lt;/h2&gt;
&lt;p&gt;The first session was integrations: Sonarr (anime + TV), Radarr, both Prowlarrs, qBittorrent, Seerr, then Plex, Jellyfin, Home Assistant, Glances, Uptime Kuma, and Speedtest Tracker. Thirteen in total.&lt;/p&gt;
&lt;p&gt;Almost immediately we hit a quirk worth its own lesson. Homarr’s MCP integration_create tool would return a long, alarming schema-validation error — a wall of red about invalid_union and “expected string, received undefined.” Every single time. And every single time, the integration had &lt;em&gt;actually been created&lt;/em&gt;. The error was in how the tool serialised its success response, not in the action itself.&lt;/p&gt;
&lt;p&gt;A human reads that error and assumes failure. The agent did something better: it ignored the error text and &lt;strong&gt;asked the system what was true&lt;/strong&gt; — called integration_all and looked for the new integration in the list. There it was, connection-tested and saved. From then on, every “failed” create was verified by a read, not a retry.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;A tool’s error envelope is not the source of truth. The running system is. When a write “fails,” read the state back before you believe it.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The same discipline caught a real failure too. Uptime Kuma’s integration genuinely wouldn’t connect — a clean 404, not the cosmetic one. With no credentials it defaults to reading a status page at slug default, which didn’t exist. Rather than guess, the agent read Uptime Kuma’s own SQLite database on disk, found a published status page with the slug main and 24 monitors, and pointed the integration at that. No trial and error — it looked.&lt;/p&gt;
&lt;h2 id=&quot;the-board-nobody-wants-one-big-wall-ofwidgets&quot;&gt;The board nobody wants: one big wall of widgets&lt;/h2&gt;
&lt;p&gt;By the time all thirteen integrations had widgets, the single board was a mess — calendar, indexers, downloads, requests, streams, Docker, weather, all fighting for space. Exactly the state that made me abandon Homarr the last three times.&lt;/p&gt;
&lt;p&gt;So we restructured into five boards: an &lt;strong&gt;overview&lt;/strong&gt; home board, then &lt;strong&gt;media&lt;/strong&gt;, &lt;strong&gt;downloads-indexers&lt;/strong&gt;, &lt;strong&gt;infrastructure&lt;/strong&gt;, and — the fun one — a &lt;strong&gt;home&lt;/strong&gt; board for Home Assistant.&lt;/p&gt;
&lt;p&gt;Here’s where the API depth paid off. Homarr’s MCP can &lt;em&gt;add&lt;/em&gt; items to a board but not &lt;em&gt;remove&lt;/em&gt; them, and it can’t create the collapsible “category” sections at all. The agent went and read Homarr’s source on GitHub, found the saveBoard tRPC route and its exact Zod schema, and reconstructed the whole board — sections, items, per-widget grid layouts — as a single validated payload. It didn’t guess the shape; it fetched the schema and matched it.&lt;/p&gt;
&lt;p&gt;And it did the &lt;em&gt;cautious&lt;/em&gt; thing without being told: instead of destructively rewriting my live board, it built the four new boards fresh and left the old one untouched as a backup for me to delete once I was happy. Judgment I’d expect from a careful colleague, not a script.&lt;/p&gt;
&lt;h2 id=&quot;129-smart-home-tiles-generated-from-the-houseitself&quot;&gt;129 smart-home tiles, generated from the house itself&lt;/h2&gt;
&lt;p&gt;The home board is my favourite artifact of the whole project. Home Assistant knows my house as &lt;strong&gt;areas&lt;/strong&gt; — Bar, Deck, Garage, Kitchen, Living Room, Master Bedroom, Study, and so on. I wanted one Homarr category per area, each filled with toggle tiles for that room’s lights, switches and fans.&lt;/p&gt;
&lt;p&gt;Doing that by hand across 1,261 entities is precisely the work that kills a someday project. The agent did it by asking Home Assistant to describe itself: it hit HA’s REST &lt;strong&gt;template API&lt;/strong&gt; with a Jinja snippet using areas() and area_entities(), got back every area and its controllable entities as JSON, and then generated the board — 14 categories, ~129 tiles — via that same saveBoard route. Lights, switches and fans became clickable toggles; media players became read-only status tiles. Five tiles a row, laid out programmatically.&lt;/p&gt;
&lt;p&gt;Then it read the board back and counted the tiles per category to confirm the whole thing landed. Verification as the last step, again.&lt;/p&gt;
&lt;h2 id=&quot;moving-devices-home-assistant-wouldnt-let-me-move-from-theui&quot;&gt;Moving devices Home Assistant wouldn’t let me move from the UI&lt;/h2&gt;
&lt;p&gt;One area was a near-duplicate — my study existed twice under two slightly different names, a legacy split, with three stray entities stranded on the wrong copy. I asked the agent to merge them. It found the three (a Samsung Odyssey monitor and an air purifier, both from SmartThings), realised HA’s area assignment is only editable over the &lt;strong&gt;WebSocket&lt;/strong&gt; API, not REST — and, finding no WebSocket library installed on the box, wrote a minimal one from the Python standard library: raw socket, the upgrade handshake, frame masking, just enough to authenticate and call config/device_registry/update. Moved the devices, verified the source area was now empty, rebuilt that board section.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;When the tool you need isn’t installed, sometimes the fastest path is to implement the ten lines of protocol you actually use. An agent doesn’t get bored writing a throwaway WebSocket client.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;where-it-got-real-two-migrations-mid-project&quot;&gt;Where it got real: two migrations mid-project&lt;/h2&gt;
&lt;h2 id=&quot;the-abandoned-image&quot;&gt;The abandoned image&lt;/h2&gt;
&lt;p&gt;I’d started a Speedtest Tracker container a while back and asked the agent to just “add it to Homarr.” It checked, and pushed back: the image was henrywhitaker3/speedtest-tracker — the original author’s build, now archived — installed through CasaOS, which I’d stopped using when I moved to ZimaOS. That’s why I couldn’t find the API-token screen: it wasn’t in that old UI.&lt;/p&gt;
&lt;p&gt;I told it to bin the lot. It listed every container first, deleted exactly the three I named — the dead Speedtest image and two orphaned CasaOS containers — and pointedly left the unrelated openspeedtest and other stale containers alone. Then it deployed the maintained image, lscr.io/linuxserver/speedtest-tracker (v1.14), as a clean stack.&lt;/p&gt;
&lt;p&gt;Getting metrics out of it was a chain of small, verified problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The production image ships &lt;strong&gt;no Tinker and no&lt;/strong&gt; &lt;strong&gt;speedtest:run command&lt;/strong&gt;, so there was no obvious way to mint an API token or trigger a test. The agent bootstrapped the Laravel framework in a throwaway PHP script to mint a Sanctum token directly, and dispatched a test through the app’s own action class.&lt;/li&gt;
&lt;li&gt;Homarr’s connection test hits /api/v1/results/latest, which &lt;strong&gt;404s until at least one test has completed&lt;/strong&gt; — so the integration refused to save until a real speed test had run. The agent curled the endpoint, read the 404, understood &lt;em&gt;why&lt;/em&gt;, ran a test, and tried again. It never assumed the token worked; it checked.&lt;/li&gt;
&lt;li&gt;Speedtest Tracker has a built-in Prometheus exporter at /prometheus, but it’s &lt;strong&gt;gated behind a database setting&lt;/strong&gt; and returns “no data available” until a test completes &lt;em&gt;after&lt;/em&gt; the exporter is enabled (it keys off a cache entry). Enable, run one more test, then metrics flow.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-one-character-trap&quot;&gt;The one-character trap&lt;/h2&gt;
&lt;p&gt;Wiring vmagent to scrape the new exporter produced a genuinely instructive bug. I edited the scrape config with sed -i, and the target stubbornly stayed “down” — reading the &lt;em&gt;old&lt;/em&gt; config. The config file is bind-mounted into the container as a single file, and sed -i doesn’t edit in place; it writes a new file and renames it, swapping the inode. The container’s mount still pointed at the &lt;em&gt;original&lt;/em&gt; inode. The fix was to rewrite the file with cat &amp;gt; (which truncates in place, same inode) and restart the agent.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Lesson:&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;Never&lt;/em&gt; &lt;em&gt;sed -i a bind-mounted config file. The rename swaps the inode and the container keeps reading the file you thought you replaced.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After that, one more networking wrinkle — vmagent scrapes its neighbours by container name, and the new container was on its own Docker network — so the agent attached it to the monitoring network and scraped it by name. Then, of course, it queried VictoriaMetrics for the exact metric names to confirm the 22 speedtest_tracker_* series were landing before declaring victory.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-2/2.png&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;h2 id=&quot;a-few-smaller-traps-for-therecord&quot;&gt;A few smaller traps, for the record&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Tautulli&lt;/strong&gt; (Plex history and stats) came up pre-seeded to skip its setup wizard — but the LinuxServer image appends its own default config on first boot, and autodiscovery quietly reset the Plex address to 127.0.0.1. The fix was to pin pms_url_manual and dedupe the config, killing the container rather than stopping it gracefully so it wouldn’t save the broken state on the way out.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Timezones&lt;/strong&gt; in Speedtest Tracker are set by an app-wide DISPLAY_TIMEZONE env var, not a per-user preference — a five-minute rabbit hole that ended in one line of compose.&lt;/li&gt;
&lt;li&gt;The host has &lt;strong&gt;broken IPv6&lt;/strong&gt;, which made Homarr’s weather widget hang on a location lookup. We worked around it with hard-coded coordinates and logged the host issue as &lt;em&gt;deferred&lt;/em&gt; — it’ll likely evaporate in the Proxmox migration, and chasing it now wasn’t worth it.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-it-looks-likenow&quot;&gt;What it looks like now&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;13 integrations&lt;/strong&gt; in Homarr, each connection-tested.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Five boards&lt;/strong&gt;: an overview home page, media, downloads/indexers, infrastructure, and a Home Assistant control board with a category per room and ~129 live tiles.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Container management&lt;/strong&gt; from the dashboard via the socket-proxy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grafana dashboards&lt;/strong&gt; for the *arr apps (Scraparr), Docker/host (cAdvisor), and internet speed (Speedtest Tracker) — feeding the same VictoriaMetrics as the smart-home data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Everything documented&lt;/strong&gt; — a repo, a knowledge-base project, and the agent’s own persistent memory of the gotchas, so next session starts warm.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;if-you-want-to-replicate-it&quot;&gt;If you want to replicate it&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Pick a dashboard with a real API.&lt;/strong&gt; Homarr’s OpenAPI + tRPC + MCP is what made delegation possible; a YAML-config dashboard would have just moved the tedium.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reuse your metrics backend.&lt;/strong&gt; If you already have Prometheus/VictoriaMetrics + Grafana, don’t stand up a second one — point Scraparr and the exporters at what you have.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Front the Docker socket with a proxy.&lt;/strong&gt; Never hand a dashboard the raw socket.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verify every write with a read.&lt;/strong&gt; Integrations, tokens, metrics, board layouts — confirm against the running system, don’t trust the success message.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prefer maintained images, and check who maintains them&lt;/strong&gt; before you build on top.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Default dashboards to LAN-only&lt;/strong&gt; and keep credentials out of your logs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Log the things you’re deliberately &lt;em&gt;not&lt;/em&gt; fixing&lt;/strong&gt; (like the host IPv6 quirk) so “done” is honest.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;the-honestpart&quot;&gt;The honest part&lt;/h2&gt;
&lt;p&gt;I want to be clear about what actually unblocked this, because it isn’t a tool recommendation. I already knew &lt;em&gt;what&lt;/em&gt; I wanted — I’ve known for a year. Homarr didn’t get easier. VictoriaMetrics didn’t get easier. What changed is that the work between “I want this” and “this exists” — the API keys, the schema-matching, the inode trap, the WebSocket client, the six different ways a token can be almost-but-not-quite valid — could be delegated to something that doesn’t find that work tedious and doesn’t lose the thread when the fourth rabbit hole opens up.&lt;/p&gt;
&lt;p&gt;And the thing it did &lt;em&gt;best&lt;/em&gt; wasn’t writing code. It was refusing to guess. Over and over, when a component misbehaved, it didn’t theorise — it asked the running system. It read the 404 instead of assuming the token was fine. It queried the database for the metric name instead of hoping. It read the board back to count the tiles. It checked which image was actually deprecated before deleting anything. That’s the difference between a script and a collaborator: a script assumes; this verified.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-2/3.png&quot; alt=&quot;&quot; loading=&quot;lazy&quot;&gt;
&lt;/figure&gt;
&lt;p&gt;Not direction. Execution — with judgment.&lt;/p&gt;
&lt;p&gt;The dashboard I’d abandoned three times took an afternoon of my attention and a lot of the agent’s. Now every arrow in my media pipeline is on one screen, every container is a click from a restart, and the whole thing trends in Grafana alongside the smart-home dashboards from Part I.&lt;/p&gt;
&lt;p&gt;Build the boring glue once — or rather, get it built. Then go actually use the thing you kept meaning to finish.&lt;/p&gt;
</content:encoded><category>ai-agent</category><category>self-hosted</category><category>devops</category><category>docker</category><category>homelab</category><category>Writing</category></item><item><title>The Someday Project — How an AI Agent Turned My Smart Home Into a Real Data Platform</title><link>https://geekconsulting.au/writing/someday-project-part-1/</link><guid isPermaLink="true">https://geekconsulting.au/writing/someday-project-part-1/</guid><description>A project that sat on my someday-list for years — Home Assistant → VictoriaMetrics → Grafana — done in an afternoon, because the AI didn’t just tell me what to do, it did the work. Here’s the architecture, the config, and every gotcha that tried to stop us.</description><pubDate>Thu, 09 Jul 2026 01:49:06 GMT</pubDate><content:encoded>&lt;figure&gt;&lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-1/1.png&quot; alt=&quot;The end result: a home’s real-time energy flow — solar, battery, house, and grid — collated from Home Assistant into Grafana.&quot;&gt;&lt;/figure&gt;&lt;p&gt;My house was &lt;strong&gt;monitored but not measured&lt;/strong&gt;. Home Assistant had quietly accumulated 868 entities across 70-odd integrations — solar, a battery, temperature and humidity sensors, smart plugs, cameras, printers — and it showed almost none of it in a way I could actually reason about. Home Assistant’s Recorder keeps a few days of detailed history and coarse long-term statistics, which is fine for a glance and useless for “how did the study’s humidity trend across winter?” or “what did the towel warmers actually cost last month?”&lt;/p&gt;
&lt;p&gt;I wanted three things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Keep the data for years&lt;/strong&gt;, not days.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Visualise it properly&lt;/strong&gt; — real time-series charts, computed metrics, thresholds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Not babysit it.&lt;/strong&gt; Set it up once, have new devices show up automatically.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Here’s the honest part: I’d wanted this for years and never built it. Not because I couldn’t — because it’s a dozen fiddly, easy-to-abandon steps, and it lived permanently on my someday list. What finally moved it wasn’t a better tutorial. It was &lt;strong&gt;Claude Code&lt;/strong&gt;, Anthropic’s terminal coding agent, actually &lt;em&gt;doing&lt;/em&gt; the work: it SSH’d into the server, wrote the configs, and — the part that mattered most — &lt;strong&gt;debugged the pipeline live by querying the databases and APIs directly&lt;/strong&gt; rather than guessing. Not direction. Execution. I’ll flag the moments where that changed the outcome.&lt;/p&gt;
&lt;p&gt;This is the playbook. Steal it.&lt;/p&gt;
&lt;h2 id=&quot;the-stack-andwhy&quot;&gt;The stack (and why)&lt;/h2&gt;
&lt;p&gt;The pattern I landed on is the one the Home Assistant community has converged on for long-term storage:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Home Assistant ─ (InfluxDB line protocol) ──▶ VictoriaMetrics  ──▶ Grafana&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Why VictoriaMetrics and not InfluxDB 2?&lt;/strong&gt; VictoriaMetrics is a single, lightweight Go binary that ingests the InfluxDB line protocol (so Home Assistant’s &lt;em&gt;built-in&lt;/em&gt; influxdb integration talks to it with zero add-ons), stores time series with excellent compression, and answers PromQL/MetricsQL so Grafana queries it through the standard Prometheus datasource. InfluxDB 2.x is in maintenance mode and InfluxDB 3 shifted APIs again; VictoriaMetrics sidesteps that churn and sips resources.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Where it runs matters.&lt;/strong&gt; A lot of guides run the database &lt;em&gt;on&lt;/em&gt; the Home Assistant box. If you’re on Home Assistant Green or a Pi, don’t — you’ll grind its little eMMC. I put VictoriaMetrics and Grafana in Docker on a &lt;strong&gt;separate always-on Linux box&lt;/strong&gt; (a spare mini PC) and pointed Home Assistant at it over the LAN. The HA appliance stays lean; the data lives on hardware built to churn it.&lt;/p&gt;
&lt;h2 id=&quot;the-build&quot;&gt;The build&lt;/h2&gt;
&lt;p&gt;Everything is declarative — a Docker Compose file plus provisioning, so the whole stack is reproducible and version-controllable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;docker-compose.yml:&lt;/strong&gt;&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;name: monitoring&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;services:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  victoriametrics:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    image: victoriametrics/victoria-metrics:latest&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    container_name: victoriametrics&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    restart: unless-stopped&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    command:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - &quot;--retentionPeriod=10y&quot;       # keep a decade&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - &quot;--storageDataPath=/victoria-metrics-data&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    ports: [&quot;8428:8428&quot;]              # HA writes here over the LAN&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    volumes: [&quot;vm-data:/victoria-metrics-data&quot;]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;grafana:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    image: grafana/grafana-oss:latest&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    container_name: grafana&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    restart: unless-stopped&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    depends_on: [victoriametrics]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    ports: [&quot;3000:3000&quot;]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    environment:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - GF_SERVER_ROOT_URL=https://grafana.home.example&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - GF_ANALYTICS_REPORTING_ENABLED=false&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    volumes:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - grafana-data:/var/lib/grafana&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - ./provisioning:/etc/grafana/provisioning&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;volumes:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  vm-data:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  grafana-data:&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Provision the datasource as code&lt;/strong&gt; — no clicking around in the UI. provisioning/datasources/vm.yml:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apiVersion: 1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;datasources:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  - name: VictoriaMetrics&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    uid: victoriametrics&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    type: prometheus          # VM speaks the Prometheus query API&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    access: proxy&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    url: http://victoriametrics:8428&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    isDefault: true&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    jsonData:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      httpMethod: POST&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;The Home Assistant side&lt;/strong&gt; is one block in configuration.yaml. Crucially, filter what you send — you do not want all 868 entities, and you want to control how metrics are named:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;influxdb:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  measurement_attr: entity_id     # &amp;lt;-- makes metric names predictable (see gotchas)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  tags_attributes:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    - friendly_name&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  include:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    domains: [sensor, binary_sensor, climate]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    entity_globs:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - sensor.*temperature*&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - sensor.*humidity*&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - sensor.*power*&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - sensor.*energy*&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  exclude:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    entity_globs:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - sensor.*battery*&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - sensor.*signal_level&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I ran Grafana behind a reverse proxy I already had (Caddy), on an &lt;strong&gt;internal-only&lt;/strong&gt; hostname. That last word is deliberate — more on that below.&lt;/p&gt;
&lt;h2 id=&quot;the-gotchas-this-is-the-part-youre-herefor&quot;&gt;The gotchas (this is the part you’re here for)&lt;/h2&gt;
&lt;p&gt;A clean architecture diagram hides a dozen small traps. Here’s every one we hit, and the fix.&lt;/p&gt;
&lt;h2 id=&quot;1-docker-works-but-wontpull&quot;&gt;1. Docker “works” but won’t pull&lt;/h2&gt;
&lt;p&gt;The server had Docker installed, but the active &lt;strong&gt;context&lt;/strong&gt; pointed at a Docker Desktop socket that wasn’t running, and ~/.docker/config.json referenced a credential helper that didn’t exist — so image pulls died with docker-credential-desktop: executable file not found. The fix without disturbing the user’s setup:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export DOCKER_HOST=unix:///var/run/docker.sock   # use the real daemon&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mkdir -p ./.docker &amp;amp;&amp;amp; echo &apos;{}&apos; &amp;gt; ./.docker/config.json&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export DOCKER_CONFIG=$PWD/.docker                 # a clean, helper-free config&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; on a machine with both Docker Desktop and a native daemon, always confirm &lt;em&gt;which&lt;/em&gt; socket you’re talking to before you debug anything else.&lt;/p&gt;
&lt;h2 id=&quot;2-the-influxdb-integration-only-loads-on-a-fullrestart&quot;&gt;2. The influxdb integration only loads on a full restart&lt;/h2&gt;
&lt;p&gt;I added the YAML, hit &lt;strong&gt;Quick reload&lt;/strong&gt;, and… nothing arrived. No data, no errors, the integration simply wasn’t loaded. It turns out the classic influxdb integration has &lt;strong&gt;no reload support&lt;/strong&gt; — a “Quick reload” doesn’t initialise it. Only a full Home Assistant restart does.&lt;/p&gt;
&lt;p&gt;This is where working with an agent that &lt;em&gt;checks&lt;/em&gt; paid off: instead of assuming, it queried GET /api/config and saw influxdb was absent from the loaded components, then triggered a real restart via POST /api/services/homeassistant/restart. Data appeared seconds later.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; after adding a YAML-only integration, do a full restart and &lt;em&gt;verify it loaded&lt;/em&gt; — don’t trust the reload button.&lt;/p&gt;
&lt;h2 id=&quot;3-metric-names-have-dotsquery-them-the-rightway&quot;&gt;3. Metric names have dots — query them the right way&lt;/h2&gt;
&lt;p&gt;With measurement_attr: entity_id, a sensor lands in VictoriaMetrics as a metric named like sensor.study_temperature_value (the numeric state becomes the _value field). Those dots aren’t valid in a bare PromQL identifier, so you select by name:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;{__name__=&quot;sensor.study_temperature_value&quot;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Set measurement_attr and keep it set — if it silently disappears (see gotcha #6), every metric name changes and your dashboards go blank.&lt;/p&gt;
&lt;h2 id=&quot;4-sparse-sensors--instant-queries--phantom-nodata&quot;&gt;4. Sparse sensors + instant queries = phantom “no data”&lt;/h2&gt;
&lt;p&gt;My temperature/humidity sensors report every ~10–15 minutes. VictoriaMetrics’ default instant-query lookback is 5 minutes. So between updates, a plain {__name__=“…”} instant query returns &lt;strong&gt;empty&lt;/strong&gt; — the value is fine, it’s just older than the lookback window. Stat tiles flicker to “No data”; computed panels break entirely.&lt;/p&gt;
&lt;p&gt;The fix is last_over_time to carry the last reading across the gap:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;last_over_time({__name__=&quot;sensor.study_temperature_value&quot;}[30m])&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; match your query windows to how often the &lt;em&gt;device&lt;/em&gt; actually reports, not how often Grafana refreshes.&lt;/p&gt;
&lt;h2 id=&quot;5-computing-dew-point-across-twoseries&quot;&gt;5. Computing dew point across two series&lt;/h2&gt;
&lt;p&gt;I wanted a mould-risk read: dew point per room, from the Magnus formula, which needs &lt;strong&gt;both&lt;/strong&gt; temperature and humidity. Those are two different metrics with different labels, so they won’t join directly — and each is sparse. VictoriaMetrics’ MetricsQL WITH expressions plus a sum(last_over_time(…)) trick solve both problems at once (the sum() strips labels so the two series match; last_over_time bridges the reporting gaps):&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;WITH (&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  t = sum(last_over_time({__name__=&quot;sensor.study_temperature_value&quot;}[30m])),&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  h = sum(last_over_time({__name__=&quot;sensor.study_humidity_value&quot;}[30m])),&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  g = ln(h/100) + 17.625*t/(243.04+t)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;) 243.04*g/(17.625-g)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That one expression turns two raw sensors into a genuinely useful comfort/health metric — no extra template sensors in Home Assistant.&lt;/p&gt;
&lt;h2 id=&quot;6-then-the-platform-changed-underneath-us&quot;&gt;6. Then the platform changed underneath us&lt;/h2&gt;
&lt;p&gt;Two Home-Assistant-2026 curveballs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;“Add-ons” were renamed “Apps.”&lt;/strong&gt; The old /hassio/store URL 404s and the menu item moved. If you can’t find the add-on store, that’s why — look for &lt;strong&gt;Apps&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The&lt;/strong&gt; &lt;strong&gt;influxdb &lt;em&gt;connection&lt;/em&gt; settings are being removed from YAML&lt;/strong&gt; (breaking in 2026.9). Home Assistant auto-imports your host/port/credentials into a UI config entry; you then delete those keys from YAML and keep only the filters, tags, and measurement_attr. We verified the imported config entry was loaded &lt;em&gt;before&lt;/em&gt; removing the keys, restarted, and confirmed data still flowed and metric names were unchanged.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; pin your assumptions to the &lt;em&gt;current&lt;/em&gt; docs, not last year’s blog posts. This is one place an agent that can fetch and read live documentation genuinely earns its seat.&lt;/p&gt;
&lt;h2 id=&quot;7-the-green-line-that-wasntmissing&quot;&gt;7. The green line that wasn’t missing&lt;/h2&gt;
&lt;p&gt;On the energy dashboard, the “solar production” line vanished. It wasn’t missing — production and household consumption were nearly identical (the home battery was soaking up all the solar, so net-to-grid was ~0), and the consumption line was drawn &lt;em&gt;directly on top&lt;/em&gt; of the green one. The fix was cosmetic: explicit per-series colours and lines-only rendering so overlapping series stay legible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; before you debug a “missing” series, check whether it’s simply hidden behind another. Query the raw numbers — they told the whole story instantly.&lt;/p&gt;
&lt;h2 id=&quot;8-a-guardrail-i-was-glad-tohit&quot;&gt;8. A guardrail I was glad to hit&lt;/h2&gt;
&lt;p&gt;When I went to expose Grafana, the agent’s safety layer &lt;strong&gt;blocked&lt;/strong&gt; an attempt to add a public, internet-facing reverse-proxy route I hadn’t actually asked for, flagging it as widening exposure beyond an internal service. It was right. We added an &lt;strong&gt;internal-only&lt;/strong&gt; hostname instead, and only later — once I confirmed the domain resolves purely on the LAN via local DNS with no port-forward — did we add the second route.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; default your dashboards to LAN-only. A Grafana full of your home’s occupancy and energy patterns is not something to casually publish.&lt;/p&gt;
&lt;h2 id=&quot;adding-the-home-battery-cloud-when-local-wont-cooperate&quot;&gt;Adding the home battery: cloud when local won’t cooperate&lt;/h2&gt;
&lt;p&gt;I wanted the Powerwall’s flow — solar, battery, house, grid, and state-of-charge — on the dashboard. The &lt;strong&gt;local&lt;/strong&gt; integration needs a gateway password printed inside the unit, which I didn’t have. Rather than crack open hardware, I used a &lt;strong&gt;cloud&lt;/strong&gt; path: a third-party service (Tessie) with &lt;strong&gt;read-only energy access&lt;/strong&gt;, which exposes the energy site through a proper Home Assistant integration.&lt;/p&gt;
&lt;p&gt;The nice part: because my InfluxDB filter already includes the whole sensor domain, the battery’s power and state-of-charge sensors &lt;strong&gt;flowed into VictoriaMetrics the instant the integration connected&lt;/strong&gt; — zero pipeline changes. That’s the payoff of filtering by domain rather than hand-listing entities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lesson:&lt;/strong&gt; local-first is the right default, but a read-only cloud integration is a perfectly good pragmatic fallback for a stubborn device.&lt;/p&gt;
&lt;h2 id=&quot;the-outcome&quot;&gt;The outcome&lt;/h2&gt;
&lt;p&gt;Two Grafana dashboards, provisioned as code:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Environment&lt;/strong&gt; — temperature, humidity, computed dew-point/mould-risk, and (after connecting SmartThings) PM2.5, PM10, air-quality index, and odor per room.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-1/2.png&quot; alt=&quot;Environment: every room’s comfort and air quality in one view, with dew point computed on the fly.&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;Environment: every room’s comfort and air quality in one view, with dew point computed on the fly.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Energy&lt;/strong&gt; — solar production, per-appliance smart-plug power and daily kWh, and the battery’s whole-home flow with a state-of-charge gauge.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure&gt;
  &lt;img src=&quot;https://geekconsulting.au/writing/someday-project-part-1/3.png&quot; alt=&quot;Energy: production, per-appliance draw, and the battery — years of it, kept and queryable.&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption&gt;Energy: production, per-appliance draw, and the battery — years of it, kept and queryable.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;A decade of retention. A couple of thousand live series. And the property that makes it genuinely low-maintenance: &lt;strong&gt;new devices just appear.&lt;/strong&gt; When I later connected SmartThings, three air purifiers’ worth of air-quality sensors showed up in VictoriaMetrics automatically — I only had to build the panels, not touch the pipeline.&lt;/p&gt;
&lt;h2 id=&quot;replicate-itthe-shortversion&quot;&gt;Replicate it — the short version&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Put &lt;strong&gt;VictoriaMetrics + Grafana&lt;/strong&gt; in Docker on an always-on Linux box (not your HA appliance).&lt;/li&gt;
&lt;li&gt;Provision the &lt;strong&gt;Prometheus-type datasource&lt;/strong&gt; at &lt;a href=&quot;http://victoriametrics:8428&quot;&gt;http://victoriametrics:8428&lt;/a&gt; as code.&lt;/li&gt;
&lt;li&gt;Add the &lt;strong&gt;influxdb&lt;/strong&gt; block to Home Assistant with measurement_attr: entity_id and an include/exclude filter. In 2026, put the &lt;em&gt;connection&lt;/em&gt; in the UI and keep only filters in YAML.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Full-restart&lt;/strong&gt; Home Assistant and verify data lands: curl ‘http://&lt;vm-host&gt;:8428/api/v1/label/__name__/values’.&lt;/li&gt;
&lt;li&gt;Query metrics as {__name__=“…_value”}; wrap sparse sensors in last_over_time(…[30m]).&lt;/li&gt;
&lt;li&gt;Build dashboards; keep them &lt;strong&gt;LAN-only&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;reflection-direction-is-cheap-nowexecution-is-theshift&quot;&gt;Reflection: direction is cheap now — execution is the shift&lt;/h2&gt;
&lt;p&gt;Let me be honest about why this article exists at all. I didn’t lack the &lt;em&gt;knowledge&lt;/em&gt; to build this — I lacked the appetite to grind through a dozen fiddly steps and their inevitable dead-ends. Every blog post could tell me &lt;em&gt;what&lt;/em&gt; to do. None of them did it. So for years, it didn’t get done.&lt;/p&gt;
&lt;p&gt;What changed isn’t that the instructions got better. It’s that the agent &lt;strong&gt;did the work&lt;/strong&gt; — SSH’d into the box, wrote the Compose file, edited the YAML, and, crucially, &lt;strong&gt;debugged against the live system&lt;/strong&gt; every time something broke. The silent Quick-reload, the empty dew-point panel, the “missing” green line, the 2026 platform rename — each fix came from the agent &lt;em&gt;reading the actual state of things&lt;/em&gt;: querying the time-series database by exact metric name, hitting Home Assistant’s REST and WebSocket APIs, checking which components had loaded, fetching current docs when the platform had moved. It didn’t theorise about why data wasn’t flowing; it asked the database.&lt;/p&gt;
&lt;p&gt;That distinction — between an assistant that &lt;em&gt;advises&lt;/em&gt; and one that &lt;em&gt;acts&lt;/em&gt; — is the whole story. Advice I’ve had for a decade. Having the tedious middle actually executed, verified, and handed back working is the genuinely new thing, and it’s why a someday-list project became a done one in an afternoon. If you want a phrase for it: this is what the age of AI actually feels like — not a chatbot with suggestions, but a collaborator that does the work.&lt;/p&gt;
&lt;p&gt;And there’s a transferable lesson buried in it, whether or not you ever use an agent: &lt;strong&gt;verify against the running system at every step.&lt;/strong&gt; Does the integration show as loaded? Does the metric exist &lt;em&gt;by its exact name&lt;/em&gt;? Is the series empty, or just outside the lookback window? Is the line missing, or behind another one? Nearly every gotcha above was invisible in the config file and obvious from one query.&lt;/p&gt;
&lt;p&gt;Build the pipeline. Then go ask it questions for years.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Setup: Home Assistant 2026.7 · VictoriaMetrics · Grafana OSS 13 · Docker · Caddy. All identifiers in this article are placeholders — keep your real hostnames, IPs, and tokens private, and keep dashboards off the public internet.&lt;/em&gt;&lt;/p&gt;
</content:encoded><category>ai-agent</category><category>home-automation</category><category>home-assistant</category><category>grafana</category><category>self-hosting</category><category>Writing</category></item><item><title>Project: The Someday Project</title><link>https://geekconsulting.au/projects/someday-project/</link><guid isPermaLink="true">https://geekconsulting.au/projects/someday-project/</guid><description>A smart home turned into a data platform, a dashboard and automations that pay for themselves.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><category>Project</category><category>Homelab</category></item><item><title>Project: Salesforce Sandbox Data Seeder</title><link>https://geekconsulting.au/projects/salesforce-data-seeder/</link><guid isPermaLink="true">https://geekconsulting.au/projects/salesforce-data-seeder/</guid><description>Fills a Salesforce sandbox with realistic, relationship-aware data that satisfies the org&apos;s own validation rules.</description><pubDate>Fri, 22 Aug 2025 00:00:00 GMT</pubDate><category>Project</category><category>Salesforce</category></item><item><title>Project: G3DK Printing</title><link>https://geekconsulting.au/projects/g3dk-printing/</link><guid isPermaLink="true">https://geekconsulting.au/projects/g3dk-printing/</guid><description>Website and customer order portal for a one-person 3D printing service in Sydney.</description><pubDate>Thu, 20 Mar 2025 00:00:00 GMT</pubDate><category>Project</category><category>Business</category></item></channel></rss>