A Year of Building with AI
BMAD, Claude Code skills, Diagrammo, a nightly software factory and the token bill — what building with AI agents looked like for me in 2026.
At the start of the year my Generative AI page said I was “doing agentic development with BMAD Method and Claude Code for a desktop app” at home. That one line turned into most of my year. This is a look back at how it went, mostly from the git history, because the commits remember better than I do.
BMAD: structure for agents
I started Diagrammo in February with the BMAD Method, an agile-ish framework where AI agents play the roles: analyst, product manager, architect, developer. You write a PRD, an architecture doc, epics and stories, and the developer agent works the stories one at a time.
It was the right training wheels. Coming from managing engineering teams, it felt familiar: clear requirements, small stories, a definition of done. The agents did much better work when they had a story with acceptance criteria than when I just described what I wanted. The BMAD output folder took 518 commits this year and peaked at about 170 a month over the summer.
Then it tapered off. In August I moved the backlog to GitHub Issues, and by September the BMAD folder had gone quiet. I didn’t make a big decision to drop it. The ceremony that helped when I didn’t trust the agents started costing more than it saved once I did. The habits stuck: an issue still has to say what done looks like.
Skills: teaching the agent how I work
The second thread was Claude Code itself. At first my setup was a settings file and a status line. By the end of the summer it was a real project in my dotfiles, about 107 commits, with a CLAUDE.md of rules and a handful of skills: instructions the agent loads only when the task calls for them.
Mine are about how I want to work more than what to build. One makes the agent stop and give me options before it starts editing when I describe a problem rather than ask for a fix. One covers how it uses git, so I never have to think about branches. One makes it show a mockup before writing paragraphs about a UI. Most of the value came from corrections: every time it did something I didn’t like, the fix became a rule so it wouldn’t happen again.
The surprise came later, when I measured. Across 882 sessions, skills had been launched 11 times, and 59 of the 64 available had never fired at all. What I’d written mattered, but mostly through the always-loaded CLAUDE.md, not the clever skill library.
Diagrammo: the thing I built
Diagrammo is diagrams as text: a small language, dgmo, that renders 50 chart types. The diagrams in this post are written in it. It started as a desktop app, and over the year it grew a CLI, an online editor, an MCP server so AI agents can draw, Diagrammo Cloud for sharing, and plugins for Obsidian, Astro, Docusaurus, VitePress and others. The npm package went from 0.0.1 in February to 0.90.0 yesterday.
About 96% of those 1,653 commits carry a Claude co-author line. I don’t write much of the code by hand any more. My job looks a lot more like my day job: decide what matters, write it down clearly, review the work and say no to the wrong things.
The software factory
The newest thread is what I’ve been calling the software factory. A home server named anchor runs an agent I call bob on a schedule. Every night at 9pm bob picks up open issues and works them for five hours. Later passes check the work, write a retro and triage new issues, and at 4am I get a Telegram message saying what closed, what’s unfinished and what needs me.
In October I added a dashboard, the factory itself, that records each run: what closed, what each run cost in tokens and how releases went. It keeps me honest. The last dgmo release took five attempts under four titles, and without the record I’d have remembered it as one.
Watching the bill
All of this burns tokens, so in September I sat down and measured instead of guessing. A few things I learned:
- Before I typed a word, a session already carried about 47k tokens of instructions and tool definitions. Diagrammo’s own instruction file alone was 38k, and I cut it by about 80%.
- 96.5% of input tokens were cache reads, which are cheap. Output was where the money went, so short replies matter more than short context.
- Tool results were almost 90% of a big session’s context, and reading whole files was three quarters of that. Now the agent reads the part of the file it needs.
- Nearly everything ran on the biggest model, greps included. Searches now go to a smaller, cheaper one.
What I’d tell someone starting
Start with structure; BMAD or something like it teaches you to write work down so an agent can do it. Then turn every correction into a rule, and keep the rules short. Measure before you optimize. And expect your job to change: I spend less time typing code and more time doing what I do at work, which is deciding what good looks like and holding the work to it.