· Jesse Edwards · Case Studies · 6 min read
A Year of Building With AI: What Worked and What Didn't
Thirteen hours a day, eight of them with AI. Why AI is not like a bug you can fix, what building with it really looks like, and how much to trust it depending on what is at stake.

I average about 13 hours a day, seven days a week, and probably 8 of those are spent working with AI: building with it, using it to work through ideas, and using it for marketing and outreach. I use it everywhere.
I am used to programs that do not lie. When they are wrong, it is a bug: same input, same wrong answer. You find it, you fix it, and it stays fixed. Even the hard ones follow that rule. An interrupt landing at just the wrong moment could give you a different result once in a while, but recreate the same timing and you got the same failure. It had a cause you could hunt down and fix. The program never “believes” anything. It does exactly what the code says.
AI models do not work that way. Give it the same input twice and it might be right once and confidently make something up the next. It is fluent and sure of itself either way. And there is no line of code to fix. The model is trained to produce likely text, and training pushes it toward being accurate, but nothing inside it checks a claim against a source unless you build that around it.
The models keep getting more accurate and answering more like a human. I would say overall they are better, but they are still far from perfect. If you need answers that are 100% accurate, you need to build a guardrail around it.
After a year of using it that much, and running it in production at RenovationRoute, I have a pretty clear picture of where it belongs and where it does not. This is the first of a few posts on that. This one is about building with AI. The next ones cover what went wrong in production and how we catch it, where AI belongs and where it does not, and why I still pay for AI instead of running it myself.
Building with AI
I wrote embedded C for years, C# for a short bit, then a little bit of most things along the way. Settled mostly with Python and Rails. Now I build mostly with Claude. I do not see myself going back to embedded C, I rarely write Python by hand anymore, and I still drop into VS Code sometimes. Overall it makes me faster, but it is not a total win.
Some days it does the opposite, and I want to throw my computer. I have spent 8 hours on a task that should have taken 2, because every fix felt like progress while it kept breaking simple things. Switching to a bigger model or the latest model did not help. It just kept messing up. The lesson I keep relearning: when it starts going in circles, take a step back. Write a design document if there is not one already. Break the work into small tasks, have it work on one at a time, and check each one. If it still keeps getting it wrong, shrink the scope again and build back up, or build the part it keeps messing up yourself. I would say the learning curve keeps becoming less. Writing C code was a learning process. Python was less. At the time of writing this, building with AI is still not tell it what to do and it does it. Maybe it will get there, but just like knowing python to build bigger systems with AI you have to know the software process. It can help you do these, but without it, good luck. That is architecture, designs, coding standards, testing. Software best practices essentially.
Models make mistakes at every level:
- Architecture mistakes. It might suggest an overall structure that seems reasonable but has hidden flaws or inefficiencies. If not reviewed carefully, these mistakes can propagate throughout the system. Overall would rate it very well, but just like normal software development this stage hits hard when wrong.
- Design mistakes. It fills in what you did not specify and does not mark where it guessed. If not reviewed carefully, you might end up with a program that is designed incorrectly or has hidden assumptions. Again, getting this stage wrong hurts.
- Coding mistakes. It is rarely code that looks wrong. The problem is volume. If not kept in check it will generate 10k lines of code that should have been 1k. More code likely more bugs. In a 10k line change where 9,998 lines are right, the two that are not do not stand out. Leftover dead code or a quietly dropped function you catch in the diff. The small ones ship. This is where testing and other reviews should come in.
- Testing mistakes. When it writes both the code and the test, the test checks what the code does, not what it should do. It cannot think outside the implementation it just wrote.
It is confident about every one of them. Years of software development and writing software without it is what lets me make sure it follows a process and be able to know when it is wrong.
How much I review depends on what is at stake. The money transfers in RenovationRoute? You can bet I wrote most of that code myself and had the AI review it, not the other way around. A UI change? I will let it do the work and barely check it. Most features land somewhere in the middle of those two. Either way, you need enough base knowledge to know what is right and wrong, or you are just shipping its mistakes faster, and they compound quickly.
The short version
AI is a component, not a coworker you trust blindly. It drafts and it extracts. Code checks and decides. A person signs off on anything that matters. And you need enough of a base to know when it is wrong, because it will not tell you.
One more thing
Even this post. I wrote it paragraph by paragraph, misspellings everywhere, just getting my thoughts down, and AI cleaned it up into what you are reading. Writing it properly myself in college would have taken me an hour or two. Now that I rarely write, it would be closer to 3 hours.



