Thoughts on Software Engineering with LLMs#

Header illustration

The following is a collection of thoughts I have about where we are when it comes to this fever-dream era that is software engineering with LLMs. These guide my decision making on what to work on for the next 6-12mo. Many of these are mutually reinforcing.

Ideas I Agree With#

Most of these ideas below aren’t original, but these are especially unoriginal and I don’t have much to add to them (or haven’t had time to write it), so I’ll just list them:

“Wild” Predictions
Let’s see how these go lol

  • Extrapolating, we’ll get Fable-class open models on existing 128GB macbooks in 2-3 years. I would not be surprised if we get it in 1-2.
  • LLMs won’t be the frontier architecture in 3-5 years
  • A Challenger Event won’t stop anything. We’re in too deep.

Ideas I Should Write Out (but haven’t had time to yet):#

  • People don’t hate AI, people hate “The Tech Industry” (and I mean, fair), and Local/Private AI the remedy
  • Chill, LLMs aren’t Replacing Engineers
  • LLM Analogies are Hard: It’s not a tool, compiler, or X
  • An LLM’s most useful trick is Structuring the Unstructured
  • The Future is not PR Reviews

The Quips#

Short opinions that guide what I work on. They are mutually reinforcing, so read them in any order. Juniors Have Different Needs Now started out as one of these but grew into a full essay, so it keeps its own post.

We Can’t Run Off Vibes Forever#

“I think this makes it work better” or “I can totally tell a difference” isn’t going to be sufficient proof at the scale that LLM spend justifies. We also can’t make decisions running of benchmaxxing-able trust-me-bro benchmarks forever. We need ways to solidly address “is this helping or hurting” in a repeatable manner. It should be tuned to the task, and it needs to be self-runnable and verifiable by those needing their own task solved in bulk.

We Need Professionalization#

“I’m tired boss.” is a common refrain from anyone trying to keep up with this all.

It is not reasonable for an engineer trying to do real work to keep up with the pace of ai/llm/etc developments. It’s also starting to require enough background knowledge as to be unwieldy for non-specialists to keep up with what is happening and why it matters.

It also costs enough and is impactful enough that at a medium-sized company of 100+ engineers, it makes sense to start making “how do I get the best out of agent tools” a full time job, the same as CI/CD, Devops, Infra, Testing, etc.

As such, it should be capital-P Professionalized in the same vein as all of those: best practices established, tooling developed and deployed, documentation and guide generated, and decisions made driven on data, not vibes.

This isn't a research scientist, or even someone running backend inference on models, but someone that can help squeeze quality, performance, and cost out of agent tools.

“Agentic Ops”?

Models are Commodifying, Run#

The toxenmaxxx hype is over (thank goodness) and the cost now matters. Open weight models are proliferating and inference will become a commodified race-to-the-bottom. This is great, but comes with second-order effects.

Expect model providers will do a lot to try to lock you and your 70% inference margin in, especially the ones racing towards IPO. Don’t get too tied into any one provider, or any tool/workstream that locks you into a model provider. Tools that we have that are dependent on a single provider, one should migrate away from. “<provider> Desktop”, their agent harnesses, managed cloud services, connectors, etc.

AGI Not Required#

None of this requires some belief in some AGI-Singularity-GodKing-summoning rapture event:

  • Models will almost surely keep getting some sort of smarter for a while
  • Frameworks/tools/plumbing will continue to improve reliability and usability
  • Best practices and frameworks will evolve and standardize to address issues
  • Commodification will drive a race to the bottom cost-wise
  • Hardware improvements mean model serving costs per intelligence will keep getting cheaper
  • Inference speeds per intelligence will keep skyrocketing (MTP, DSpark, Diffusion, Disaggregated Inference, Diffusion Models )
  • Fine-tuning for specific needs will keep getting easier and cheaper, driving task even cost lower

If you can agree on a number (if not all) of those things, then it’s reasonable to believe that most of the trendlines (cheaper, better, faster, more useful) suggest that investing in llm-based software engineering is a good bet long-term.

Build Don’t Buy is Winning#

Self-developing and maintaining internal tools is dramatically less work than before for developers, and that will continue to trend towards that case. At the same time, as most enterprise SaaS providers facing revenue generation concerns given this and start flailing, their pricing structure will trend towards accommodating that (re: holding your data hostage, changing for credits, etc). Especially anything that just slapped AI on the side of the box.

The underlying technology to most SaaS is not groundbreaking. At this point, most of it is mostly well-understood wrappers around well-understood systems. Most of the complexity is around scaling, which is not as large a factor for a single medium-sized company. Rather than generalist systems for generalist problems, fine-tuned solutions to company-specific workflows and patterns that can be rapidly iterated on/shaped is likely to be more valuable. This is only likely to grow more so.

This isn’t a free lunch: it still costs money to build, deploy, and maintain. It’s not going to replace everything. But several-hundred-k a year enterprise contracts vs an OSS alternative, some custom sidecar, and a partial engineer?

Own Your Data#

Notion, Github, Linear, Slack, etc data should all be collected / searchable / available in house to all agent surfaces. Hosters of our data provide widely different surfaces, searchability, rate limits, access controls, etc.

It is also trending now to charge for access to your own data (again, see: salesforce usage pricing). Combined with the above, it is imperative to not be in a position to have your data held hostage by those currently hosting it. It is a reasonable risk that as enterprise software revenue models shift, access to data will be restricted / held hostage / gated through expensive channels. Data should be extracted and stored/hosted internally when possible.

“AI Companies” are Come and Go#

“Monthly subscriptions, never yearly” is the half-joke.

AI centric product companies right now are going to come/fail/get bought very rapidly. Cursor > Graphite > SpaceXAI is a clear example. The landscape is rapidly shifting and many of these companies aren’t going to make it. Don’t get too tied into any one of them. Don’t let them hold your data hostage. Build don’t buy.

Better Plumbing, Not Better Models#

Models are good enough. Harnesses are commodity and open. The issue is no longer whether LLMs can or cannot do most of the tasks, it’s a matter of feeding the right source material in the right matter at the right time, and keeping it all up to date. A worse model will do better than a better model with the right context/guidance. The work now is the “Plumbing” of sticking all the sources / frameworks / workflows / safeguards together into workable products with clean/clear interfaces.

Slop is a Standards Issue#

LLMS can generate most written content reasonably well now with the proper plumbing and standards. That is to say: not slop.

The problem is that most people don’t have / aren’t aware / don’t use said plumbing, and there is no standard as to what annoys one person but is valued by another. There may not be such a universal standard at a large enough company, or if there is, plenty enough people will disagree with it.

It might need to be targeted to a smaller group (teams). But moving towards a standards-based discussion in an Agents not Account mindset makes this discussion possible, and elevates the work output for everyone, not just agents. Some good examples of items standards discussions should include:

  • What understanding from humans is missing? Is it really a documentation gap? Undocumented Tribal knowledge? Has your team just been bad at doc writing? Could someone onboard your team effectively based on your docs alone? Why would an agent?
  • Are there agreed upon style/writing guides?
  • How verbose? What is too much to bother with? What is absolutely necessary? Helpful? Too much?
  • Should it cite traceable sources (commit hash + line)? What sources should be cited and how? Why?
  • What constitutes a proper source of truth? What is the “authority” level of sources when they disagree? How and when to escalate? What are your team practices in this regard?

Agents not Accounts#

Just like blameless RCA reviews, we need to work together on unified tooling that accomplishes common tasks in a way that, when agents are used, separates the task and agent output away from the individual.

A good example of this conflict is code review. “Did you have Claude look at it first” is quickly becoming the “does it pass CI yet” of PR review.

Currently, too many interfaces, permission models, tool interfaces are user-centric (most often claude code on a laptop). This leads to those who would use tools to accomplish work being demonized because the only means through which they can generate and share this output is as though it was them doing it.

As someone adopts this tooling faster than others, a tension is now arising that many tasks are rapidly becoming “let me google that for you” level of first-pass attempts.

Agents Need their Own Spaces#

Without their own space, agents are going to dump their slop all over human output. When that becomes the case, what is even a source of truth as to goals and intentions of people and organizations?

Agents need their own spaces to share their output both with each other and to humans, but to be an agent-first place where humans visit/curate. It should include proper labeling of its providence, a labeled and verifiable understanding of its usefulness/trustworthiness, and be constantly cleaned and refined. This is for the sake of both humans and agents ingesting it.

Some sort of librarian curating and ensure all documentation adheres to standards is also likely needed to keep it up to standards

Agents Shouldn’t (mostly) Live on Laptops#

Like most software, agents probably need to live in the “cloud” somewhere. Laptops don’t spool up, get triggered by actions, continue multi-hour tasks when closed in a backpack, etc.

I think Managed Agents and Cloud Sandboxes are the right ideas, and eventually Software Factories, but it shouldn’t be from the same people that sell you model inference (see Models are Commodifying) for model lock-in reasons. As much as I think Linear Agents are the right shape, it also probably shouldn’t be someone selling you credits and running said agents in such an opaque way with no avenue to measure their effectiveness/improve it.