Dyego Maas - Blog

Generative AI Consultant and Software Architect

Getting the Most Out of Vibe Coding

Getting the Most Out of Vibe Coding

Why is Vibe Coding amazing, yet not enough to build professional-grade software?

9 min read

Vibe Coding is a new way to bring our ideas to life and finally ship that project that’s been sitting in a drawer for months. Free from the rigor of traditional software engineering, the practice brings accessibility, inviting and empowering people all over the world to CREATE. But as we’ll see below, there are limits to how far this approach can take you, and there are also ways to get the most out of the experience.

Vibe Coding

The practice went viral and gained popularity quickly after January 2025, when Andrej Karparthy, former head of AI at Tesla, published this post:

Super Accessibility

Two years ago it would have been unthinkable for someone with no technical expertise to build an application from scratch, with no help at all, and now that’s entirely possible.

Agentic tools like bolt.new, Lovable, and Gemini opened the door to a new kind of rapid prototyping, unlike anything that came before. You don’t need to write a single line of code by hand!

Lovable's home page, with a simple prompt box suggesting you build a complete application
The Lovable.dev interface
UI prototype for a financial dashboard built in Lovable with a single prompt
Project built in Lovable with a single prompt

This means new audiences can now code applications:

  • Kids, who used to have to learn basic programming languages or use some low-code tool
  • Professionals from other fields, like education, healthcare, and retail, who used to depend on developers to build their digital solutions

For tech professionals, tools like these, along with agentic IDEs like Cursor and Windsurf, bring a lot of autonomy and new superpowers:

  • Product Owners, designers, coordinators, managers: they all now have far more autonomy and no longer depend on the development team
  • Backend developers can now build the frontend on their own with the help of an AI agent
  • Likewise, frontend developers can now build backends on their own with the help of an AI agent

Nota

This video was presented by Tom Wolf, from HuggingFace, and shown in Andrej Karpathy’s talk “Software Is Changing (Again)”.

Flight game made by Pieter Levels with Vibe Coding
Flight game made by Pieter Levels entirely with Vibe Coding

Fun and the Luck Factor

Vibe Coding can be a ton of fun, because working freely with an AI agent involves an element of luck: when it works, it feels like winning a prize.

AI tools are making programming more fun for me. It’s like having an unpredictable genie. I wake up in the middle of the night thinking: ‘Oh, that’s how I should have instructed my agent!’ […] I’m having more fun programming now than ever.

— Kent Beck, on The Pragmatic Engineer podcast (paraphrased)

But notice Kent Keck’s wording: an unpredictable genie. Just like in the stories, your wish is granted, but not always the way you pictured it, and often with unintended consequences.

Losing at the Game of Luck

The biggest problem in Vibe Coding sessions is that, at some point, things will start going wrong.

The non-deterministic nature of agents built on LLMs (Large Language Models) is the source of this variation in results and in the tools’ reliability, but it isn’t the only factor that introduces luck into the mix.

Other important sources of variation are:

  • The prompt you provide
  • The context you provide (or don’t)
  • The chosen model (LLM)
  • The project structure

Let’s explore how each of these introduces uncertainty into the process.

Vague or Incomplete Prompts

An overly simplistic prompt has two immediate effects: it gives the agent a lot of freedom, and at the same time it leaves the agent in the dark.

Contrary to what a shallow analysis might suggest, too much freedom gets in the way of good results instead of boosting them. To illustrate, here are a few examples:

  • It’s the deadline constraint that gets many projects delivered: without a deadline, projects slip indefinitely.
  • Small teams perform better than very large ones: oversized teams add a lot of communication noise and administrative bureaucracy. That’s the reason behind Jeff Bezos’s famous golden rule for squad size: the “Two-pizza Team”.
  • Every game is built on rules: it’s the constraints those rules impose that make the game possible.

In other words, it’s detailed, well-written prompts, free of ambiguity, with clear and well-thought-out requirements, that let the agent do what we ask, the way we expect, with effective results.

So high-quality prompts are essential for good results. Some traits of these prompts are:

  • Clear definition of functional and non-functional requirements
  • Clear goals
  • Unambiguous language
  • Rules to follow, like team, project, or company guidelines, making it clear what the agent can and can’t do

Context Provided (or Omitted)

Just like a poor prompt, poor context lets the agent work against our expectations and wishes.

Even when we’re building a project from scratch, it’s important to give the agent context. That context can be provided in several ways:

  • Relevant details in the prompt about the company’s reality and about the end user
  • Images attached to the prompt provide valuable visual references. bolt.new, for example, lets you attach Figma designs so it faithfully follows existing designs.
  • Links to library and framework documentation help the agent implement things more correctly
  • Project and company documentation, guidelines, examples: all of this helps get the agent on the same page as us and align expectations.

Another way to add extra context, or to let the agent gather more context on demand as needed, is to give the agent tools through MCP servers. The MCP (Model Context Protocol) has seen explosive adoption since March 2025, when Anthropic launched it.

Agentic coding tools like Cursor, Windsurf, GitHub Copilot, and Claude support this feature, and agents can take advantage of it to enrich their context dynamically.

Some powerful MCP servers for enriching context are:

  • Context7, to fetch up-to-date implementation examples for any library or framework
  • Perplexity, for deeper searches that require research to support decision-making
  • Pupeteer or Playwright, which let the agent test the application visually by opening the browser, taking screenshots, reading the console, etc.

In short, well-defined context can be the difference between a bad, low-quality, or even invalid result and a high-quality one.

Picking the Wrong Model

Another factor that can heavily influence the results and the vibe coding experience is the choice of model (LLM).

While high-level Vibe Coding tools like bolt.new and Lovable pick the models for you, other tools like Cursor and Windsurf give the user the power to choose.

Agent prompt box in the Cursor IDE, with the grok-4 model selected
Agent model selector in Cursor

You can always turn on “Auto” mode and leave the choice to the tool. But in many cases the selected model may be a simpler, cheaper, or older one. These models can be efficient and well suited for simple tasks, but not for the task at hand.

Each model has its own capabilities, weaknesses, and strengths, which can make all the difference in the outcome of an implementation. Here are a few examples:

  • Older models, like OpenAI’s GPT-4, were trained on competitive-programming-style tasks, but aren’t very good at real, everyday development work.
  • Anthropic’s Claude 3.7 Sonnet was trained specifically for real-world programming tasks, which is why it’s one of the most popular models for coding, but it isn’t great at following instructions to the letter: it often takes liberties and frequently forgets or ignores requirements.
  • Claude 4 Sonnet and GPT 4.1 were trained to follow prompt instructions more closely.
  • Reasoning models, like Claude 4 Sonnet Thinking and Google’s Gemini 2.5 Pro, can think before answering or taking action, producing better results.
  • Older models like Grok 3 don’t handle tool use, such as MCP servers, very well.
  • Claude 4 Opus is a very powerful reasoning model for planning, but it’s also extremely expensive to use. It’s very useful for solving complex problems, but financially unviable for most day-to-day tasks.

At the same time, LLMs develop a kind of personality during training, which can influence the results and how agents react to problems.

Claude 3.7 and Claude 4, for example, tend to panic and start making mistakes.

Claude reporting that it deleted everything, even what it shouldn't have.
Claude deleted everything, even what it shouldn’t have
Claude delivering bad news
Claude delivering bad news

Models can also show a sense of humor, as you can see in the next image.

Claude answering in metaphors.
Claude delivering bad news

In short, knowing the available models lets us pick the best model for each task.

The Big Three

These three elements need to be well aligned for us to get even minimal success in any agent-driven development.

All three elements need to be nailed down to get good results: Prompt, Context, and Model (LLM).
Recipe for good results

From Vibe Coding to Augmented Coding

Exciting and fun as it is, Vibe Coding is nowhere near a recipe for success. Especially in long-running projects, which is the case for most real projects.

It’s only a matter of time before things start going wrong.

True story of a project gone wrong

Below is another example that’s a natural consequence of Vibe Coding, something we could call runaway bloat.

To avoid this kind of trap, we need to turn to more robust software engineering practices. That’s where Augmented Coding comes in.

Augmented Coding

Augmented Coding goes far beyond Vibe Coding, applying well-established software engineering practices to ensure good results.

Some examples of software engineering practices that can be applied:

  • Architecture planning
  • Development techniques like AI-assisted TDD (_Test-Driven Development)
  • Periodic refactoring to raise quality, removing code instead of always adding more, which is the tendency of today’s LLMs
  • Applying design patterns
  • Automated testing: static analysis (linters, security), unit tests, integration tests, end-to-end, load tests
  • Good CI/CD pipelines (Continuous Integration/Continuous Delivery)
  • Good code review processes
  • Technical documentation
  • Others

It’s about using AI agents to augment our capabilities and produce more and better, giving the developer superpowers, but with responsibility, never losing sight of how important it is to maintain software quality.