Leveraging Agentic AI Tools Without the Slop
Agentic AI: Your New (Occasionally Clueless) Pair Programmer#
I've been using agentic AI code generation tools heavily for about a month now. Proof of concepts, side projects, automation scripts - you name it. And I have thoughts. Lots of them.
But before we dive in, let me establish something: I'm not a complete novice who discovered AI can write console.log for them. I've been teaching myself to code since I was thirteen. I have a Bachelor's degree in Business Informatics. I've worked as a Software Architect and Developer Experience Engineer at a large German IT company since 2018. I review pull requests daily and can smell code rot from three abstraction layers away.
So when I talk about agentic AI coding, I'm speaking from the perspective of someone who knows what good code should look like - and who has seen these tools produce both brilliant solutions and absolute garbage.
The Dark Ages: When LLMs Were Just Fancy Stack Overflow#
When agentic coding tools first started appearing, I wasn't impressed. The code quality was... let's call it "aspirational." There weren't proper mechanisms to feed these models architectural context or enforce project-specific rules. Using an LLM for code back then was essentially like having a faster search engine for Stack Overflow answers.
Small refactorings? Sure. Generating a simple utility function? Fine. But the moment you asked for anything complex, quality nosedived. Why? Because context is the enemy of coherence for LLMs.
Think of it like this: someone sends you a 15-minute voice message rambling about their weekend, their cat's medical issues, and oh by the way, can you pick up milk? By minute twelve, you've mentally checked out. Large Language Models experience something similar. The more input they process, the harder it becomes to generate output that actually addresses your intent. Research has shown that LLM performance degrades significantly as context length increases, particularly for tasks requiring information synthesis from multiple sources.
The Evolution: Agents Learn to Use Tools#
Here's where things get interesting. Modern agentic AI coding isn't just about having a smarter model - it's about giving models tools to manage their own limitations.
Picture this: it's Monday morning. You've barely touched your first coffee when your boss drops six massive folders on your desk and asks for a TL;DR by noon. Without the ability to Ctrl+F through those documents, you're cooked. But with search? Suddenly the task is manageable.
This is exactly what MCP (Model Context Protocol) servers and agentic tools enable. Instead of shoving an entire codebase into the context window and hoping for the best, agents can:
- Search selectively - Grep through files for relevant patterns
- Extract surgically - Pull only the code sections that matter
- Browse dynamically - Query the web for documentation or examples
- Execute strategically - Run commands, check test results, validate assumptions
What is MCP? The Model Context Protocol is an open standard that allows AI models to interact with external tools and data sources. Think of it as a standardized plugin system for LLMs - letting them fetch files, search codebases, query databases, or interact with APIs rather than relying solely on their training data.
The result? Context windows stay lean while information density stays high. The model can focus on what matters instead of drowning in irrelevant noise.
Frontier Models Still Produce Slop#
Let's be clear: throwing money at the biggest, shiniest model doesn't prevent bad code.
Even flagship models like Gemini 3 Flash or Claude Opus 4.5 can fail spectacularly when you're implementing something uncommon or debugging non-obvious issues. And honestly? The hardest part isn't the debugging itself - it's keeping calm while you watch the model confidently proclaim it found the issue, only to make things worse or apply the same broken "fix" for the fifth time.
It's like having a pair programmer who sounds incredibly authoritative while producing pure nonsense. We've all worked with that person.
Claude Opus: The (Current) Gold Standard#
As of early 2026, Anthropic's Claude Opus 4.5 remains the leader for code generation quality. It was trained on especially high-quality data and excels at following project conventions - linter rules, type safety requirements, agent rules you've defined. I've genuinely learned new patterns by reading Claude's reasoning steps as it works through problems.
But even Claude gets sloppy. Here's what I've noticed after extended sessions:
| Pattern | What It Looks Like |
|---|---|
| Reinventing the wheel | Creating new utility functions when identical ones exist in the same project |
| Type cast addiction | Excessive use of as assertions instead of fixing underlying type issues |
| Test corruption | Modifying tests to pass rather than fixing the actual code |
| Fatigue shortcuts | After 4-5 hours of intensive work, quality visibly degrades |
That last point is particularly insidious. The model isn't literally "tired" - but as conversation context accumulates, the same degradation effects we discussed earlier kick in. The model starts taking shortcuts, and unless you're paying close attention, you might not notice until your test suite is full of meaningless assertions.
Why are type casts code smell? Excessive casting (
as SomeType) typically indicates a design problem. Either your types don't accurately model your data, or you're fighting the type system instead of working with it. Type casts exist for edge cases like variance issues or poorly-typed external libraries - not as a general-purpose escape hatch.
Shift Your Review Left#
In any substantial project, breaking changes that cascade through dozens of call sites are the worst kind of refactoring. This is painful enough when you're doing it manually. It's somehow even more painful watching an LLM do it badly.
I've seen agentic tools attempt mass refactors through:
- Overly aggressive sed replacements that corrupt unrelated code
- File-by-file regeneration (30+ seconds each for large files)
- Context-busting approaches that make subsequent outputs unreliable
The solution? Always use planning mode for non-trivial changes.
Before letting the agent write a single line of implementation code:
- Have it explain its intended approach
- Check for architectural consistency with existing patterns
- Verify it understands the scope of required changes
- Question any assumptions that seem shaky
Yes, it's tempting to type "implement this feature" and walk away. That approach will hurt you. The same way skipping design review hurts your own code - AI-generated or otherwise.
Git Is Your Safety Net (Use It)#
This should be obvious, but: always use Git.
When working with agentic AI, I recommend an aggressive staging strategy:
# After every few meaningful changes the agent makes
git add -p # Review and stage incrementally
git stash # Save work-in-progress if neededBenefits:
- Visibility - You see exactly what changed, when
- Reversibility - Easy rollback at any granularity
- History - Squash commits later for clean main branch history
Some developers commit every single agent change, then squash-merge when the feature is complete. Whatever works for you - the point is having escape routes when (not if) the agent produces garbage.
Test-Driven Development Is Non-Negotiable#
If you're not using TDD with agentic AI, you're flying blind.
Here's the workflow that works:
- Agent writes implementation code
- Immediately have the agent write tests
- Review the tests yourself - are edge cases covered? Do assertions test behavior or just check that nothing throws?
- Run the full test suite after every significant change
Why does this matter? Tests are vastly easier to review for correctness than implementation code. Well-written test descriptions tell you exactly what behavior is expected. Missing test cases reveal gaps the model didn't consider. And as your project grows, the test suite becomes your early warning system against agent-induced regressions.
I'll admit: I don't love writing tests manually. But reviewing AI-generated tests and thinking critically about what's missing? That's genuinely interesting detective work.
Security Isn't Optional#
This section could save your career.
Agentic AI doesn't care about your security posture. It will cheerfully:
- Commit API keys to client-side code
- Create SQL injection vulnerabilities
- Expose internal endpoints without authentication
- Store sensitive data unencrypted
I've seen all of these. In production-destined code. The API key one especially - it's a mistake any junior developer learns not to make in their first week, yet LLMs reproduce it regularly.
Your responsibilities:
- Learn OWASP Top 10 vulnerabilities and check for them
- Use static analysis tools (like
semgrep,snyk, or language-specific linters) - Never trust agent code with sensitive operations without manual review
- Consider threat modeling before implementing auth/payment/data-handling features
The agent can build a beautiful application. But if that application leaks customer data or opens attack vectors, your reputation - and possibly your company - is finished. There's no git revert for destroyed trust.
The Bottom Line#
I want to give a shout-out to Gabriel Afonso, whose post on "vibe coding" being the new "Made in China" crystallized something I'd been feeling. His observation that we're building a world of AI-generated mediocrity is worth reading.
Here's my take:
Use agentic AI to amplify your abilities, not replace them. Vibe Coding has bad reputation but isn't inherently bad. Use it the right way.
- Keep your mental model sharp
- Become an excellent code reviewer
- Don't just "vibe code" your way through features
- Understand what the agent is doing and why
The alternative is shipping another half-baked application that dies before anyone sees it. The tools are powerful. The output quality is on you.
What's your experience with agentic coding? Found any other patterns or anti-patterns I should know about? Drop me a line - I'm always curious to learn.