All projects
Client & Production Projects

CodeMate Agent A local AI coding agent for VS Code that turns free-text tasks into reviewed, safely-applied code changes across .NET and React codebases.

An AI coding agent engineered with the same reliability and safety discipline as production software — policy-guarded code generation, automatic post-apply verification and rollback, circuit breakers, and a CI pipeline that includes mutation testing.

Illustration of CodeMate Agent turning a chat task into a reviewed code diff, verified by policy, build and typecheck checks

Client / Project

-

Industry

Developer tooling / AI engineering

My Contribution

  • Backend agent architecture (Python/FastAPI)
  • VS Code extension (TypeScript)
  • Safety systems: policy engine
  • verification & rollback
  • circuit breakers
  • AI orchestration: intent classification
  • self-review
  • memory
  • CI/CD
  • release process & mutation testing
  • Architecture decision records (ADRs)
CodeMate Agent chat panel in VS Code with context scope and intent controls and a task ready to send

The Brief

CodeMate Agent is a local AI coding agent for VS Code, built as a client engagement to cut down time spent on repetitive engineering work across .NET and React codebases — designed to act as a second developer that can be handed a task, rather than a generic chatbot bolted onto an editor.

Built with: Python, FastAPI, TypeScript, VS Code Extension API, Qdrant, GitLab CI

Problem

General-purpose AI coding assistants are good at producing plausible code and much weaker at being trusted with it: no real guardrails against unsafe changes, no verification that generated code actually builds, no automatic recovery if something goes wrong, and no memory of what failed last time. The goal here wasn't a smarter autocomplete — it was an agent someone could reasonably let touch a real codebase.

Solution

Built end-to-end as a Python/FastAPI backend paired with a VS Code extension with a Copilot-style chat panel. Free-text requests are classified by intent (code, tests, build-error diagnosis, chat, explain) with a confidence threshold — low-confidence requests are surfaced back to the user for confirmation instead of guessed. Every generated change runs through a bounded self-review loop, then a runtime policy engine before it's allowed to touch disk (blocking forbidden paths, dangerous operations, and hardcoded secrets), then a post-apply build/typecheck verification step with automatic rollback if it fails. Long or interrupted tasks checkpoint per stage so a run can resume instead of restarting from scratch.

Key Features

Every change is shown as a diff preview first and applied only after approval.

CodeMate Agent diff preview in VS Code showing generated email validation next to the original component, awaiting approval

Technical Approach

CodeMate Agent request flow from the VS Code extension through intent classification, generation, self-review and policy checks to diff preview, apply and verification with rollback

  • Backend: Python + FastAPI, with an LLM-provider abstraction so the underlying model can be swapped without touching orchestration, memory, or the code-generation pipeline.
  • Orchestration: intent classification with confidence gating, task planning, a bounded self-review loop, and an autonomous re-planning path that feeds prior failure patterns (stored in Qdrant across sessions) back into the next attempt.
  • Safety: a runtime policy engine blocking forbidden paths, dangerous operations, and hardcoded secrets before any change is applied; post-apply build/typecheck verification with automatic rollback on failure; per-stage checkpointing so interrupted runs can resume.
  • Reliability: shared circuit breakers across LLM and memory calls, an admin API exposing health/session/metrics data, and structured logging correlated by request ID.
  • Extension: a VS Code chat panel with a context-scope toggle (active file / open files / workspace), run history, and per-stage failure diagnostics.
  • Process: architecture decisions recorded as ADRs, a mutation-testing gate in CI (kill-rate threshold, not just line coverage), diff-coverage checks, and a documented release process behind multiple shipped extension versions.

Challenges

Trusting an agent with real edits

the hard part wasn't generating plausible code — it was building enough guardrails (policy checks, verification, rollback, human approval) that letting an LLM touch a real file stopped feeling reckless.

Failing usefully instead of guessing

confidence-gated intent classification means the agent asks for clarification on an ambiguous request instead of silently doing the wrong thing.

Making the safety net actually hold

circuit breakers and mutation testing only matter if they're enforced, not just present — wiring a kill-rate mutation gate and shared circuit breakers into CI meant proving the safety net held up under real failures, not just existing in principle.

Extending the agent without breaking trust in it

adding new intents, re-planning, and checkpointing while the agent was already handling real tasks meant every change had to preserve existing safety guarantees rather than just adding capability on top.

My Role

Worked across the backend agent, the VS Code extension, and the CI/CD and release process — architecture, implementation, and ongoing iteration across multiple shipped versions. Also wrote the ADRs behind key decisions and drove the project's testing and reliability practices, from unit and mutation testing to the verification/rollback and circuit-breaker systems described above.

Conclusion

The result behaves less like a code-generation demo and more like production infrastructure: every generated change passes through intent classification, self-review, policy checks, and post-apply verification before it's trusted, with automatic rollback and resumable checkpoints if something breaks along the way. Reliability and engineering discipline — not just prompting — turned out to be most of the actual work.

Real products, practical engineering, and problems worth solving.

Selected work across .NET, full-stack SaaS, AI integrations, developer tooling, React, React Native, and more.

Get In Touch