All projects
Personal & Open Projects

Lumen Lumen A multi-tenant RAG and MCP server on .NET, built stage by stage with every design decision researched and recorded.

An in-progress .NET service that ingests documents, indexes them with vector search, and exposes retrieval as MCP tools an AI agent can call directly — with tenant isolation proven by tests, not just assumed.

an illustrated cover: tenant documents → vector database → MCP tool → AI agent

Industry

AI infrastructure / developer tooling — RAG & MCP

My Contribution

  • Architecture (Onion Architecture
  • .NET)
  • Multi-tenancy design & verification
  • Database & vector-search technology decisions
  • CQRS backend (EF Core
  • MediatR
  • PostgreSQL)
  • Testing strategy (unit + integration
  • tenant-isolation proofs)
  • Project direction & decision log
Lumen's Onion Architecture layers, from the API down to the domain, with unit and integration test projects

The Brief

Lumen is a personal project: a multi-tenant Retrieval-Augmented Generation (RAG) and Model Context Protocol (MCP) server built on .NET — ingesting documents, indexing them with vector search, and exposing retrieval as tools an AI agent (like Claude Code) can call directly. It's still in active development, built stage by stage, with its own running decision log instead of decisions made and forgotten.

Built with: C#, ASP.NET Core, .NET, EF Core, PostgreSQL, pgvector, Azure OpenAI, Microsoft.Extensions.AI, Semantic Kernel, MediatR, AutoMapper, JWT, xUnit, Testcontainers, Docker

Problem

The goal was a project that goes deeper than wiring an LLM API to a chat interface — something that tests real backend engineering under a constraint that's easy to fake and hard to actually get right: multi-tenancy that holds up under scrutiny, not just code that compiles. It also needed a realistically messy document domain to index against, rather than clean sample text, and to end somewhere concretely usable — an MCP server an AI agent can query directly, not just a demo script.

Solution

Built stage by stage rather than all at once, with each stage getting its own research, design, and implementation pass before the next one starts. Stage 1 built the foundation: an Onion Architecture .NET solution (Domain → Application → Infrastructure → Api), self-issued JWT authentication, and multi-tenancy enforced through an EF Core global query filter with exactly one deliberate, documented exception — the login lookup, which has to run before a user's tenant is even known. Every non-trivial decision is written down in a running log with the reasoning and, wherever possible, a measured number behind it, rather than just "we chose X." Development itself runs through a spec-driven, agent-assisted workflow: each stage gets a written design and a task-by-task plan, an AI agent implements each task, a separate review pass checks it against the spec and code quality, and nothing moves forward until it's verified with a real build and test run — the architecture and the decisions behind it stay under direct control throughout, not delegated along with the code.

Stages 2 and 3 built on that foundation: a document pipeline that turns real SEC filings (HTML and PDF) into typed elements — headings, paragraphs, tables — and chunks them by token budget, then embeddings stored in PostgreSQL with pgvector and a tenant-scoped similarity search over them.

Technical Approach

  • Architecture: Onion Architecture across four layers (Domain, Application, Infrastructure, Api) plus two test projects, with dependencies pointing strictly inward; the Domain layer has zero external dependencies.
  • Multi-tenancy: a TenantId discriminator enforced as a real foreign key, an EF Core global query filter applied to every tenant-owned entity, and an unconditional server-side stamp on every write, so a caller can never set their own tenant. Exactly one legal exception exists in the entire codebase — the pre-login user lookup — and it's documented and checked rather than left implicit.
  • Auth: self-issued JWT (no external identity provider), with claims driving a dedicated tenant-resolution middleware.
  • Data & CQRS: PostgreSQL via EF Core, MediatR for CQRS command/query handling, a Repository + UnitOfWork pair over one shared scoped database context.
  • Vector search: PostgreSQL with the pgvector extension, chosen over SQL Server 2025's native vector type after a direct comparison — pgvector's indexing is stable and generally available, while SQL Server's approximate-search indexing is still preview-gated, and its container image has a documented crash under Docker on Apple Silicon. POST /api/search embeds the query (Azure OpenAI, or a deterministic offline fake for local development) and runs a cosine-distance query over stored chunk embeddings — tenant-scoped by hand in raw SQL, because a vector query can't compose with EF Core's global filter.
  • Document ingestion: real SEC filings rather than clean sample text. HTML and PDF are read into typed DocumentElements (heading / paragraph / table); prose is chunked at 250 tokens with a 60-token overlap (cl100k tokenizer), while every table stays one indivisible chunk, split by rows only if it alone exceeds the budget. On Apple's FY2025 10-K (1.4 MB of HTML) that produced 296 chunks — 209 prose, 87 tables — with p95 at 236 tokens, none over budget, and the whole upload pipeline running in 0.75 s.
  • Testing: 137 tests across unit and integration suites, including two independent tenant-isolation tests — one at the database-query level, one a full HTTP round-trip with real issued tokens — plus integration tests running against a real PostgreSQL container rather than a mock.

Tenant-isolated request flow in Lumen, from login to a document query filtered to the caller's tenant

Challenges

Proving isolation instead of assuming it

multi-tenancy through a query filter is easy to get subtly wrong; it only counts as solved once it's backed by tests at two different levels and a codebase-wide check that exactly one documented exception exists.

Choosing infrastructure on its actual operational merits

the vector-database choice came down to real, reproducible constraints (a container crash on one platform, a still-preview feature on another) rather than a general "which is better" debate — decided and written down before any code depended on it.

Targeting a genuinely messy document domain

real filings meant real failure modes — a fixed-size splitter cut financial tables mid-row, detaching numbers from their labels. Treating tables as their own indivisible chunks fixed that for the 29% of chunks that are tables, and the decision log records when to revisit it (if later retrieval evaluation shows table questions failing).

Keeping architecture decisions human-owned in an agent-assisted workflow

using AI agents to implement and review individual tasks only works if the actual design reasoning — why this isolation boundary, why this database, why this one exception and no others — stays explicit rather than getting delegated along with the code.

My Role

Every design decision — the tenant-isolation strategy, the database choice, the layering — is researched and recorded with reasoning before implementation starts. Development runs through a spec-driven workflow where AI agents handle task-level implementation and review under this direction, with every stage verified by a real build and test run rather than taken on faith.

Conclusion

Lumen is an active, multi-stage project rather than a finished product — what's built so far — the foundation, document ingestion, and embeddings with tenant-scoped vector search — is the first half of a longer roadmap that still includes the RAG endpoint with eval-driven tuning, the MCP server itself, and deployment. What's meant to stand out isn't how much is done, but how it's being done: every stage starts with real research, every non-trivial decision is written down with its reasoning and a measured number behind it, and nothing ships without independent verification — including a tenant-isolation guarantee backed by two different kinds of tests, not just a design that looks right on paper.

Real products, practical engineering, and problems worth solving.

Selected work across .NET, full-stack SaaS, AI integrations, developer tooling, React, React Native, and more.

Get In Touch