Building a Personal RAG System from Scratch, Part 1: Why I'm Building My Own
I'm building a personal RAG system around selected Obsidian notes using C#, Ollama, SQL Server 2025, MCP, and Open WebUI. Part 1 explains the problem, the planned architecture, and what I want to learn.
I already keep homelab documentation in Obsidian. The next thing I want to build is a way to ask questions about those notes and get an answer that points back to the information it used.
That might be a question about a Home Assistant automation, a network decision, or a service running in the homelab. A general-purpose chat model can explain those technologies, but an explanation of how something usually works is different from an answer about my setup.
In Documenting My Homelab, Part 1, I wrote about keeping structured Obsidian notes and using a script to compile them into a master document for AI-assisted planning and troubleshooting. This project takes that idea further: retrieve the relevant sections when I ask a question, rather than rely on a compiled document as the source of context.
I’m going to build a personal Retrieval-Augmented Generation system, or RAG system, around a selected collection of those notes. I’ll write the ingestion and retrieval logic in C#, generate embeddings locally with Ollama, and store them in SQL Server 2025. Open WebUI will provide the chat interface.
This first post is the plan. I’m still at the architecture and research stage, so there are no retrieval benchmarks or end-to-end results to report yet. The point of the series is to work through those pieces and show what I learn along the way.
What I Want to Ask My Notes
The starting collection will be homelab documentation. Potential topics include Home Assistant, UniFi networking, virtual machines, Docker services, NAS configuration, Plex, and DNS.
The kinds of questions I want to support are:
- Which notes explain how a particular service is configured?
- Why did I create a separate ingress network for my reverse proxy services?
- Where is the configuration example for an automation?
- What troubleshooting steps did I record for a problem?
Those are examples of intended use, not questions the system has answered already. I’ll start with a small collection that I can inspect and questions whose answers are actually in the notes.
An ingress-network note is one example of the material I want to search. It records the purpose of that network alongside configuration details. A question about why the network exists needs the explanation, not just a subnet or VLAN number.
An example of the homelab documentation I want my RAG system to search.
I also want the original source to remain easy to reach. A summary is useful, but I should be able to open the document and check the configuration, the heading, and any qualifications around it. For technical documentation, the line after a code block can matter just as much as the code.
What RAG Means in This Project
RAG stands for Retrieval-Augmented Generation. The basic idea is to give a language model relevant source material as context when answering a question. Instead of expecting it to know how I configured my network or why I made a particular decision, I can retrieve that information from my own documentation.
For this project, there will be two main stages:
- Prepare the documentation: read selected Markdown notes, split them into smaller chunks, and store those chunks with their source information and embeddings.
- Answer a question: search for relevant chunks, then give the matching text and source references to a language model so it can generate an answer.
An embedding is a numerical representation of text that helps compare related content. The plan is to use the same embedding model for document chunks and questions, so the system can look for relevant notes even when my question doesn’t use their exact wording. The details of chunking and generating those embeddings will get their own articles.
The embedding model and the language model have two different jobs. One helps find the information; the other turns the retrieved context into a readable answer. I’m not training a model on my vault. The notes remain separate from the answer model, and the index will need to follow changes to them.
A convincing answer won’t prove that the right information was retrieved. I want to see the supporting document and section, and I want the system to acknowledge when the notes don’t contain an answer. That will be part of judging whether this is useful for my own documentation.
Why Build This Myself?
Existing applications already provide document ingestion, search, and chat. If the only goal were to get a knowledge-base interface running, a ready-made application would be a reasonable place to start.
I want to see what happens when Markdown becomes chunks, when chunks become vectors, and when a query returns something that is related but doesn’t actually answer the question. Building those pieces gives me a way to follow the information through the system and investigate where an answer went wrong.
I’ll use Pluralsight courses and other online resources throughout the project, alongside the official documentation for the tools I’m using. I’ve always found that working on a real project while learning helps reinforce new concepts. Writing the code, testing it, and troubleshooting problems along the way makes what I’ve learned stick. This project gives me somewhere to apply each topic as I go.
It’s the same practical learning approach behind my UniFi Doorbell MQTT project: build something useful while working through concepts that are easier to understand with real code and a real problem. I’ll write the ingestion and retrieval services while reusing a database, model hosting, and a chat interface. That keeps the work focused on the concepts I want to learn.
The Planned Architecture
These are my initial choices. I’ll revisit them if the implementation gives me a reason to change something.
| Component | Planned role |
|---|---|
| Obsidian and Markdown | Keep the original documentation and source material |
| C# and ASP.NET Core | Read, chunk, index, and retrieve notes through reusable application services |
| Ollama with a Qwen3 embedding model | Generate embeddings locally for document chunks and questions |
| SQL Server 2025 | Store document text, metadata, indexing state, and vectors; perform similarity search |
| Official MCP C# SDK | Expose selected retrieval operations as tools over Streamable HTTP |
| Open WebUI | Provide chat sessions and connect the model to the retrieval tools |
| OpenAI API initially; Ollama chat models later | Generate answers and compare tool-calling behavior |
The application will have an indexing path and a question-answering path. This diagram shows the planned responsibilities; it does not represent a tested deployment.
The planned architecture, with separate paths for indexing documentation and answering questions.
On the indexing side, my application will prepare the selected notes, ask Ollama for embeddings, and store the results in SQL Server. The source text needs to stay attached to each result so I can trace an answer back to the note it came from.
On the question side, Open WebUI will connect the chat model to my search tools through MCP, the Model Context Protocol. Those tools will call the retrieval service and return source excerpts. MCP provides the tool interface; the application still owns the search logic.
Open WebUI is a deliberate part of that separation. I’ll use it for conversations and tool calling while building my own knowledge index and retrieval pipeline. Its built-in RAG functionality would take over much of the work I want to understand.
The initial design keeps the notes, index, and embedding generation local. Answer generation through the OpenAI API means that retrieved excerpts sent as context leave the local environment. I’ll need to select and sanitize the notes with that boundary in mind. Local Ollama chat models will be a later experiment.
Exact models, service placement, and the database search approach still need to be selected and tested. I’ll start with hardware I already have and record those choices in the implementation posts, along with any limitations I encounter.
What’s Next
Part 2 will focus on preparing the initial knowledge base: selecting notes, removing sensitive material, preserving Markdown structure, and trying the first chunking approach. I want the source material to be manageable before connecting it to a chat model.
Parts 3 and 4 will cover local embeddings with Ollama and vector storage and search in SQL Server 2025. Parts 5 and 6 will expose the retrieval service through an ASP.NET Core MCP server and connect it to Open WebUI. Part 7 will look at keeping the index current as the notes change. Part 8 will bring the experiments together to evaluate retrieval, answers, and performance.
I expect some of these choices to change once I have results. I’ll document those changes and the reasons for them, including attempts that don’t work. A small system I can understand and troubleshoot is the first goal; a useful assistant for my homelab documentation is the result I’m working toward.
Want to share your thoughts or ask a question?
This blog runs on coffee, YAML, and the occasional dad joke.
If you’ve found a post helpful, you can
support my work or
☕ buy me a coffee.
Curious about the gear I use? Check out my smart home and homelab setup.
