How I documented a legacy VB.NET system using AI in 3 days.
TL;DR — A VB.NET module the size of a small book, untouched since 2014, with no docs and the original author long gone. We documented it in three days with a VS Code plugin running a structured agent, a prompt that demanded specific outputs, and a human review pattern that scaled. The trick is not the AI — it's the structure around the AI.
What we were looking at
Jobscope's VB.NET codebase came with the kind of documentation backlog every long-running enterprise app has. About 240,000 lines, organised across modules that nobody on the current team had originally written. The code worked. Nobody knew exactly why it worked, in the sense of "what business decision in 2008 led to this branch existing." The team's estimate to document it the conventional way — a developer reading code, writing wiki pages, peer-reviewing — was four to six months.
The constraint we had: a .NET 8 + React migration was queued behind the docs. Every week the docs slipped, the migration slipped. Six months of slippage was not on the table.
What we shipped instead was a VS Code plugin with an internal name nobody loved ("Code Gini"), a strict prompt template, and a review process that ran in parallel to the agent. Three days in, the first 30 modules were documented to a higher standard than the team would have produced in a sprint. Three weeks in, the whole thing was done.
What the agent actually did
The agent didn't "read code and write docs". That's the demo version. The production version was structured: it produced four specific outputs per module, in a fixed format, every time.
- The summary — one paragraph, what the module is for, in plain language
- The interface — the public methods, their parameters, the types coming in and out
- The business rules — extracted from the code as numbered statements, each with the file:line reference
- The known unknowns — places the agent was uncertain, with the specific line and the question it would have asked the original author
The fourth one is the one that made it work. By making the agent name what it didn't know, we turned the document from "AI-generated paragraph soup" into "a list of things a human still needs to verify". The human review became targeted — read the unknowns, spot-check the business rules, ignore the parts the agent was confident about.
The prompt template
This is the version we converged on after about a day of iteration. Not the first prompt, not the prettiest — the one that produced the most consistent output across modules.
You are documenting a legacy VB.NET module. Your job is to produce
a structured analysis. You are NOT writing prose; you are filling
in a template.
INPUT: the file at {path}
OUTPUT: the four sections below, in this exact order, with
these exact headings.
## 1. Summary (one paragraph)
What does this module do, in plain language a non-developer can
understand? Limit: 4 sentences.
## 2. Public interface
List every Public Sub, Function, or Property. For each:
- Name
- Parameters (name : type)
- Returns (type)
- One-line description of what calling it does
## 3. Business rules
Extract the business rules encoded in the code as numbered
statements. Format:
1. [Rule, plain language] — see ModuleName.vb:123
2. [Rule] — see ModuleName.vb:145
For each rule, include the line reference where it is enforced.
Skip rules that are pure plumbing (null checks, error returns).
## 4. Known unknowns
List specific places where the code's intent is unclear and a
human would need to verify with the original author or a stakeholder.
Format each as a question, with a line reference. Examples of what
counts as unclear:
- Magic numbers without context
- Branches that can never fire (or appear to)
- Comments that contradict the code
- Calls into other modules whose behaviour can't be inferred
If a section has no content (e.g., no business rules in a pure
data module), write "None." Do not invent.
Do not produce sections beyond these four.
Do not editorialise about code quality.
Do not suggest refactors.
Every word of "do not" in that prompt was added because of an output we got and didn't want. The "do not editorialise" line came from a 600-word digression on naming conventions in the second module's output. The "do not invent" line came from a "business rule" the agent confidently asserted that turned out to be wrong.
The VS Code plugin part
The plugin was thin: it knew where the code lived, it knew which modules had been documented, it ran the agent against the next module, it wrote the output to a markdown file in a parallel docs/ tree. The substantive part is what was deliberately not in the plugin:
- No "review the output" step in the plugin. Reviews happened in PRs against the docs tree, like normal code review.
- No auto-merge. Every doc was reviewed by a human before it became canonical.
- No "agent fixes its own output". The agent's job was first draft, exactly once. Corrections were human.
The temptation with these tools is to have the agent loop until it's "happy". Don't. The agent doesn't have judgement; loops just generate more text to read. One pass, structured output, human review.
The review pattern that scaled
This is where the project would have collapsed without discipline. The first day, two reviewers tried to read each generated doc end-to-end. They were exhausted by lunch. We changed the pattern:
- Skim the summary — does it match the reviewer's mental model of the module? If yes, continue. If no, the agent is confused; investigate before approving.
- Read the known-unknowns first. These are the places the agent self-flagged. Treat each as a real question the documentation needs to answer; spend time here.
- Spot-check 3 random business rules against the file:line references. If they match, trust the rest. If they don't, push back to the agent with a more specific prompt.
- Skim the public interface for completeness. Easy to verify with the IDE's outline panel.
Average review time per module dropped from ~45 minutes (read everything) to ~12 minutes (this pattern). Quality went up, because the time saved went into resolving the known-unknowns instead of skimming paragraphs the agent had already written confidently.
What the docs looked like, end of week 1
Per module, we ended up with a markdown file that read roughly like this (sanitised excerpt):
## 1. Summary
Computes the eligible discount band for a given customer order
based on customer tier, order volume, and prior commitments. The
output drives the price-quote generation step. Used by the order
entry screen and the bulk-import workflow.
## 2. Public interface
GetDiscountBand(customerId : Integer, orderTotal : Decimal,
priorCommitments : Decimal) : DiscountBand
Returns the discount band the customer is eligible for given
the inputs.
ResetCustomerCommitments(customerId : Integer) : Boolean
Resets the running 'prior commitments' counter to zero.
Used during fiscal year rollover.
## 3. Business rules
1. Tier-1 customers are eligible for the 'platinum' band only
above ₹500,000 — see DiscountBand.vb:142
2. Prior commitments older than 365 days are excluded from the
running total — see DiscountBand.vb:188
3. The discount band cannot exceed the customer's contractual
ceiling stored in tbl_CustomerContract.MaxDiscountPct —
see DiscountBand.vb:212
## 4. Known unknowns
- Line 156: there is a hardcoded list of customer ids
(1024, 1031, 1899) that bypass the volume threshold.
Why? Were these grandfathered in?
- Line 230: the function returns 'silver' on exception.
Is this intentional graceful degradation, or a bug?
- Line 271: a commented-out branch references a 'Q4_2018'
flag. Was this a temporary policy?
The "known unknowns" section turned every module review into a productive conversation with the team's senior developers — who knew, or could find, the answers. The module above led to two product-team conversations, one historical price agreement that was no longer valid, and one bug that had been silently degrading prices for two years.
Where this approach doesn't work
- If the code is heavily reflective or generated. The agent reads the visible code; if half the behaviour is in T4 templates, generated by a tool the agent doesn't have access to, you'll get plausible nonsense. Document the generator first, then the generated code.
- If your codebase has no naming discipline. Modules called
Utils,Helper,Common,Miscdefeat the agent. The agent assumes meaning from names. Without it, every output is a guess. - If the team doesn't have the bandwidth to review. The bottleneck of this approach is human review time, not agent time. If you can't dedicate one or two engineers for the review window, the docs will pile up un-validated and the project will fail in a different way than the conventional one would have.
The numbers, end of project
- Modules documented: 142
- Total agent time: ~38 hours of LLM inference
- Total human review time: ~52 hours across two reviewers
- Estimate before AI: 4–6 months for one developer
- Actual elapsed time: 3 weeks
- Bonus discoveries: 14 historical bugs, 6 obsolete features still running, 2 commercial agreements no longer valid
The bonus discoveries — the things the team learned about its own codebase that no one had bothered to write down — are arguably the bigger win than the documentation itself. We weren't just writing docs. We were running an audit.
Sitting on a legacy codebase nobody has time to document?