All posts
GuideKnowledge Base

How to Build a Proposal Knowledge Base That Actually Wins Bids

A practical, structure-first guide to the corpus behind AI-generated proposals — what to upload, what to keep out, how guidelines differ from knowledge, and the maintenance ritual that keeps answers current.

Alex GertzFounder, BidAuthorAugust 27, 20265 min read

Every AI-generated proposal answer has a ceiling, and that ceiling is not the model — it is the corpus the answer is grounded in. Feed a grounded generation system a thin, stale, contradictory knowledge base and it will faithfully produce thin, stale, contradictory answers with tidy citations. Feed it a deliberate one and the first drafts start landing close enough to final that your team's hours move to strategy.

This guide is the practical version of that sentence: what goes in, what stays out, how knowledge differs from guidelines (they are different inputs doing different jobs), and the small maintenance ritual that keeps the whole thing trustworthy. It expands steps one and two of the BidAuthor workflow, but the principles hold for any grounded system.

The two-corpus model: knowledge vs. guidelines

The single most common setup mistake is dumping everything — facts, style guides, old proposals, do's-and-don'ts — into one pile. Grounded systems work dramatically better when you separate two kinds of material, because they are consumed differently:

Knowledge is what is true about your organization: capabilities, methodology, security posture, staffing, past performance, certifications. Knowledge is retrieved per question — the staffing question pulls staffing material — and quoted with citations.

Guidelines are how your organization answers: voice and tone, structural rules ("every answer leads with the direct answer, then evidence"), claims you never make, phrasing legal has approved, page-limit discipline. Guidelines are not retrieved selectively; they steer every answer, the way a style guide sits on every writer's desk.

Rule of thumb: if a document should influence one answer, it's knowledge. If it should influence every answer, it's a guideline.

BidAuthor makes this split physical — knowledge-base documents and guideline documents live in separate workspaces — but even in a folder on a shared drive, the separation pays.

What belongs in the knowledge base

Ranked by return on upload effort:

  1. Your best past proposals — the wins. The densest source of evaluator-tested language about what you do. Two or three strong, recent, winning responses beat twenty mediocre ones; volume of weak material actively dilutes retrieval.
  2. Capability and methodology documents. The how-we-deliver material: implementation methodology, project governance, QA approach, escalation paths. These answer the "describe your approach to…" family, which dominates most narrative solicitations.
  3. Security and compliance documentation. SOC 2 summaries, data-handling policies, business-continuity plans, insurance certificates. Security questionnaires and DDQs are the most template-able solicitation type — this is where automation rates run highest.
  4. People material. Bios, role descriptions, org overviews, key-personnel resumes. "Identify the personnel assigned to this engagement" appears in nearly every services RFP.
  5. Case studies and past performance. Client-approved, outcome-focused, with concrete scope and (where permitted) numbers.
  6. Commercial scaffolding. Rate-card structure, pricing methodology explanations, standard assumptions and exclusions — the prose around pricing, which pricing questions always request.

What to keep out

Exclusions matter as much as inclusions, because retrieval treats everything in the corpus as a candidate truth:

  • Losing proposals — unless you know why they lost and the content wasn't the reason.
  • Drafts and superseded versions. Two versions of the same methodology document is a contradiction generator. One canonical version per fact.
  • Anything client-confidential you don't have reuse rights to.
  • Marketing fluff without factual content. "We are passionate about excellence" retrieves well and says nothing; it teaches the system your voice is empty calories.
  • Aspirational capability claims. If you wouldn't defend it in an oral presentation, it does not belong in the corpus. Grounded generation means the system will use what you gave it.

Writing guidelines that actually steer

Good guidelines read like instructions to a smart new proposal writer, not like a brand book. The patterns that work:

  • Be imperative and specific. "Open every answer with a direct one-sentence response to the question, then supporting detail" beats "we value clarity."
  • Encode your red lines. "Never commit to indemnification language; flag those questions for legal." "Never claim certifications not listed in the certifications document." These become machine-enforced discipline.
  • Give before/after examples. One rewritten answer teaches tone better than three paragraphs describing it.
  • Keep each guideline document focused. A tone guide, a structure guide, a claims policy — separable rules a team member could follow, not one 40-page wall.

In BidAuthor, guidelines are compiled and applied to every generation pass, and per-run steering ("more formal, emphasize the incumbent transition risk") layers on top for a specific pursuit without editing the standing rules.

Structure and hygiene that improve retrieval

A few mechanical habits raise answer quality noticeably:

  • Descriptive filenames and headings. Retrieval sees document structure; "Implementation_Methodology_2026.pdf" with real section headings outperforms "final_v3_FINAL.pdf."
  • One topic per document where possible. Focused documents retrieve cleanly; omnibus documents drag unrelated context into answers.
  • Prefer text-native formats. Clean PDFs, DOCX, and Markdown ingest with structure intact; scanned images and slide-art lose fidelity.
  • State facts with their scope. "As of 2026, we operate three delivery centers (Denver, Austin, Kraków)" ages visibly; "we have several offices" ages invisibly, which is worse.

The maintenance ritual: 30 minutes a quarter

A knowledge base is not a project, it is an asset with a decay rate. The teams whose answer quality rises over time run a small loop:

  1. After every submitted proposal, promote the best new answers into the knowledge base — the corpus should compound with each pursuit.
  2. Quarterly, re-verify the volatile facts: headcount, certifications, client references, anything with a date. Retire superseded documents the same day their replacement lands, never "later."
  3. Watch the gaps the system reports. Every "no source found" or partial answer is a free audit of what your organization has not written down. That list is your KB backlog, prioritized by real demand.

The compounding effect is the whole point: every answered solicitation should make the next one cheaper. Teams that skip the ritual plateau at "decent first draft"; teams that run it converge on "review and ship."

Starting from zero, this week

If you have no knowledge base at all: collect your three best winning proposals, your security/compliance pack, your methodology document, and your team bios. That is genuinely enough to run a real solicitation end to end and see where the gaps are — the system will tell you what's missing faster than any planning exercise. Write two guideline documents (tone + red lines), upload, and generate against a live RFP on the free tier.

The gap report from that first run is your roadmap. From there, the ritual takes over, and the asset starts compounding.

See it on your own RFP

Upload a real solicitation, watch every requirement get extracted, and judge the generated answers against your own knowledge base — free.