Inventory
Before anything is indexed we map the corpus as it actually is: which repositories exist, who can see what, and what state the documentation is in. We flag superseded documents, duplicates where two versions are both live, and scanned PDFs that need OCR before they can be read at all. The output is a list of sources that go into phase one, the ones that stay out, and the clean-up work to do first, which shapes the final result more than anything else.