AI for Knowledge Management: Why Your Firm Search Still Fails
A firm buys an AI search tool, points it at twenty years of work, and asks it a question. It returns seven documents. Nobody can tell which one is current, so the person who asked walks down the hall and asks a colleague, which is what they did before.
TL;DR
- The archive is not the problem. The signals are. Retrieval can only read what somebody recorded, and most firm documents record nothing about their own standing.
- Four signals decide whether a found document is usable: currency, authority, context and reach.
- The last-good-version problem is the specific failure. A firm holds many versions of the same thing and no marker for which one is the last good one.
- Better retrieval makes a weak archive worse, because it surfaces confidently what used to stay buried.
- Fix the signals on the twenty things people actually reuse before buying anything. That work is boring, cheap and the reason the tool later works.
- Knowledge management is a maintenance commitment, not a project with an end date. Decide who owns it before you start.
The question your archive cannot answer
Ask a partner which of two documents is the one to reuse and they answer in seconds. They remember the matter, they know who wrote it, and they know the second one was drafted after the rules changed. That knowledge is real and it lives nowhere except in that partner.
Now ask the same question of the file server. It can tell you a modification date, which is often the date somebody opened a file and pressed save by accident. It can tell you a folder name chosen by whoever set up the matter. It cannot tell you which document is the one to reuse, because nobody ever wrote that down.
This is the last-good-version problem, and it is the actual constraint on firm knowledge. A firm suffers from having a great deal written down with no way to tell which of it still holds.
Better search makes a weak archive worse
This is the part that surprises firms, so it is worth saying plainly before anything else.
When search is poor, people do not trust it. They use it as a starting point, then verify what it gives them, then ask a colleague anyway. The weakness of the tool keeps a human check in the loop.
When retrieval becomes good, the check quietly disappears. The system returns a confident, well-written answer assembled from documents nobody has evaluated. A junior professional has no way to know that the precedent it drew from was superseded two years ago, because the document does not say so and the system has no way to infer it. The output reads exactly as well as a correct one would.
So a firm that improves retrieval without improving signal has not made its knowledge more available. It has made its stale knowledge easier to reach and harder to question.
Worth keeping
- Retrieval quality and archive quality are separate problems, and buying the first does not fix the second.
- A confident wrong answer is more expensive than a slow right one, because it removes the moment where somebody would have checked.
- If your current search is bad and people work around it, that workaround is a control. Do not remove it without replacing it.
The four signals
A retrieval system ranks documents using whatever the documents and their surroundings tell it. If your archive does not record these four things, no tool can weigh them, however good the model behind it is.
Is this still true today? A modification date does not answer that. A review date with a named owner does, and it also tells you when nobody has looked at something for three years, which is information in itself.
Who stood behind this? A draft that a partner approved and a draft somebody abandoned look identical to a file system. Recording an approver separates the firm's position from one person's working file.
What did this assume? Work product is correct relative to a set of facts, a jurisdiction and a date. Strip those away and you have text that reads authoritative and may be wrong for the matter in front of you.
Who is allowed to see it? Confidentiality obligations do not relax because a search tool made retrieval easier. Permissions have to hold at the source, so that a system cannot surface something to somebody who should never have reached it.
Reach deserves its own warning. A retrieval system inherits whatever access model sits underneath it, and firms regularly discover during a rollout that their folder permissions have been broken for years without anyone noticing, because nobody was searching across the whole archive. Nothing exposed the gap until something finally read everything at once.
What to fix before you buy anything
The instinct is to fix the archive. Twenty years of documents, every one of them missing all four signals. It is an unbounded task and the reason knowledge management initiatives are abandoned.
Do not fix the archive. Fix what people reuse.
-
Find the twenty things people actually reuse
Every firm has a small set of documents that gets reused constantly: a handful of templates, a few methodologies, some standard sections, a pricing structure. Ask five people what they open when they start something new and the same items appear on every list. That list is short, and it is where all the value is.
-
Name an owner for each one
One person, by name, who decides what the current version is. Not a committee and not a practice group. If nobody will accept the name against an item, that item is not actually a firm standard, which is worth learning now rather than after a tool has been ranking it as one.
-
Mark the last good version and retire the rest
For each item, the owner picks the version that stands, records the date they decided, and moves the alternatives somewhere retrieval does not reach. This is the step that gets resisted, because retiring a document feels like losing something. Keeping nine versions is what loses it.
-
Write down what it assumes
Two or three lines at the top of each item: what it is for, what it assumes, and when somebody should look at it again. This is the highest-value writing in the whole exercise, because it is the only signal that tells a reader when a document does not fit their situation.
That is a week of work for most firms, and it can be done before any procurement decision. If a tool is later bought, it lands on an archive with something to read. If no tool is ever bought, the firm is still better off, because the humans doing the asking now get better answers too.
The test of a knowledge base is not whether the document can be found. It is whether the person who finds it can tell, without asking anyone, whether they are allowed to rely on it.
What changes about how people work
Two things change, and only one of them is technical.
The technical change is that retrieval starts to work on the material that matters, because that material now carries signal. Questions that used to return seven undated versions return one current document with an owner and a date.
The behavioural change is larger and slower. Somebody has to keep the signals true. A review date that passes without a review is worse than no review date, because it asserts something false. This is why knowledge management is a maintenance commitment rather than a project: the value decays on a schedule unless somebody is accountable for renewing it.
Firms that get this right tend to make it small and specific. One owner per item, one review a year, fifteen minutes each. Firms that get it wrong tend to announce a programme, build a taxonomy nobody uses, and quietly stop.
Worth keeping
- Scope the work to what people reuse, not to what the firm has stored. The first is a week and the second never ends.
- An owner by name is the load-bearing part. Everything else is recording what that person decided.
- Budget the annual review before the rollout, because that is the cost the business case usually omits.
Common mistakes
- Buying retrieval before fixing signal. Why it fails: the tool ranks on information that does not exist, so it returns confident answers drawn from documents nobody has evaluated. Better: spend a week on the twenty reused items first, then evaluate tools against an archive that can be ranked.
- Trying to classify everything. Why it fails: the task is unbounded, the value is concentrated in a small fraction of the material, and the programme dies before it reaches that fraction. Better: fix what people reuse and leave the rest searchable but unmarked.
- Assigning ownership to a group. Why it fails: a practice group cannot decide which version is current, so the decision never gets made and the signal stays absent. Better: one name per item, accepted by that person, recorded where the document lives.
- Treating access control as a tool setting. Why it fails: the retrieval system inherits the permissions underneath it, and those are usually looser than anyone assumes. Better: audit permissions at the source before connecting anything that reads across the whole archive.
- Setting review dates nobody honours. Why it fails: a lapsed review date asserts currency that was never checked, which is worse than admitting the document is unreviewed. Better: fewer items, real dates, and a calendar entry for the owner.
- Assuming the archive is the knowledge. Why it fails: a large part of what a firm knows is why something was done, and that was never written anywhere. Better: capture the reasoning on the items you are already touching, while the people who remember it are still there.
What this article does not claim
It does not claim a figure for time saved. Numbers circulate for this subject and the ones that can be traced lead back to product marketing, so none is published here.
It does not claim that any particular tool solves this. The four signals are a property of your archive, and every retrieval system has to read them from somewhere.
It claims the first useful week is short and cheap. The maintenance after it runs indefinitely.
- Knowledge management
- The practice of making what an organisation already knows available to the people who need it, reliably enough that they act on it.
- Retrieval
- The step where a system finds candidate documents in response to a question, before anything is generated or summarised from them.
- Signal
- Anything recorded about a document that lets a system or a person judge its standing, such as an owner, an approval, a review date or a stated assumption.
- Currency
- Whether a document is still true now, as distinct from when its file was last modified.
- Authority
- Whether somebody with standing in the firm approved the document as the firm's position.
- Context
- The facts, jurisdiction and date a document assumed, without which its conclusions cannot be transferred safely.
- Last good version
- The version an accountable owner has designated as the one to reuse, recorded with the date of that decision.
- Taxonomy
- A classification scheme for organising material. Useful when small and maintained, and the most common place a knowledge programme stalls when it is neither.
Questions firms ask
Do we need to clean up twenty years of documents first?
No, and attempting it is the most common reason these projects fail. The material people actually reuse is a small fraction of what is stored. Fix that fraction, leave the rest searchable, and revisit only if a specific gap shows up in use.
Will AI not just work out which document is current?
It can infer from whatever is recorded, such as dates in the text or explicit supersession notes. Where nothing is recorded, there is nothing to infer from, and the system will rank on similarity to your question instead, which has no relationship to whether a document still holds.
Who should own this in a firm without a knowledge role?
The people who own the individual items, which is usually whoever is already the informal authority on each. A central role helps coordinate later, and waiting to hire one is a common way to never start.
What about confidentiality between matters and clients?
It has to be enforced where the documents live, not in the search layer, because any system reading across the archive inherits whatever permissions are already in place. Audit those before connecting anything, and expect to find gaps that were invisible while nobody could search widely.
Is this worth doing if we are not buying an AI tool?
Yes. The four signals are what a colleague uses when they answer the question in the corridor. Recording them helps people directly, and it is the same work you would have to do before any tool could help.
How do we stop it decaying again?
One named owner per item and one short review a year, scheduled rather than intended. Decay is the default, and the only thing that resists it is somebody whose name is attached to the item.
What breaks first when this is skipped?
Trust. A confident answer built on a superseded document is discovered once, and after that people go back to asking a colleague, which returns the firm to where it started with a licence cost attached.