Ready to Build a Private LLM?
Private LLM projects almost never fail on model capability. They fail on document permissions nobody mapped, a use case nobody defined precisely, or an inability to tell whether the output is any good. Fourteen questions on the things that actually decide it.
Answer for the organisation as it is today. A low score is common and fixable, and knowing which dimension is weak is worth far more than a flattering result.
Private LLM Readiness Assessment
Fourteen questions about your organisation, not about models.
A self-assessment to structure an internal conversation, not a technical audit. Nothing entered is recorded or transmitted.
What Actually Determines Success
The technology is largely a solved problem. What varies between organisations — and what decides whether a project delivers — is the state of the material the system has to work with.
Permissions are the hardest part
Ensuring users can only retrieve documents they were already entitled to see is consistently the most underestimated component of an enterprise deployment. Get it wrong and you have built an extremely efficient tool for surfacing information people should not have.
Vague use cases produce vague results
"An AI assistant for staff" is not a use case. "Answer policy questions from the HR handbook with citations" is. Precise scope is what makes a project evaluable, and evaluable projects are the ones that get improved rather than abandoned.
If you cannot measure it, you cannot improve it
Without an evaluation set — real questions with known good answers — nobody can say whether a change made the system better or worse. Teams without one end up tuning on impressions and arguing about anecdotes.
The Five Dimensions
Fourteen questions. Permissions and use case clarity carry the most weight because they cause the most project failures.
Content and data
Whether the documents exist in accessible formats, are reasonably current, and can be retrieved programmatically.
Permissions and access
Whether document-level entitlements are known and enforceable, or whether access has been managed informally.
Use case definition
Whether the intended questions, users and success criteria are specific enough to build against and evaluate.
Governance
Whether the organisation has decided on acceptable use, data handling and who signs off on AI deployments.
Evaluation capability
Whether subject-matter experts are available to define correct answers and judge output quality.
The Four Preparation Steps Worth Doing First
Each is achievable in weeks, each reduces build cost, and each is worth doing whether or not the project proceeds.
Map where the documents actually are
Most organisations discover their knowledge is spread across a document management system, several shared drives, a wiki, an email archive and a number of personal folders. Establishing what exists and where, before scoping, prevents the most common cause of mid-project scope expansion.
- Inventory the repositories that hold relevant knowledge
- Identify which are authoritative and which are stale copies
- Check formats — scanned PDFs need different handling from text
- Estimate volume, since it drives both cost and architecture
Establish document-level entitlements
Before any retrieval system is built, establish who is permitted to see what and whether that is enforceable programmatically. Organisations that have relied on obscurity rather than permissions face a genuine problem here, and it is far better discovered now.
- Confirm existing permissions are accurate, not merely present
- Identify content protected only by nobody knowing where it is
- Decide how permissions will be enforced at retrieval time
- Plan for permission changes to propagate to the index
Write down fifty real questions
Collect fifty questions actual users would genuinely ask, in their own words, and get subject-matter experts to write the correct answers. This becomes your evaluation set, your scope definition and your acceptance criteria all at once, and it takes a few days.
- Collect real questions from real users, not hypothetical ones
- Have experts write the correct answer and cite the source document
- Include questions the system should decline to answer
- Keep the set stable so results are comparable over time
Decide the deployment boundary early
Whether the model runs in a public cloud API, a private cloud tenancy or on your own infrastructure changes cost, timeline and vendor options substantially. Deciding late means re-architecting; deciding early means the whole project is scoped against a real constraint.
- Establish any data residency or sovereignty requirement up front
- Confirm whether public API processing is acceptable for this content
- Check contractual obligations to clients about data handling
- Document the decision and the reasoning behind it
Next Steps
RAG vs Fine-Tuning Decision Tool
Work out which architecture your requirements actually call for.
Choose an approach →AI Data Sovereignty Checklist
The residency and governance questions to settle before any data moves.
Open the checklist →Build vs Buy Calculator
Check whether building is the right call at your user count.
Compare the options →Frequently Asked Questions
Permissions, by a considerable margin. A retrieval system that ignores document-level entitlements will happily surface salary information, disciplinary records, commercially sensitive material or board papers to anyone who asks a well-phrased question. Organisations that have historically relied on content being hard to find rather than properly restricted discover this the moment the system works well. The second most common cause is an undefined use case, which produces a system nobody can evaluate and therefore nobody trusts.
Far less than most teams assume, provided you scope narrowly. Start with a single well-defined domain — the HR policy set, the technical documentation for one product, the standard operating procedures for one function — rather than attempting the entire organisational corpus. A focused corpus of a few hundred current, authoritative documents produces a genuinely useful system, and it demonstrates value quickly. Attempting everything at once means spending months on data preparation before anyone sees a working result.
You need decisions rather than a document, though writing them down is sensible. At minimum, settle what data may be processed and where, who approves AI deployments, what the system may and may not be used for, and how you will handle the fact that outputs are probabilistic rather than deterministic. A single page covering those points is sufficient to begin, and it can be expanded into a fuller policy once you have practical experience to base it on. Starting with no decisions at all means they get made implicitly by whoever configures the system.
A spreadsheet is entirely adequate. Fifty to two hundred rows, each containing a real question a user would ask, the correct answer, and the source document that supports it. Include questions the system should refuse or escalate, since knowing when not to answer is as important as answering well. Run it after every material change to the system and record the scores. This is unglamorous and it is the single highest-value artefact in the project, because it converts arguments about quality into measurements.
Yes, and it is strongly the better approach. A narrow first deployment over one document domain for one user group delivers value in weeks, surfaces the permission and data issues while they are still cheap to fix, and builds the organisational confidence needed for anything larger. Broad first deployments tend to spend months in preparation, arrive late, and disappoint because expectations have had time to inflate. Narrow and working beats broad and pending.
No. It runs entirely in your browser, nothing is transmitted, and there is no email gate. The questions ask only about organisational processes and capabilities, and no confidential detail is requested. You are welcome to screenshot the result for an internal discussion, and equally welcome to use the tool without ever contacting us.
Know Where You Stand?
Tell us your weakest dimension and the use case you had in mind. We will tell you what to fix first — and if the honest answer is that you are not ready, we will say so.