Is This Use Case Worth Building?
Most organisations have a list of things they could do with a private LLM and no way to rank them. This scorecard takes one candidate at a time and scores it on the seven things that decide whether the build delivers: data, volume, error tolerance, source documents, onshore need, a measurable outcome and a human review path.
Pick one use case and answer for it alone. Run the scorecard again for each candidate and compare the scores. A use case that scores low is not a bad idea, it is one that needs preparation before money is spent.
LLM Use Case Scorecard
Eight questions about one use case. Answer for how things are today.
A ranking tool to structure an internal decision, not a technical assessment. Nothing entered is recorded or transmitted.
What the Score Means
The score is a ranking tool, not a guarantee. It rewards use cases where the inputs exist, the output can be checked and the failure cost is manageable.
Build now
The documents exist, the task repeats often enough to matter, an occasional wrong answer is caught before it causes harm, and you can measure whether it worked. These are the use cases that pay for a first deployment and build internal confidence.
Pilot first
Most of the pieces are in place but one or two are uncertain, usually error tolerance or the availability of clean source documents. A short pilot on real questions with a fixed evaluation set will settle the doubt cheaply.
Not yet
The use case depends on data that is not written down, cannot tolerate mistakes, or has no one checking the output. Building now would spend most of the budget on preparation. Do that preparation first, then score it again.
How the Scoring Works
Eight questions, each scored zero to three. Every question carries the same weight, so the maximum is 24 and the bands are set as a share of that.
Answer each question for one use case
The options run from the most favourable condition for a build at three points to the least favourable at zero. Pick the one closest to the truth today, not the one you hope will be true after the project.
Points are added
No weighting. A use case cannot compensate for having no source documents by scoring well on volume, and the flat scoring makes that visible.
Score becomes a percentage of 24
Seventy percent and above is build now. Forty to sixty-nine percent is pilot first. Below forty is not yet.
Compare candidates
Run the scorecard once per use case. The highest scorer with a measurable outcome is normally the right first deployment, even if it is not the most exciting one.
Why These Seven Factors
Each factor maps to a way private LLM projects fail in practice. The scorecard exists so those failures are found on a web page rather than in month three of a build.
Source documents and data sensitivity
A retrieval system can only answer from what is written down and indexed. If the knowledge is in people's heads, the model has nothing to retrieve. Sensitivity matters separately: sensitive data can be handled in a private deployment, but only if you know who may see what before anything is indexed.
- Documents must exist, be current and have a known owner
- Permissions must be enforceable at document level
- Personal information triggers Australian Privacy Principles obligations
- Mixed sensitive and general content in one repository needs separation first
Volume and error tolerance
Volume decides whether the build is worth it. Error tolerance decides whether it is safe. A task done fifty times a day where a mistake is caught on review is close to ideal. A task done twice a month where a mistake reaches a customer is not, however impressive the demo.
- High volume justifies the build cost through the value calculators
- Low tolerance for error needs a review step in the workflow
- Draft-and-review use cases tolerate error better than auto-send ones
- Score the failure cost honestly, including reputational cost
Onshore processing
Some organisations must keep processing in Australia because of contracts, regulator expectations or policy. That is not a reason to avoid AI, it is a design constraint, and it is settled if you already have an Australian-region or on-premises option. The score drops when nobody has asked the question yet.
- Check contract clauses on data location before choosing a platform
- Region of inference matters as much as region of storage
- A private deployment in an Australian region satisfies most requirements
- Use the AI data residency quiz if you are unsure where data goes today
Measurable outcome and human review
If you cannot say what a good result looks like, you cannot evaluate the system, and you will not know whether to expand or stop. A named reviewer who sees the output before it has effect is the single cheapest control available and it makes low-tolerance use cases viable.
- Write down the metric before the build starts
- Collect fifty real questions with expert answers as the evaluation set
- Name the reviewer and put review inside the workflow, not beside it
- Log what the reviewer changes: it shows where the system is weak
Next Steps
Private LLM Readiness Assessment
Score the organisation rather than the use case: content, permissions, governance and evaluation capability.
Assess readiness →RAG vs Fine-Tuning Decision Tool
Once a use case scores well, work out which architecture fits it.
Choose an architecture →Document Search Time Cost Calculator
Put a value on the most common first use case: finding things in your own documents.
Value the use case →Frequently Asked Questions
One at a time. The scorecard is designed to compare candidates, and mixing two in one pass hides the weak one behind the strong one. Write the score down, refresh the page or change your answers, and score the next candidate. Then rank them.
No. Sensitive data is one of the main reasons organisations choose a private deployment over a public tool. The score drops when sensitivity is high and permissions are unclear, because that combination means the preparation has not been done. Sensitive data with accurate, enforceable permissions scores reasonably well.
Score it as it is today and treat writing the documents as the preparation step. Capturing expert knowledge into a maintained document set is worthwhile regardless of the AI project, and it usually takes weeks rather than months. Re-score when the documents exist.
Because it converts a use case that cannot tolerate error into one that can. A model that drafts a response for a person to check and send is safe in most contexts. The same model sending the response itself is not. The review step is the difference, and it costs almost nothing to include.
The readiness assessment scores the organisation: where content lives, whether permissions are enforceable, who approves AI deployments. This scorecard scores one use case within that organisation. A ready organisation can still pick a poor first use case, and a well chosen use case can still be blocked by organisational gaps. Use both.
No. Your answers stay on the page and are not recorded or transmitted. Refreshing the page clears them.
Sources and further reading
- Australian Privacy Principles (Office of the Australian Information Commissioner)
Scored a Use Case at Build Now or Pilot First?
Send us the use case, the score and what you answered on source documents and review. We will come back with a scoped pilot or build, a price within our published range and the evaluation set we would use to judge it.