Skip to content
Skip to content

Is This Use Case Worth Building?

Most organisations have a list of things they could do with a private LLM and no way to rank them. This scorecard takes one candidate at a time and scores it on the seven things that decide whether the build delivers: data, volume, error tolerance, source documents, onshore need, a measurable outcome and a human review path.

Pick one use case and answer for it alone. Run the scorecard again for each candidate and compare the scores. A use case that scores low is not a bad idea, it is one that needs preparation before money is spent.

LLM Use Case Scorecard

Eight questions about one use case. Answer for how things are today.

0 of 8 answered0%
1.How sensitive is the data this use case would touch?

Sensitivity is fine in a private deployment if permissions are known.

2.How often does this task happen?
3.What happens if the output is wrong?
4.Do the source documents the system would answer from exist?
5.Is there a requirement to process this data in Australia, and is it settled?
6.Can you measure whether it worked?
7.Who checks the output before it has effect?
8.How well defined is the task itself?
Answer all 8 questions to see your score and a tailored recommendation.

A ranking tool to structure an internal decision, not a technical assessment. Nothing entered is recorded or transmitted.

What the Score Means

The score is a ranking tool, not a guarantee. It rewards use cases where the inputs exist, the output can be checked and the failure cost is manageable.

Build now

The documents exist, the task repeats often enough to matter, an occasional wrong answer is caught before it causes harm, and you can measure whether it worked. These are the use cases that pay for a first deployment and build internal confidence.

Pilot first

Most of the pieces are in place but one or two are uncertain, usually error tolerance or the availability of clean source documents. A short pilot on real questions with a fixed evaluation set will settle the doubt cheaply.

Not yet

The use case depends on data that is not written down, cannot tolerate mistakes, or has no one checking the output. Building now would spend most of the budget on preparation. Do that preparation first, then score it again.

How the Scoring Works

Eight questions, each scored zero to three. Every question carries the same weight, so the maximum is 24 and the bands are set as a share of that.

1

Answer each question for one use case

The options run from the most favourable condition for a build at three points to the least favourable at zero. Pick the one closest to the truth today, not the one you hope will be true after the project.

2

Points are added

No weighting. A use case cannot compensate for having no source documents by scoring well on volume, and the flat scoring makes that visible.

3

Score becomes a percentage of 24

Seventy percent and above is build now. Forty to sixty-nine percent is pilot first. Below forty is not yet.

4

Compare candidates

Run the scorecard once per use case. The highest scorer with a measurable outcome is normally the right first deployment, even if it is not the most exciting one.

Why These Seven Factors

Each factor maps to a way private LLM projects fail in practice. The scorecard exists so those failures are found on a web page rather than in month three of a build.

Source documents and data sensitivity

A retrieval system can only answer from what is written down and indexed. If the knowledge is in people's heads, the model has nothing to retrieve. Sensitivity matters separately: sensitive data can be handled in a private deployment, but only if you know who may see what before anything is indexed.

  • Documents must exist, be current and have a known owner
  • Permissions must be enforceable at document level
  • Personal information triggers Australian Privacy Principles obligations
  • Mixed sensitive and general content in one repository needs separation first

Volume and error tolerance

Volume decides whether the build is worth it. Error tolerance decides whether it is safe. A task done fifty times a day where a mistake is caught on review is close to ideal. A task done twice a month where a mistake reaches a customer is not, however impressive the demo.

  • High volume justifies the build cost through the value calculators
  • Low tolerance for error needs a review step in the workflow
  • Draft-and-review use cases tolerate error better than auto-send ones
  • Score the failure cost honestly, including reputational cost

Onshore processing

Some organisations must keep processing in Australia because of contracts, regulator expectations or policy. That is not a reason to avoid AI, it is a design constraint, and it is settled if you already have an Australian-region or on-premises option. The score drops when nobody has asked the question yet.

  • Check contract clauses on data location before choosing a platform
  • Region of inference matters as much as region of storage
  • A private deployment in an Australian region satisfies most requirements
  • Use the AI data residency quiz if you are unsure where data goes today

Measurable outcome and human review

If you cannot say what a good result looks like, you cannot evaluate the system, and you will not know whether to expand or stop. A named reviewer who sees the output before it has effect is the single cheapest control available and it makes low-tolerance use cases viable.

  • Write down the metric before the build starts
  • Collect fifty real questions with expert answers as the evaluation set
  • Name the reviewer and put review inside the workflow, not beside it
  • Log what the reviewer changes: it shows where the system is weak

Next Steps

Private LLM Readiness Assessment

Score the organisation rather than the use case: content, permissions, governance and evaluation capability.

Assess readiness

RAG vs Fine-Tuning Decision Tool

Once a use case scores well, work out which architecture fits it.

Choose an architecture

Document Search Time Cost Calculator

Put a value on the most common first use case: finding things in your own documents.

Value the use case

Frequently Asked Questions

Sources and further reading

Scored a Use Case at Build Now or Pilot First?

Send us the use case, the score and what you answered on source documents and review. We will come back with a scoped pilot or build, a price within our published range and the evaluation set we would use to judge it.