AI / The TecBase journal
Everyone Wants AI. What Would You Actually Hand Over?
Treating AI like a new colleague produces a more useful question than asking which tool is the most impressive.

Start with a handover
Suppose a capable new colleague starts tomorrow. Would you tell them to ‘improve the business’ and leave them alone? Or would you give them a defined task, the information they need, an example of a good result and a way to ask for help?
That is a useful way to frame an AI project. The demonstration may be impressive, but the everyday value depends on the handover: what the system receives, what it is expected to produce and who decides whether the result is good enough to use.
The colleague comparison has a limit: an AI system does not acquire responsibility simply because it produces human-sounding language. Someone in the business still owns the task and its consequences. Thinking in terms of a handover is useful because it forces you to specify the work, not because the software should be treated as an accountable member of staff.
Choose a task you can explain
Candidate tasks might include organising incoming requests, finding relevant passages in approved material or preparing a draft from supplied notes. The first exercise is to describe one task without mentioning AI. What starts it? Which information is necessary? What should the next person receive?
If the description is still vague, a model is unlikely to make the workflow clearer by itself. You may first need consistent records or an agreed process. And if the task is an exact calculation or a fixed rule, ordinary software may be the more straightforward choice.
Use four questions to find a sensible first task
Ask how often the task occurs, how clearly a good result can be described, how costly a mistake would be and how easy the result is to check. A frequent task with a recognisable answer and a manageable review process is a more useful candidate to investigate than a rare decision that requires extensive judgement.
For example, a team might trial suggested labels for incoming support messages. Staff can compare a suggestion with the message, correct it and keep the final routing decision. That is a different starting point from letting the same system promise refunds or alter customer accounts. The output may look equally simple on screen, but the authority involved is not.
Write the task as a sentence with a boundary: given these approved inputs, prepare this result for this person to review; if this information is missing, stop and ask. If that sentence becomes a page of exceptions, narrow the task or address the underlying process first.
Decide what a mistake would cost
A rough internal summary and a customer-facing commitment deserve different treatment. Think about who sees the result, whether a mistake is reversible and how quickly someone would notice it. Those questions help define the level of review the workflow needs.
NIST’s generative-AI profile identifies the risk of confidently stated false content. A polished answer therefore needs more than a polished interface. For a task that depends on source material, make that material available to the reviewer. Decide when the system should flag uncertainty, request missing information or stop for a person’s decision.
Also distinguish suggesting an action from being allowed to perform it. Preparing a response does not require permission to send it. Reading a record does not require permission to change it. Give a first trial the smallest set of capabilities needed to test its value, then make each expansion in authority a deliberate decision.
Reference: NIST: Generative AI Profile
Include the time spent checking
A draft produced in seconds is not necessarily a task completed in seconds. Someone may need to verify details, correct the format and move the result into the place where work actually happens. Count that effort when comparing a trial with the existing process.
Use a small, representative set of examples, including incomplete requests and awkward cases. Compare usefulness, corrections and total completion time. Agree on who can access the information used in the trial and avoid sharing data that the task does not need. The goal is to learn whether the whole workflow improves.
Here is a hypothetical calculation. One hundred tasks at eight minutes each take 800 minutes. If an assisted version needs one minute of preparation and three minutes of review per task, that becomes 400 minutes. Add sixty minutes for corrections and exceptions, and the total is 460 minutes: a saving of 340 minutes, before ongoing maintenance or other costs. The point is to count the whole task, not to predict that another business will get the same result.
Build a test set that is allowed to disappoint you
Choose examples before deciding whether the trial is successful. Include ordinary requests, missing details, conflicting information and cases that should be passed to a person. Remove unnecessary personal or confidential information, and use approved access arrangements. Keep some examples aside until you are ready to evaluate a change, rather than repeatedly adjusting the system to the same demonstration.
For each result, record whether it is usable as supplied, needs a minor correction, contains a material error or should have stopped. Define those categories for the task. A misspelled internal label and an incorrect customer commitment should not disappear into the same average score.
Ask someone who regularly performs the work to review the results. Where possible, have them assess the output before hearing how it was produced. The question is whether the result meets the standard of the job. A convincing explanation from the person running the demonstration is not a substitute for that standard.
Give the person reviewing it a real decision
A review step only helps if the person has the information, time and authority to change the outcome. Put the relevant source material beside the draft, make meaningful changes visible and let the reviewer reject or send it back. Avoid an interface that presents approval as the only convenient action.
Watch what happens during the trial. If people spend longer searching for the evidence than checking the draft, improve the review experience. If they approve every result because the queue is too long, the workflow may be undermining the safeguard it claims to provide.
Make it possible to continue without the AI step when necessary. Decide who handles failed runs, how a task returns to the ordinary queue and where corrections are recorded. A dependable process includes a workable fallback, not just a successful path.
Connect the useful part to everyday work
If the trial works, the next challenge is often practical. How does the approved input reach the system? Where does the output go? Who handles a failed run? Can a person pick up the task with the relevant context intact? Those details determine whether a promising experiment becomes a dependable tool.
A good first project does not need to be ambitious in every direction. It needs a clear task, a sensible boundary and an observable improvement. Before asking what AI could do for your business, choose one piece of work you would feel comfortable handing over with the right instructions and review.
Give the trial a named owner and a review point. That person should be able to explain what is being measured, what would cause the team to stop and which decision comes next. A useful outcome may be to expand the scope, improve the inputs or abandon the idea. Learning that a simpler solution is better is still a successful use of a small experiment.


