Choose one piece of work
Pick a bounded task with a known input and a useful output. For example, preparing a draft summary for a person to review is easier to evaluate than asking an assistant to manage an entire customer relationship. Be explicit about who will use the result.
Record the current process
Understand how the team completes the task today, where delays occur and how mistakes are found. This baseline helps you compare the pilot with the work it is intended to improve. Faster output is only useful if the result remains suitable for its purpose.
Build a representative test set
Use examples that reflect the variety of the task, including incomplete, ambiguous and unusual inputs. Decide how a reviewer will judge each output. Keep some examples aside so that changes to the pilot can be assessed against cases it was not tuned around.
Set data and action boundaries
Agree which information the pilot may access and which actions require review. Start with the minimum access needed. If the pilot drafts an email, generating the draft and sending it should be treated as separate decisions.
Make failure visible
Give people a way to recognise an uncertain result, correct it and continue without the automation. Record the types of error that matter to the task. A system that produces a plausible answer to every input can be difficult to supervise.
Decide what earns a wider rollout
Set a review point and an owner. Look at quality, time saved, operating cost and the effort required to maintain the workflow. A useful pilot can lead to a production build, a simpler automation or a decision to keep the process as it is.
