"Hire an AI employee" is a phrase that now sells a great many things, from a chat window with a persona to a workflow that breaks the first time a customer asks something unexpected. This guide is the hiring checklist: what to decide before you hire, how the first two weeks should run, what to measure, and which criteria separate an employee that finishes work from a demo that only talks. If you still need the definition, start with what a digital worker is; this guide assumes you have it.
Before you hire: three decisions on paper
Hiring goes wrong for the same reason it does with people: the role was never written down. Three decisions, each a few lines, before you open any product.
- The job, in one sentence. Input, action, output. "Reads incoming customer messages, answers the ones covered by our order records and opening hours, flags the rest." If the sentence needs a second one, you are hiring for two roles.
- The permission tier for that job. Read only (it looks and summarises), draft only (it prepares, you send), or complete with stop rules (it finishes, and stops on named conditions). Most first hires belong in draft only; the tier is raised later, job by job. We covered the logic in approval gates.
- The tools it has to reach. List the systems the job touches: your WhatsApp number, your inbox, your order system, your calendar, your spreadsheet. The employee must work inside these, not in a copy of them.
Write the three on one page. That page is the job description, and it is also the test you will hold every provider against.
The employment contract: identity, permissions, skills
An AI employee is defined by the same three things as a person you onboard, and each one is a setting you control:
- Identity. A name, a role, a tone, and the rules it never breaks (no prices, no promises about delivery dates, no medical or legal advice). Written the way you would brief a new colleague, in plain language.
- Permissions. Which systems it may read, which it may write to, and where it must stop and ask. The default in a well-built platform is closed: a connected tool opens nothing until you allow it.
- Skills. The connections to your tools, switched on one at a time. Your own channels and records, not the provider's.
If a product cannot show you these three as separate, editable things, it is not offering an employee; it is offering a script.
The first two weeks
Onboarding is where most hires either take hold or quietly die. The sequence below is the one that works.
| Days | What happens | What you do |
|---|---|---|
| 1-2 | The job description, red lines and the twenty most frequent questions from your last month go in. Tools are connected. | Write, connect, check that it can see the right records. |
| 3-7 | Draft mode. The employee finishes each job and leaves the last step to you. | Read every output. Correct mistakes by editing the job description, not by re-doing the work. |
| 8-10 | Draft mode, read fast. Note only what you had to change. | Score the job with the attention you will have in month three, not on day one. |
| 11-14 | Completion rights on the jobs that went through untouched. The rest stays in draft. | Widen permissions one job at a time. Record the first three numbers (below). |
Two rules keep this honest. First, no completion rights before a week of clean drafts. Second, a correction goes into the job description so it never has to be made twice; an employee you correct by hand every day is one you have not finished onboarding.
What to measure from week two
Three numbers, kept weekly. They are the difference between a hire you can defend and a feeling.
- Runs. How many times the employee did the job this week. Flat means the job is too rare to matter.
- Checking time. Minutes per output, on average. Over a third of the job's own duration means the output's shape is wrong; ask for a figure that has to add up, not a paragraph.
- Undone outputs. How many you corrected or reversed. Not zero, but not rising.
The savings figure follows: runs times the job's old duration, minus runs times checking time. When that figure is positive for four weeks, hire the second role. The where to start guide walks through the same numbers day by day.
Six criteria for choosing where to hire
Comparing providers by feature lists is the wrong instrument. Hold each one against the page you wrote, and check six things:
- It works inside your tools. Your WhatsApp number, your inbox, your records. If the demo happens in the provider's window with the provider's sample data, ask to see it on yours.
- Permissions are real, not a prompt. A rule written into an instruction is a request; a permission enforced at the tool is a control. Ask where "it may not send" is enforced.
- Scheduled and long-running jobs survive. Monday's report runs without you opening anything; a job interrupted halfway resumes rather than restarts.
- There is a run log. For every job: what it read, what it did, what it sent, what it stopped on. Without this you cannot check, and what you cannot check you will stop trusting.
- The price is a salary, not a meter. A fixed monthly fee per employee with a shared pool is a budget; a charge per conversation is a bill that grows with your success. The cost guide compares the models.
- A second role is a second employee, with hand-off. When the inbox employee needs the reporting employee's numbers, the platform should let them pass work between roles under the same permissions. Otherwise you will be the messenger between two bots.
Five red flags
- The demo only talks. Impressive conversation, no record touched, no message sent. That is a chat window, whatever it is called.
- "Fully autonomous from day one." Nobody serious hires that way. Draft mode and stop rules are signs of a platform built for real work.
- Per-message or per-credit pricing for routine work. It punishes exactly the behaviour you want: delegating more.
- No answer to "what did it do last Tuesday?" No log, no accountability.
- The pitch is replacing your team. The value is in the routine leaving people's hands so they can do the judgment work; a provider selling headcount reduction has misunderstood the job, and probably the risk.
After the hire
An AI employee is not finished when it is switched on; it is finished when its job description stops changing. That usually takes a month. From there, growth is a hiring decision each time: the same one-sentence test, the same permission tier, the same three numbers. Roles accumulate into a small team with clear responsibilities, which is the point, and the reason the word "hire" fits.
Frequently asked questions
How long does hiring take?
The paperwork, a page, takes an afternoon. Onboarding to completion rights on the first job takes two weeks in draft mode. A stable job description takes about a month.
Should the first hire be customer-facing?
Only in draft mode. Customer-facing jobs have the most expensive mistakes, so they earn completion rights last. A reporting or reconciliation job is the safer first hire; the inbox in draft mode is a close second.
Do I need someone technical?
No. Every step in this guide is writing a description, choosing a permission and reading outputs. The only technical knowledge needed is which system holds which data, and the person who does the job today knows that.
What if it makes a mistake with a customer?
In draft mode it cannot, because you send. After completion rights, the run log shows what happened and the stop rules narrow. That is why permissions widen one job at a time.
Can one AI employee do everything?
One employee owns one role. Adding a second role means a second employee, and the two hand work between them. A single "does everything" bot is the same mistake as the "does everything" vacancy: hard to check and quick to fail.

