An AI training curriculum turns into a tool list quickly. People create accounts, tour the screen and copy a few example prompts. By the end of the session everyone has opened the tool. The trouble shows up the following Monday: should this task go to the AI, and how would anyone notice a wrong answer? The session never covered either.
This is a practical guide, not a report on a client's training. It covers one part of the AI education and workplace automation design DMS.Labs works on: how to order a training program, using published research. The core idea is to move the center of training from the tool to the task boundary, the line between work you can hand to AI, work you hand over but always check, and work you keep for now.
What the research shows: the effect is uneven
Evidence that generative AI helps at work has piled up. Read closely, though, it says the help varies by task and by person. Three studies are worth walking through.
The first is Noy and Zhang's experiment, published in Science in 2023. They gave 453 college-educated professionals occupation-specific writing tasks and exposed half of them to ChatGPT. Average time fell by 40% and graded quality rose by 18%. Two more findings matter for training. Workers who had used ChatGPT in the experiment were twice as likely to report using it in their real job two weeks later, and 1.6 times as likely two months later. A single hands-on experience carried into daily work. The same paper reports that weaker writers gained more, so the gap between workers narrowed.
The second is Brynjolfsson, Li and Raymond's study of 5,179 customer support agents, who got a generative AI assistant in a staggered rollout. Issues resolved per hour rose 14% on average. For novice and low-skilled workers the gain was 34%, while experienced, highly skilled workers saw minimal impact. The same tool did very different things for different people.
The third is a preregistered experiment by Dell'Acqua and colleagues with Boston Consulting Group. In it, 758 consultants were randomly assigned to no AI, GPT-4, or GPT-4 plus a short prompt engineering overview. On 18 tasks the researchers designed to sit inside the AI frontier, people using AI completed 12.2% more tasks, finished 25.1% faster on average, and delivered significantly better quality. On a complex managerial task designed to sit outside the frontier, people using AI were 19% less likely to produce a correct solution than people working without it. The authors call this the jagged technological frontier. Two tasks that look equally hard to a person can land on opposite sides of it, so AI helps on one and hurts on the other.
The settings differ. Tasks, populations and model versions are not the same, so the numbers cannot be added together or carried straight into another company's work. What carries over is the direction. How much AI helps depends on the kind of task and the skill of the person, and some areas produce wrong answers that look convincing.
A hand holding a pencil over one of several task cards laid out on a wooden table.View original
The gap that tool-first training leaves
Teaching the tool is not wasted effort. If the screen is unfamiliar, people will not even try, so basic operation has to be covered. But when usage instructions fill the whole program, participants tend to drift toward one of two extremes.
One is overconfidence. The demo ran smoothly in the session, so the same approach must work for everything. The outside-the-frontier result in the Dell'Acqua experiment, where people using AI were less likely to be correct, shows how a tool that usually works can fail without warning.
The other is giving up. After one or two strange answers, someone concludes the tool does not fit their job. Because they cannot tell which task went wrong, they stop using it for all of them.
Both come from the same gap. Neither person has ever sorted their own work into tasks that sit inside the AI's reach and tasks that sit outside it. Building that experience is the main job of the training.
Building a task boundary map
I suggest that participants produce one concrete deliverable in the session: a task boundary map. It is one table, not a formal document. These are the steps.
1. Break the job into small pieces. Do not write "reporting." Write gathering material, choosing the key sentences, drafting, checking figures and explaining the result to a manager. A job is usually made of several kinds of work, and only some of them suit AI. The boundary runs between those pieces, not around the job title.
2. Note four things for each piece. How often you do it, how costly a mistake would be, whether you can recognize the right answer on sight, and whether the input contains information that must not leave the organization. There is no need to score these. High or low is enough.
3. Sort into three columns. Work that is frequent, cheap to get wrong and easy for you to judge goes in "try handing over." Work where a mistake is costly or the result is hard to judge goes in "hand over, then check." Work with sensitive inputs, or with standards you cannot put into words, goes in "keep for now." Pieces you cannot place stay as "unsure" and move to a test list.
4. Test the unsure column yourself. This is the most valuable part of the session. Each participant picks one piece, runs it through AI with de-identified material from their own work, and compares the output against an answer they already know. They add one line to the table: where it went wrong, and how confident the AI's wording sounded when it did. As those lines accumulate, the map fits the real job more closely.
An open paper notebook with a pencil-drawn wavy line and blank slips of paper on both sides.View original
What to check during a test
Keep the test light but consistent. Record the instruction and material you used, the output you received and what a person changed. Running the same task two or three times over two days also shows how much the results wobble.
Three things are worth checking. The first is accuracy: figures, dates, names and quotations compared directly against the source, plus whether anything was invented that the material does not contain. The second is omission. Wrong statements stand out, but missing items do not, so someone who knows the answer should count whether every expected item is present. The third is review time. If AI drafts something in five minutes and a person needs thirty to verify and fix it, the task may not be worth handing over. Our guide on measuring AI adoption explains why review time belongs in the same ledger.
Record test results as the conditions under which the output was right, not as simply right or wrong. For example: accurate when the source was short and clearly structured, but swapped columns once several tables were mixed. With the conditions written down, the boundary becomes specific.
Split the training by skill level
In the Brynjolfsson study the effect varied sharply with skill, and that is a direct hint for training design. A less experienced person may gain a lot from AI but may not yet have a standard for judging whether the output is correct. A highly skilled person recognizes errors quickly, but may gain little. Run the same session the same way for both and it is risky for one group and dull for the other.
Within a single session, roles can differ. Newer team members check AI output against a senior colleague's criteria sheet. Experienced members write out their own judgment criteria in words. Writing the criteria down surfaces tacit know-how, and those criteria become raw material for instructions to the AI and for review checklists. This guide offers no evidence that the approach raises a team's results. Treat it as a design hypothesis.
Design training to run until two weeks later
Noy and Zhang did not stop at the moment of the experiment. They asked about real use two weeks and two months later. Training outcomes are the same: satisfaction on the day tells you little. You have to ask, after some time has passed, whether people are using it at work and where they stopped.
The follow-up proposed here is simple. About two weeks after training, ask participants to take out their map again. Did anything from "try handing over" actually get handed over? Were the checks kept for "hand over, then check"? Has evidence appeared that would move something out of "keep for now"? The two-week gap is my suggestion, not a standard from the research. A team whose work runs on a monthly cycle should wait a month.
The map is not a fixed document. When models change or work changes, items move between columns. A task that produced one serious error moves down a column, and one that stayed stable for a long time moves up. How to widen permissions in steps is covered in Giving AI automation permission in three steps. If the map decides what to hand over, that guide decides with what permissions.
Questions for the person planning the training
Here are some questions to ask while designing the program. Do participants get time in the session to test with their own work material? Does it include examples where the result is deliberately wrong? Are the rules for information that must not leave the organization stated up front? After the session, is there a schedule and an owner for checking in at the workplace? If any of these is empty, shorten the tool tour a little and fill that slot instead.
The aim of AI training is not to increase the number of people who know a tool. It is to increase the number of people who can decide for themselves how much of their own work to hand over and where to look directly. This guide is general advice built on published research and training design principles, not a result from any specific organization.
References
- Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.
- Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at Work. NBER Working Paper 31161.
- Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. Navigating the Jagged Technological Frontier. Organization Science. (HBS Working Paper 24-013, SSRN)