Imagine running a business with 150 employees. Finance wants fewer hours spent on invoices. HR wants fewer repetitive questions. Your developers want an AI assistant that can build and fix software. Each proposal sounds reasonable. The difficult decision is how much responsibility to give the technology—and how to know whether it is actually helping.
I would start with the work itself: where does time disappear, what does a mistake cost, and who can put things right? Those questions make it easier to separate a useful implementation from an impressive demonstration.
Here are four cases worth examining. I use medium-sized in a practical sense: tens to a few hundred employees, with small specialist teams—not a formal SME classification. The two implementation examples fit that scale. The two failure cases come from different settings, identified below, and illustrate risks a medium-sized business can learn from. They are examples of useful and harmful practices, not a ranking of entire companies.
Harbor Compliance: give routine questions a reliable home
ClearFeed’s Harbor Compliance case study describes two People Ops staff supporting over 200 employees. AI answers from an existing Google Drive knowledge base were combined with ticket routing and tracking inside Slack. The vendor reports a 40% reduction in resolution time. This is a reported result for the combined workflow; the study does not isolate AI’s contribution or provide an independent evaluation.
What interests me is the modest scope. An internal assistant can help people find an approved policy without deciding an employee’s entitlement. I would keep that boundary explicit: a question about where to find the leave policy is different from a disputed leave request.
For a similar rollout, I would give each policy an owner and a review date, check that access permissions carry through to answers, and make sensitive cases easy to hand to a person. Faster responses are useful only if employees can trust the answer and reach someone when it does not fit their situation.
Nordell: a promising design, with benefits still to prove
Made Smarter’s July 2026 Nordell case study describes a plastics manufacturer with over 80 employees. Its AI invoice-processing investment connects to the existing ERP system, checks invoices against records, and flags discrepancies for review. Leadership and staff training accompany the project. The projected saving exceeds 60 administrative hours a month. That is an expectation, not a published measurement of realised savings.
This is the kind of project I would take seriously: a recurring task, records to compare against, and an exception process. I would still test whether simpler rules or conventional automation could do the job at lower cost. Buying AI should follow from the problem, rather than become the objective.
Thought experiment 1 · Hypothetical
The hours that never reach the bottom line
Suppose an invoice assistant saves 60 hours a month, but checking and correcting its work takes 15. At an assumed fully loaded labour cost of NOK 600 an hour, 45 net hours represent NOK 27,000 of capacity. Subtract NOK 8,000 in monthly software and operating costs, and the potential value is NOK 19,000 before implementation costs.
That is not automatically NOK 19,000 in cash savings. If payroll stays the same, the benefit depends on what people do with the time. Can they clear a backlog, reduce overtime, or take on more work? I would agree on that use before approving the project. These figures are illustrative assumptions, not Nordell’s economics.
Replit and SaaStr: the cost of excessive access
In a July 2025 incident account, Replit acknowledged that its agent deleted data from an application database being built by SaaStr co-founder Jason Lemkin. Replit says the database was fully restored. It subsequently introduced separate development and production databases by default. This was a founder’s application-building project; the account does not establish a medium-sized company rollout or an organisation-wide data loss.
My lesson is about authority. Instructions in a conversation should sit alongside technical permissions that limit what an agent can actually change. A system that proposes an update needs less access than one that can apply it to live customer records.
I would separate experimentation from production, restrict credentials, and require approval for consequential changes. I would also test restoration before depending on it. A backup is valuable when someone can find it, restore the right version, and verify that the business can continue.
Thought experiment 2 · Hypothetical
The same agent, two sets of keys
Give two identical agents the task of cleaning up duplicate customer records. One works on a copy and produces a proposed change list. The other can delete records in the live CRM. Both misunderstand the same instruction.
The first creates work for a reviewer. The second may interrupt sales, invoicing, or customer service. The difference is the access we chose to grant. Before expanding an agent’s role, I would ask: what is the largest mistake this permission allows, and can we contain it?
Air Canada: an answer is part of the customer experience
Air Canada is a large-company counterexample. In Moffatt v. Air Canada, 2024 BCCRT 149, a Canadian tribunal found negligent misrepresentation after a customer relied on incorrect chatbot advice about bereavement fares. A link to a contradictory policy page did not resolve the problem. The decision does not establish what technology powered the chatbot, so it should not be presented as a proven generative-AI failure.
The management question still applies to an AI service: who owns the promises made through your website? I would test answers about refunds, prices, eligibility, and deadlines especially carefully. A citation is useful, but someone must check that the answer actually agrees with it.
Thought experiment 3 · Hypothetical
A confident answer to an ambiguous policy
Your sales assistant finds two documents: an old discount policy and a current one. A customer asks whether a large order qualifies. Would you rather have an immediate, confident answer, or a short explanation that the case needs checking?
I would make escalation a successful outcome when the evidence conflicts. Then I would measure how quickly a person resolves it. If the dashboard rewards only conversations closed without human help, the team may optimise away the very review that protects the customer.
What I would ask a management team to show me
These cases have different evidence behind them: a vendor-reported improvement, projected savings, a provider’s incident account, and a tribunal finding. They do not establish a universal return on AI. They do help frame a practical pilot.
- One task and a baseline. Measure the current completion time, error rate, and rework before introducing the assistant.
- People who know the exceptions. Involve the staff doing the work. Ask them for the awkward cases a polished demo leaves out.
- A defined data boundary. Decide which records the tool needs, who can see them, and what the provider may retain or use.
- Separate permission to suggest and to act. Document who can approve payments, change records, or make commitments to customers.
- A fair evaluation. Test normal and difficult cases against the existing process. Count review time, unresolved requests, errors, and operating costs—not just adoption.
- A stop rule and an owner. Agree on conditions for pausing, the person responsible, and a tested way back to normal operations.
I would start with a narrow, reversible pilot and expand it when the evidence supports doing so. The interesting milestone is a team completing useful work more reliably, with a clear understanding of where the system still needs help. That is the result I would want to bring to a board meeting.
Sources and scope
Reviewed on 11 October 2026. The interpretations and hypothetical scenarios are my analysis of the public material.
- ClearFeed: Harbor Compliance customer case study — vendor account; publication date not stated.
- Made Smarter: Nordell — programme case study, 3 July 2026; forecast benefits.
- Replit: Doubling down on our commitment to secure vibe coding — provider account, 29 July 2025.
- Moffatt v. Air Canada, 2024 BCCRT 149 — Civil Resolution Tribunal, 14 February 2024; see paragraphs 14–17 and 24–32.
