Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Sunday, July 5, 2026

Pilots Succeed and Programmes Still Fail

Pilots Succeed and Programmes Still Fail. A pilot proves feasibility under favourable conditions.

AI pilots fail at deployment rather than at proof. A pilot is built to remove the conditions that make production difficult, including messy data, absent ownership, undefined governance and ongoing maintenance. Success under those removed conditions predicts very little. The programme then fails on precisely what the pilot was designed to exclude.

A Pilot Is a Controlled Exclusion

The purpose of a pilot is to answer one narrow question. Can this technique produce a useful result on this kind of problem at all.

Answering that question quickly requires removing everything that would slow it down. Scope is narrowed, data is cleaned by hand, edge cases are set aside, and the best available people are assigned.

All of that is correct practice. A pilot loaded with production constraints would take as long as production and would prove nothing sooner.

The error is in the interpretation, not the design. A successful pilot is read as evidence that the capability works, when it is evidence that the capability can work under conditions nobody intends to maintain.

Feasibility and viability are separate questions. The pilot answers the first one honestly and says nothing at all about the second.

Programmes stall because the second question is never asked as a question. It is assumed to have been answered by the first.

The pattern repeats across technologies. Automation projects, analytics platforms and workflow systems have all produced convincing demonstrations that did not survive the move into daily operation.

What is different now is the speed of the demonstration. A capable result can be produced in days, which shortens the gap between an impressive demonstration and a funding decision.

What the Pilot Quietly Removed

The exclusions are worth naming individually, because each one becomes a failure mode later.

Data condition is the first. Pilot data is usually selected, cleaned and often corrected by hand, while production data arrives incomplete, inconsistent and formatted by whoever entered it.

Staffing is the second. Pilots are staffed by capable and interested people who volunteered, and production is staffed by whoever holds the role that week.

Attention is the third. A pilot runs with a sponsor watching, which resolves blockers in hours that would otherwise take weeks through normal channels.

Scope is the fourth. Pilots select the cases where the technique is likely to work and exclude the difficult remainder, which is often the majority of real volume.

Time is the fifth and least visible. A pilot runs for a fixed short period, during which nothing in the surrounding business changes enough to break it.

Each exclusion is defensible on its own. Together they describe an environment that has almost nothing in common with the one the tool will actually live in.

Integration is the sixth exclusion and often the largest. A pilot can accept input by hand and produce output into a spreadsheet, while production has to read from and write to systems that were not designed for it.

Deployment Tests Ownership Before It Tests Accuracy

The first thing production asks is who owns this. That question has no answer in most pilots, because the pilot team owned it by default.

Ownership means several distinct things at once. Someone decides when the output is wrong, someone fixes it, someone approves changes, and someone answers when a customer is affected.

In a pilot these roles collapse into the project team. In production they belong to different functions with different priorities and different reporting lines.

The gap shows up as delay rather than as failure. Output degrades, several people notice, and nobody has the standing or the time to act on it.

Unowned systems do not stop working suddenly. They drift, produce results that staff learn to distrust, and are quietly worked around until they are abandoned.

A clear structured route from a working pilot to a supported deployment exists precisely to force these assignments before the tool reaches live use.

Naming an owner is cheap at the start and expensive later. After launch, the assignment becomes a negotiation between functions that have already absorbed other commitments.

A useful test is to ask who would be interrupted if the output were wrong on a Friday afternoon. If the answer is the project team, the deployment is not ready.

The second test is to ask who has authority to switch it off. Systems without a named person able to suspend them tend to keep running through problems that should have stopped them.

Governance Is the Second Test

Governance sounds like a compliance concern and is actually an operating one. It determines what happens in the cases the pilot never encountered.

Every production system generates situations outside its tested range. A customer request phrased unusually, a document in an unexpected format, a case that sits between two categories.

Governance is the set of standing answers to those situations. What the tool is permitted to decide alone, what requires review, and what must be escalated to a person.

Without those answers, staff invent their own. Some accept the output uncritically, others check everything manually, and the variation between them becomes invisible until something goes wrong.

The absence is hardest to see when the tool performs well. Consistent good output removes the pressure to define boundaries, so the boundaries get defined during the first incident instead.

Organisations that move from experiment to routine use tend to follow a deliberate path from isolated pilots to standing governance rather than treating each deployment as a separate exercise.

Governance also settles the record-keeping question. What was decided, on what input, and by which version of the system, all of which become urgent the first time an outcome is disputed.

Staff behaviour is the other governance variable. People adjust their own work around the tool, and those adjustments become the real process regardless of what the documentation says.

Maintenance Is the Test Nobody Budgets

A pilot has a defined end date. A production system has an indefinite one, and that difference changes the economics completely.

Inputs shift over time. Customers phrase requests differently, suppliers change formats, regulations move, and the business itself changes what it sells.

Performance degrades gradually against all of that. The decline is slow enough that no single week justifies attention and cumulative enough to matter within a year.

Maintenance means monitoring output quality, refreshing reference material, adjusting prompts or rules, and retiring behaviour that no longer fits. None of that appears in a pilot budget.

Vendors rarely raise it either, because the maintenance burden falls on the buyer and discussing it complicates the purchase. The subscription is quoted. The standing internal effort is not.

The practical consequence is a slow reversal. Staff notice output quality falling, revert to the previous method for the important cases, and the tool survives only for work that does not matter much.

Version changes add a further layer. Underlying models and platforms are updated by vendors on their own schedule, and behaviour can shift without anything changing on the buyer side.

That makes periodic revalidation part of the standing cost. A short recurring check against a fixed set of known cases catches most of this before users do.

Designing a Pilot That Predicts Deployment

A pilot can be made more predictive without losing its speed. The adjustments are small and they change what the result means.

Use unedited data for part of the test. Cleaned data proves the technique works and dirty data proves whether the workflow around it survives contact with reality.

Include the difficult cases deliberately. A pilot that handles only the straightforward majority hides the exact volume of exception work that production will generate.

Assign the eventual owner during the pilot rather than after it. That person surfaces objections early, and those objections are the cheapest information the pilot produces.

Run at least part of the pilot without sponsor attention. How the process behaves when nobody senior is watching is the closest available preview of normal operation.

Write the governance rules before launch rather than during the first incident. Three or four written boundaries covering permitted decisions, required review and escalation are usually enough to start.

Set the success criterion at the process level, not the model level. Accuracy on a test set is a technical result, while cycle time and exception volume are business results.

The uncomfortable implication is that a pilot which fails under these conditions has done its job. It has produced the information the organisation needed at the cheapest possible moment. A pilot that succeeds under favourable conditions and is then scaled has produced confidence instead of information, and confidence is what makes the programme expensive. The question worth asking after any successful pilot is not whether the technology worked. It is which of the removed conditions the business is now prepared to put back.

None of these adjustments require a longer pilot. They require the pilot to answer a slightly different question, which costs planning attention rather than calendar time.

The choice is between paying for that information early or paying for it after the programme has been announced internally. The second option carries reputational cost alongside the financial one.

Frequently Asked Questions

What is the difference between a pilot failing and a deployment failing?
A pilot fails when the technique cannot produce a useful result, which is a technical outcome and usually cheap. A deployment fails when the technique works but the organisation cannot sustain it, which is an operating outcome and usually expensive. The second failure happens after budget, staffing and expectations have all been committed. Most programmes that stall have already passed the technical test.

Should a pilot use real production data?
At least part of it should. Cleaned data answers whether the technique is capable, while unedited data answers whether the surrounding workflow can cope with what arrives in practice. Running both gives two different and equally useful results. Using only cleaned data produces a pilot that cannot fail in the way that matters.

Who should own an AI tool once it is live?
Ownership belongs with the function that owns the underlying process rather than with the technology team. The technology team maintains the system, but the process owner decides what acceptable output looks like and when it has stopped being acceptable. Splitting these roles without naming both leaves the monitoring gap that most deployments fall into. The assignment should be made before launch.

How much maintenance does a deployed tool actually require?
Enough that it needs a named person and recurring time rather than occasional attention. The work includes checking output quality, updating reference material, adjusting rules as the business changes and retiring behaviour that no longer fits. It is ongoing rather than periodic, because the inputs shift continuously. Treating it as a one-off cost is the most common budgeting error.

What governance rules matter most at the start?
Three questions cover most of the ground at launch. What the tool may decide without review, what requires a person to check before it takes effect, and what must be escalated regardless of confidence. Writing those answers down before launch prevents each staff member from inventing a private version. The rules can be refined once real cases start arriving.

Is it worth running a pilot at all if the conditions are artificial?
Yes, provided the result is read correctly. A pilot is a cheap answer to a narrow question and remains the right first step for most organisations. The failure lies in treating that narrow answer as permission to scale rather than as one input among several. Adding a few production conditions to the pilot preserves the speed while improving what the result actually tells anyone.

Friday, June 26, 2026

Shadow AI Is a Governance Failure, Not a Tooling One

Shadow AI is a governance gap. Employees adopt AI faster than policy arrives. That gap cannot be purchased shut.

AI governance is the set of decisions about who may use which tools, on what data, with what review, and who answers when something goes wrong. It is not software and it is not a document. Companies that treat it as a purchase discover that unapproved use continues quietly, because the underlying questions were never answered.

Unapproved Use Is a Signal, Not a Violation

Shadow AI describes the ordinary situation inside most companies right now. Employees use assistants on personal accounts, on personal devices, and through browser extensions nobody approved. The behavior rarely reflects defiance of any kind. People are solving a problem faster than the organization can decide how they should solve it.

The distinction matters, because the response follows directly from the diagnosis. Treating unapproved use as a discipline problem produces a memo, a prohibition, and more careful concealment. Treating it as a signal produces a better question about what work became painful enough to route around the company.

The pattern of staff adopting assistants faster than any oversight can arrive concentrates in predictable places. The list usually includes repetitive writing, summarizing long documents, drafting client communication, and cleaning up messy data. Those are exactly the tasks where the gap between available tooling and daily demand runs widest.

A prohibition does not remove the demand that created the behavior in the first place. It removes visibility, which is the one asset the company still had. The work continues on personal accounts where no logging exists, no retention rules apply, and no review is possible.

Visibility carries practical value that extends well beyond risk reduction. Knowing which tasks employees hand to an assistant is a free map of where internal process is weakest. Companies that suppress the behavior lose that map and keep the underlying inefficiency.

Discovery is straightforward once the goal is understanding rather than punishment. A single direct question, asked without a consequence attached, usually produces a longer list than any audit tool returns. The quality of that answer depends entirely on what employees expect to happen next.

Governance Is a Set of Decisions, Not a Document

Governance gets confused with documentation because documentation is the part that becomes visible. The actual work is deciding which tools are permitted, what data may enter them, what output needs review, and who answers for failures. Everything else is formatting around those four answers.

Each of those decisions carries an owner and a cost. Naming permitted tools costs money and forces an honest comparison between options. Deciding what data may enter a model requires knowing what data the company actually holds. Skipping that inventory is why so many policies end up written in generalities nobody can apply.

The four decisions interact more than they appear to at first. A permissive tool list demands stricter data rules, and a strict data rule makes a longer tool list harmless. Deciding them separately produces contradictions that employees notice immediately. Deciding them together produces a policy that holds up under pressure.

Practical oversight built for a company still growing rather than one with a compliance department stays deliberately small. A handful of decisions, written plainly and revisited each quarter, will cover most of the real exposure. A framework designed for a regulated enterprise collapses under its own weight where one person covers finance and operations together.

The written policy still matters as the record of what was decided. Its job is to answer the question an employee has at the actual moment of use. A useful usage policy written for a smaller company rather than a legal department fits on a single page and names concrete examples. Anything longer gets skimmed once and then ignored permanently.

Cost belongs in the conversation from the beginning. Approved tools carry subscription costs that unapproved personal accounts had hidden inside individual behavior. Governance converts an invisible expense into a visible one, which feels uncomfortable and is nonetheless correct. A budget line is far easier to manage than an unknown.

Ownership of the policy matters as much as its contents. An unowned policy ages into a document that describes tools the company no longer uses. Someone has to hold the pen, watch what changes, and carry the authority to update it without convening a committee.

Policy Fails When It Raises the Cost of Thinking

Most AI policies fail in exactly the same way. They describe prohibited behavior in abstract categories and leave each employee to classify their own situation. Classification is work, and it lands at the precise moment somebody is trying to finish something else.

An employee facing an ambiguous rule has three realistic options. Ask somebody and wait, guess and hope for the best, or avoid the tool entirely. Two of those outcomes damage the company and the third damages the employee.

This is where the steady erosion of judgment that follows from making too many small calls in a day quietly undermines governance. Policies demanding constant interpretation consume the same attention the actual work requires. People stop interpreting and start defaulting, and the default is always whatever is fastest.

The remedy moves the decision from the employee back to the policy. Name specific tools rather than abstract categories, and name specific data types rather than sensitivity tiers. Specificity costs the author time and saves every reader time, which is the correct direction for that trade.

Training helps only when it teaches judgment rather than rules. Employees who understand why a data category is sensitive will handle unlisted cases sensibly. Employees who memorized a list will freeze the moment reality falls outside it. Short examples of good and poor use teach more than an hour spent reading policy.

Escalation deserves the same treatment as everything else in the policy. Employees need to know exactly who to ask and how quickly an answer will arrive. Skill in raising an issue to a busy executive in a form that produces a decision is not evenly distributed across a team. When the escalation path stays vague, confident people improvise and cautious people stall.

The Real Subject Is Decision Quality

Governance conversations drift toward data risk because data risk is easy to name. The larger exposure is quieter and sits inside the decisions that generated output influences. An assistant producing a confident summary of a market will shape a plan whether or not anyone verified the summary.

Companies with an existing habit of testing claims against evidence before acting on them absorb AI output far more safely. The assistant becomes one more source that has to survive normal scrutiny. Where no such habit exists, generated confidence passes straight into strategy without meeting any friction.

The same logic applies to how decisions are structured. A defined sequence for moving a choice from framing through commitment gives AI output a specific place to sit. It becomes an input during analysis rather than an answer at the conclusion. That placement is governance in a far more meaningful sense than any acceptable use clause.

Attribution is the quiet piece that most policies omit entirely. Once generated material enters a document, nobody later remembers which passages a person wrote and which arrived from a model. That ambiguity matters most when the document is challenged by somebody outside. A light convention for marking drafted material preserves the ability to check.

Verification has to stay proportional or it will be abandoned within weeks. Asking for a source check on every generated sentence guarantees that nobody checks anything at all. Asking for verification of the specific claims a decision rests on is achievable, and it catches what actually matters.

Review requirements should follow consequence rather than tool. Output that reaches a customer, touches money, or enters a contract needs a human name attached to it. Output that speeds up an internal draft needs almost nothing at all. Applying identical review to both trains people to treat review as theater.

Staying Current Without Chasing Every Announcement

One reason governance lags is that the ground underneath it keeps moving. New model versions, new features inside existing tools, and new default settings arrive without warning. A policy written against a specific feature set expires quietly, and usually nobody notices for months.

Constant monitoring is not the answer, because no smaller company can afford that attention. A modest habit of reading a regular scan of what is shifting for smaller companies keeps each review grounded in what changed rather than what feels urgent. Scheduled review paired with a light reading habit beats continuous anxiety.

Vendors change terms as often as they change features. Data handling commitments, retention windows, and training defaults all shift without any formal announcement. Reviewing those settings on the same schedule as the policy keeps assumptions and reality aligned. Assumptions made at signup rarely survive a year without examination.

Governance also has to survive turnover and growth. Decisions recorded only in the memory of whoever made them evaporate when that person changes roles. Writing them down is not the governance itself, but it is what lets the governance outlive the moment that produced it.

The uncomfortable part of shadow AI is that it delivers accurate feedback. Employees identified real friction and resolved it without permission, because permission was never on offer. A company that answers with prohibition buys silence and keeps every unit of the risk. A company that answers the open questions gets the productivity and the oversight together, and the tools stop being the interesting part of the conversation.

Frequently Asked Questions

What does AI governance mean for a company without a compliance function?
It means a short list of decisions that somebody owns and revisits on a schedule. Those decisions cover permitted tools, permitted data, required review, and accountability when something fails. Nothing about that requires a compliance department or a dedicated platform. The scale of the framework should match the scale of the company using it.

Is banning AI tools a reasonable response to unapproved use?
Prohibition moves activity out of sight without reducing the demand that created it. Employees continue on personal accounts where the company has no logging, no retention control, and no ability to review output. The practical effect is higher exposure combined with lower awareness. A narrow set of approved tools with clear boundaries performs better than a broad ban.

What actually belongs in an AI usage policy?
Named tools, named data categories, a rule about what output requires human review, and a named person to ask. Concrete examples do more work than abstract principles, because employees classify situations poorly under time pressure. The document should be short enough to read completely before a first use. Anything that requires interpretation will be interpreted in whichever direction is fastest.

Who should own AI governance inside a growing company?
An operating leader with authority across functions is usually the right holder. Handing it to technology alone produces rules about systems rather than rules about work. Handing it to legal alone produces caution that employees route around. The owner needs enough authority to approve tools and enough proximity to the work to know where assistants are genuinely useful.

How can a company discover which tools staff already use?
Asking directly works better than most people expect, provided the question arrives without a threat attached. Framing the request as an effort to approve useful tools produces far more honest answers than an audit does. Browser and expense records fill in the remainder of the picture. The goal is an accurate map rather than a list of names to discipline.

How often should an AI policy be reviewed?
Each quarter is a reasonable default for most companies, with an unscheduled review whenever a major tool changes its defaults. The review should examine what employees are actually doing rather than only what the document says. Policies drift out of date faster than most other internal documents. A short review held reliably beats a thorough review that never gets scheduled.

Sunday, June 21, 2026

The Tools Arrived Before the Rules Did

The Tools Arrived Before the Rules Did. Employees adopt capable tools faster than any organisation writes policy.

AI governance covers what staff may put into these tools, what must be checked before the output is used, and when the use has to be disclosed. Most organisations are writing those rules after adoption has already happened, which changes the task from prevention to correction.

The Gap Between Adoption and Policy Is Structural

A capable tool reaches a working professional through a colleague, a social feed, or a free tier that requires no approval from anyone. Trying it costs a few minutes and produces a visible result on the same afternoon.

Writing a policy about that tool involves legal review, a data protection assessment, a discussion about which functions are affected, and a decision about enforcement. Those steps take weeks at best in a business with the appetite to attempt them.

The mismatch is not a failure of diligence by anyone involved. Individual adoption runs on curiosity and immediate benefit, while organisational rulemaking runs on consensus and risk assessment, and the two operate at incompatible speeds.

Assuming the gap can be eliminated leads to the wrong programme of work. The realistic objective is a narrow gap with visibility into what is happening inside it, rather than a closed gap that has never existed anywhere.

The gap also reopens with every capable release. A policy written about one category of tool becomes partially obsolete when the same vendor adds a feature that changes what data the tool touches.

Governance therefore has to be designed as something that gets revised, not as a document that gets finished. Businesses treating it as a one-time drafting exercise find their rules describing a landscape that no longer exists.

Speed of change is not the only reason the gap persists. Staff adopting a tool are answering a question about their own work, while policy writers are answering a question about the whole organisation.

Procurement Cannot Solve a Capability Problem

The instinctive response is to control the tools through purchasing and network restrictions. Approved platforms are selected, accounts are provisioned, and everything else is blocked at the firewall.

That approach worked reasonably well for software that had to be installed on a company machine. It works poorly for capability that is available through any browser and on every personal phone in the building.

Blocking a domain removes it from the corporate network and leaves it fully available on the device in every pocket. Staff who found the tool useful will continue using it and will stop mentioning that they do.

The blocking approach therefore converts visible use into invisible use. Nothing about the underlying risk changes, and the business loses its only source of information about where the risk sits.

Restriction also carries a productivity charge that rarely gets counted. Staff who were working faster with a tool return to slower methods, and the people most affected are usually the most capable ones.

Understanding the shape of unsanctioned tool use inside organisations that never approved it tends to change the response from restriction toward disclosure. Knowing what is being used is worth more than a rule that pushes the same activity out of sight.

Approved provisioning still matters and does a different job. Paying for business accounts gives the organisation terms it can rely on and a place to send staff who want a legitimate route.

What the Gap Actually Exposes

The first exposure is data leaving the business. Staff paste customer records, supplier terms, draft contracts and internal figures into services whose retention terms nobody in the business has read.

Free consumer tiers deserve particular attention in this area. Terms for consumer accounts frequently differ from the business equivalents, and the difference usually concerns whether submitted content is retained or used for training.

The second exposure is output that reaches a customer without review. Generated text is fluent and confident regardless of accuracy, which removes the usual signals that a draft needs checking.

Errors that would have been caught in a rough draft pass through a polished one. Reviewers read for tone and structure, find both acceptable, and never test the underlying claims.

The third exposure is contractual rather than technical. Client agreements and supplier terms increasingly contain clauses about automated processing, and staff using these tools have no visibility into which agreements say what.

The fourth exposure concerns the provenance of finished work. When a piece of work is later questioned, nobody can establish how it was produced, which turns a routine query into an investigation.

A fifth exposure sits in undeclared dependency on one person. Work quietly reorganises around a tool that a single employee pays for personally, and the capability leaves the business when they do.

Rules That Can Be Written Within the Month

An adequate first policy is short enough to be read in one sitting. Long documents produce compliance theatre, since nobody consults a policy they cannot remember the shape of.

The first rule concerns what may go into the tools. A plain statement of what may never be entered into an external tool, naming customer identifiers, credentials, unpublished financials and anything covered by a confidentiality obligation.

The second rule concerns review of what comes out. Any output reaching a customer, a regulator or a decision maker has to be checked by a named person against a source. The check itself has to be recorded somewhere.

The third rule concerns which routes are approved. Naming the tools the business has provisioned, and stating that anything else requires a short conversation rather than a formal request, keeps the disclosure barrier low enough to be used.

The fourth rule concerns what clients are told about it. Deciding in advance what will be said if a customer asks whether these tools were involved prevents an improvised answer under pressure.

The fifth rule concerns ownership of the policy itself. Somebody has to be responsible for reviewing the policy on a stated cycle, because a rule set with no owner ages into irrelevance without anyone noticing.

Practical evaluation of which of these tools a smaller business should actually be running belongs alongside the rules rather than after them. Approving a small number of specific tools gives staff a legitimate route and makes the input rules concrete.

Each of these exposures is manageable once it is known about. What makes them dangerous is that all of them are invisible until something goes publicly wrong.

Amnesty Produces Better Information Than Enforcement

Businesses that discover widespread unapproved use face a choice about how to respond. Punishment is available and destroys the visibility that made the discovery possible.

A stated amnesty produces a far better result. Asking staff to declare what they have been using, with an explicit commitment that nobody will face consequences for past use, produces a map of actual practice within days.

The map is usually surprising in useful ways. Adoption tends to cluster in functions nobody expected, and the tools in heaviest use are often not the ones the business was worried about.

That information changes what the rules need to cover. Policy written against a real inventory addresses situations that exist, while policy written against imagined risk addresses situations that do not.

The declaration also identifies the informal experts inside the business. Staff who adopted early usually understand the failure modes better than anyone in management, and they make credible advocates for the rules that follow.

Repeating the exercise periodically keeps the map current. A short standing question in an existing management meeting is sufficient, and it costs less than any monitoring system.

Governance as a Habit Rather Than a Document

The written policy is the smallest part of the work. What determines whether governance holds is a set of recurring behaviours that keep the rules connected to what people are actually doing.

The first behaviour is asking about it routinely. Managers who include tool use in ordinary conversations about how work was produced normalise the topic and remove the sense that admitting to it invites trouble.

The second behaviour is reviewing the rules on a schedule. Someone reads the policy against the current tool inventory at a stated interval, and the review takes an hour rather than a project.

The third behaviour is deciding openly and quickly. When staff request a new tool, answering within days, with reasons, teaches everyone that the approved route is faster than the unapproved one.

Speed of response is the mechanism that keeps the whole system honest. A request that sits unanswered for a month trains the requester to stop asking, and one silent refusal undoes a great deal of written policy.

The fourth behaviour is correcting people without punishing them. Where a rule was broken, the useful response separates the person from the process and asks why the approved route was harder than the alternative.

The plain fact about this subject is that the organisation was never in control of the sequence. Tools capable enough to change how work is done arrived in the hands of individuals first. No amount of policy discipline could have reversed that order. What remains available is the choice between governing a practice that is visible and pretending to govern one that is not. Businesses that accept the sequence and work with it end up with usable rules. Those that insist on the sequence they wanted end up with a document and no visibility.

Frequently Asked Questions

Where should a business start if it has no policy at all?
The right starting point is an inventory rather than a document. Asking each function what tools are currently in use, under an explicit amnesty, produces the information that any sensible policy has to be built on. Writing rules before knowing the actual practice guarantees a mismatch between what the policy addresses and what staff are doing. The inventory usually takes days and the first policy can follow within the same month.

Is blocking these tools ever the right answer?
Blocking makes sense for specific tools with terms that are genuinely incompatible with the obligations the business carries. It fails as a general strategy because the capability remains available on personal devices that the business does not control. A blanket block converts a manageable visible problem into an unmanageable invisible one. Selective restriction paired with an approved alternative works considerably better than restriction alone.

Who should own this inside a smaller business?
Someone senior enough to make decisions and close enough to the work to know what is being produced. Placing it entirely with a technical function tends to produce rules about systems rather than about practice. Placing it entirely with a legal or compliance adviser tends to produce rules nobody can follow. A named operational owner, with access to both perspectives, is the arrangement that survives contact with daily work.

How detailed does a first policy need to be?
Short enough that staff can recall its main provisions without looking. A page covering inputs, review obligations, approved tools, client disclosure and ownership is enough to manage the material risks. Detail can be added once the business understands where its actual exposure sits, which becomes apparent within a few months of the policy existing. Starting with a long document delays the start and improves nothing.

What about staff using personal accounts on personal devices?
That situation cannot be prevented and can be addressed through obligation rather than through control. The rules that matter concern what information may leave the business and what has to be checked before work is delivered, and both apply regardless of which device was used. Framing the policy around information and output rather than around equipment closes the loophole. Attempting to police personal devices generally fails and damages trust in the process.

How often should the rules be revisited in practice?
On a stated cycle, with an owner responsible for the review, and additionally whenever a tool in active use changes materially. Quarterly review suits most smaller businesses, since it is frequent enough to track the pace of change and infrequent enough to be sustained. The review should compare the policy against the current inventory rather than reading the policy in isolation. Reviews that never produce a change are usually reviews that never looked at practice.

Saturday, May 16, 2026

Automation Exposes Whichever Process You Never Defined

Automation Exposes Whichever Process You Never Defined. Automating a vague process does not clarify it.

AI operations management fails most often because automation gets applied to a process nobody ever defined. A vague process run by people is corrected quietly by judgment at every step. The same process run by software repeats the vagueness at speed, and the errors arrive faster than anyone can intercept them.

Judgment Is the Hidden Correction Layer

Most business processes are far less defined than the people running them believe. The documented version describes a clean path. The actual version contains dozens of small corrections nobody records.

A clerk notices that an address looks wrong and checks it. A dispatcher sees an order that does not fit the usual pattern and calls the customer. A technician spots a part number that has changed and substitutes the correct one.

None of those corrections appear in any procedure document. They happen inside the head of an experienced person, in the space between one documented step and the next.

That correction layer is doing enormous work. It absorbs ambiguity, patches missing rules, and prevents bad inputs from travelling downstream. It is also invisible, which is why it gets removed by accident.

When a process is automated, the documented steps get encoded. The correction layer does not, because nobody knew it was there to encode.

The people who supplied the corrections rarely raise an objection either. They do not experience the corrections as work, since each one takes seconds and feels like ordinary attention rather than a task.

Asked whether the process is well understood, they answer yes and mean it. Their confidence is genuine and it describes their own competence rather than the process itself.

Speed Turns Ambiguity Into Volume

A vague process performed by hand produces occasional errors that people catch. The same process performed by software produces the same error rate applied to far more transactions.

Rate matters less than volume here. An error that appeared a few times a month becomes an error that appears a few times an hour, and the mechanism that used to catch it is gone.

Detection also degrades. Human errors are varied and noticeable, because different people fail in different ways. Automated errors are uniform, and uniformity looks like correct operation to anyone glancing at a summary.

The consequence is a delay between the start of the problem and the discovery of it. That delay is where the real cost sits, since every affected transaction has already moved downstream.

Reporting makes the delay worse rather than better. Dashboards summarise, and a summary of uniform output looks healthy right up to the point where someone opens an individual record.

Confidence in the numbers also rises after automation, because the figures now arrive without manual handling. Trust increases at exactly the moment scrutiny should have increased.

Some of those transactions reached customers. Some entered the accounting system. Unwinding them takes longer than the process took to run, which is how a productivity project becomes a cleanup project.

The Exception Rate Nobody Measured

Every process has an exception rate, meaning the share of cases that do not follow the standard path. Almost no business knows what its exception rate is, because exceptions are handled rather than logged.

The exception rate determines whether automation is a good idea. A process where nearly everything follows the standard path automates well. A process where a meaningful portion requires a decision does not.

The trouble is that people describe their work as if exceptions were rare. Asked how a process runs, they describe the standard path, because the standard path is what they think of as the process.

The exceptions live in habit rather than memory. They are handled so routinely that they stop registering as departures from anything.

Managers tend to guess low, and the people doing the work tend to guess low as well. Both groups are estimating from memory, and memory discards anything resolved without difficulty.

A short observation period fixes this cheaply. Watching the work for two weeks and marking every case that required a decision produces a number that no interview will produce.

Sound operational management of automated work treats that number as the entry criterion. Where exceptions are common, the correct move is to reduce them first, not to encode them.

What Definition Actually Requires

Defining a process is not the same as documenting it. Documentation records what people say happens. Definition establishes what should happen, including in the cases nobody wants to discuss.

A defined process states the outcome it exists to produce. Without that, no one can judge whether a variation is an error or an improvement, and the automated version will preserve both equally.

A defined process states its inputs and what makes an input acceptable. Most automation failures trace back to inputs that were always slightly wrong and always silently fixed.

A defined process states what happens when something does not fit. Not a vague escalation, but a named person, a stated timeframe, and a rule for what the system does while waiting.

A defined process states who is accountable for the output. Automated work loses ownership faster than manual work, because nobody feels responsible for a step they do not perform.

Definition also has a political dimension that documentation avoids. Writing down what should happen forces a choice between two departments who both believed their version was standard.

That choice belongs to a person with authority, and it cannot be resolved by a workshop. Processes without a decision maker end up defined by whoever is most persistent.

Producing those four statements takes days rather than months. Skipping them takes no time at all and costs considerably more later.

Where the Exposure Shows Up First

Certain processes reveal their lack of definition immediately under automation, and knowing which ones saves considerable pain.

Anything involving customer communication exposes fastest. Tone, timing, and appropriateness were all being judged in the moment by someone who knew the customer. Encoding the words without the judgment produces messages that are technically correct and situationally wrong.

Anything involving pricing or quoting exposes next. Quotes routinely include informal adjustments for relationship, urgency, or risk that never made it into any rule. Removing the adjuster produces quotes that are consistent and commercially poor.

Anything involving intake exposes soon after. Intake is where bad data enters, and it is usually the point where the most correction was happening invisibly. Automating intake without validation moves the error deeper into the business.

Scheduling sits somewhere in between. Assignment rules look simple until the informal factors appear, such as which technician a particular customer trusts and which route the dispatcher knows is slower than the map suggests.

Those factors were never written down because they change constantly and everyone doing the work already knew them. Encoding the visible rule alone produces a schedule that is defensible on paper and worse in practice.

Approval steps expose more slowly and more expensively. An approval that was always granted looks like a formality, until the case arrives where it should have been refused and nobody was watching.

Careful application of generated output inside operating work begins with the steps where a wrong result is visible immediately, not the ones where it hides in a ledger.

The Order That Works

The productive sequence is unglamorous and it has not changed. Observe, define, simplify, then automate whatever survives.

Observation comes first because description is unreliable. People are honest and still inaccurate about their own work, which is a well-established feature of how habits form.

Definition comes second, and it will surface disagreements. Two people who have run the same process for years will describe different rules, and discovering that before automation is a gift rather than a problem.

Simplification comes third and is the step most often skipped. Many steps exist because of a supplier who is gone, a system that was replaced, or a mistake made years ago. Automating them preserves history nobody needs.

Automation comes last and covers less ground than expected. A well-defined process usually reveals that only part of it is worth automating, with the judgment-heavy portion better left to a person supported by better information.

Each stage should also produce something usable on its own. Observation improves training. Definition improves onboarding. Simplification improves throughput before a single tool is purchased.

Sequencing the work that way removes the pressure to justify a spend. Value appears early and does not depend on whether the final automation step is ever taken.

That outcome disappoints people expecting a fully automated function. It also produces something that keeps working after the person who built it moves on, which the ambitious version rarely does.

The uncomfortable finding in most of this work is that automation was never the constraint. Businesses reach for tooling because the alternative is admitting that a process everyone relies on was never actually agreed. Agreeing it surfaces conflicts between departments that have been managed by avoidance for years. The technology arrives as a way to skip that conversation. It does not skip it. It stages the conversation publicly, at speed, in front of customers, which is a more expensive venue than a meeting room.

Frequently Asked Questions

How do you know whether a process is defined well enough to automate?
Ask three people who run it to describe what happens when a case does not fit the standard path. Consistent answers indicate a defined process. Different answers, or hesitation, indicate that judgment is filling a gap that no rule covers. That gap is exactly what disappears when the work moves to software.

What is an exception rate and how is it measured?
It is the share of cases that require a decision rather than following the standard path. Measuring it requires observation rather than interviews, because people underreport exceptions they handle by habit. A short logging period, where anyone touching the process marks cases needing judgment, produces a usable figure. That figure predicts automation success better than any feature comparison.

Should a business fix the process before buying anything?
Yes, and the fixing is usually cheaper than the software. Definition work requires attention rather than expenditure, and it produces benefits even when nothing gets automated afterward. Buying first creates pressure to encode the process as it currently stands, including the parts that should have been removed. That pressure is the mechanism by which bad processes become permanent.

What happens when experienced staff leave after automation?
The business loses the ability to judge whether the automated output is still correct. Judgment about acceptable output lives with people who performed the task, and that knowledge decays quickly once they stop performing it. Keeping a manual fallback and rotating someone through it periodically preserves the capacity to evaluate. Without that, errors persist until a customer reports them.

Which processes are the worst candidates?
Anything with a high exception rate, anything where a wrong result stays hidden for a long time, and anything where the standard was never written down. Approval steps are particularly risky, because their value only appears in the rare case that should be refused. Customer communication is risky for a different reason, since errors there are visible externally and damage trust immediately.

Is partial automation a failure?
Partial automation is usually the correct answer rather than a compromise. Most processes contain a mechanical portion and a judgment portion, and separating them cleanly delivers most of the benefit at a fraction of the risk. Attempts at full coverage tend to encode judgment badly and then require constant correction. A person supported by better information often outperforms a system asked to decide.

Monday, May 11, 2026

Strategy Cannot Be Delegated to a Tool

Strategy Cannot Be Delegated to a Tool. Tools execute decisions.

An AI business strategy is a set of decisions about where a company competes, what it will stop doing, and which advantage it intends to build. Tools carry out decisions. They do not make them. Organisations disappointed by AI usually bought capability before deciding what that capability was meant to change.

The Decision That Gets Skipped

Strategy answers a small number of questions. What is this business trying to be better at than its competitors. Which customers matter most. What will be given up in order to fund the answer.

Those questions predate every technology and survive every technology. They were the same questions before spreadsheets, before the internet, and before any model could write a paragraph.

What changes with a new capability is the range of available answers, not the necessity of answering. A tool widens the option set. Widening the option set does not choose from it.

The disappointment pattern is consistent. A business buys capability, distributes access, waits for improvement, and finds that activity increased while position did not.

Nothing malfunctioned. The tool did exactly what tools do, which is execute faster in whatever direction it was pointed. Direction was the missing input.

The skip is easy to understand. Deciding is unpleasant, because a decision closes options and assigns blame if it proves wrong. Buying feels like progress and defers the closing of options indefinitely.

Purchases also produce visible motion. A subscription, an announcement, a training session, and a dashboard all look like a company doing something. A written choice about what the business will refuse looks like a single page.

Amplification Is Not Direction

A useful way to think about any capability is as a multiplier applied to existing intent. Multipliers are indifferent to what they multiply.

A business with a clear position and a defined customer gets more of that position, faster. A business without one gets more of its confusion, faster and at greater volume.

This shows up most visibly in marketing output. Content production accelerates dramatically, and the acceleration reveals that nobody had decided what the business was trying to say.

The result is more material saying less, which is worse than the previous condition rather than better. Volume without a point costs attention on both sides of the transaction.

The same dynamic appears in analysis. A team that could not previously produce enough reporting now produces far too much, and the constraint moves from availability of data to willingness to decide.

Amplification also hardens whatever the business already believes. Prompts encode assumptions, and outputs written from those prompts return the assumptions in polished form. Confidence rises while accuracy stays where it was.

A business with a wrong theory about its customers will now express that theory more fluently, more often, and across more channels. Fluency is frequently mistaken for validation.

Anyone reviewing how strategic choices should be made before any tool is selected tends to find the shortage sits in decision making, not in capability.

Three Questions That Come First

Before any purchase, three questions deserve written answers. Written matters, because unwritten answers stay comfortably vague.

The first question is what the business wants to be measurably better at within a year. Better at responding to inbound enquiries. Better at quoting accurately. Better at retaining the customers already won.

Generic ambitions fail this test immediately. Being more efficient is not an answer. Efficiency is a direction of travel, not a destination anyone can recognise on arrival.

The second question is what the business will stop doing to make room. Every genuine strategy contains a subtraction, and strategies without one are wish lists.

Adoption consumes attention, and attention in an operating business is already committed. Something has to be dropped, deferred, or accepted as worse. Naming it in advance prevents the slow abandonment that otherwise follows.

The third question is what would have to be true for the capability to matter. If the answer depends on data the business does not collect, or on volume it does not have, the honest conclusion is to wait.

A fourth question is optional and clarifying. Ask what a competitor would have to see in order to be worried. Answers that no competitor would notice describe internal convenience rather than advantage.

Internal convenience is worth having and should be called by its name. Confusing it with strategy is what leads businesses to describe faster document handling as a change in market position.

Three written answers make the purchase decision straightforward. Without them, every vendor demonstration looks compelling, because every demonstration is designed to answer questions the buyer has not asked.

Why Vendors Cannot Supply the Decision

Vendors are not being evasive when they fail to provide strategy. They are structurally unable to provide it, and expecting otherwise misreads the relationship.

A vendor knows the capability well and knows the business barely at all. What a supplier can demonstrate is what the tool does. What a supplier cannot know is which of the buyer's processes is worth changing.

That judgment requires knowledge of the customer base, the cost structure, the competitive position, and the tolerance of the staff. None of that fits in a demonstration.

The incentive structure compounds the gap. A vendor is rewarded for adoption, and adoption is a poor proxy for advantage. A tool can be used constantly and still change nothing about why customers choose the business.

Buyers can still use vendors well by changing what they ask for. Requesting a description of which businesses the tool has not suited produces more useful information than any capability walkthrough.

Suppliers who can answer that question honestly are worth more than suppliers who cannot, and the answer costs nothing to request.

The same limitation applies to any adviser who arrives with a solution already selected. The order of operations matters more than the expertise. Diagnosis before prescription, or the prescription is a guess wearing confidence.

What a Strategy Actually Looks Like

A working strategy in a small or mid-market business fits on one page and can be recited by the leadership team without reading it.

It names the customer segment that matters most. It names the thing the business intends to do better than alternatives. It names the activities being reduced to fund that focus.

Only after those statements exist does technology enter the conversation, and it enters as a question. Which of these commitments can be advanced faster with the capability now available.

Sometimes the answer is none of them, and that is a legitimate outcome. A business whose advantage rests on relationships built over years may find that automating communication weakens exactly what it sells.

A one-page strategy also survives contact with staff, which longer documents rarely do. People execute what they can hold in memory during a working day, not what sits in a shared folder.

The recital test is not a stylistic preference. A leadership team that cannot state the position without notes has not agreed on it, and disagreement discovered later is far more expensive.

More often the answer is one or two commitments, applied narrowly. Narrow application against a named commitment produces evidence quickly, and evidence is what justifies the next decision.

An examination of what changes when a company organises itself around these capabilities deliberately shows the sequence running from position to process to tool, never the reverse.

The Cost of Reversing the Order

Businesses that buy first and decide later pay in three currencies, and none of them appear on an invoice.

The first is credibility. Staff watch leadership introduce something with enthusiasm, then watch it fade without explanation. The next initiative meets a colder room, regardless of its merit.

The second is opportunity. Attention spent on an undirected experiment is attention not spent on the constraint that was actually limiting growth. Small businesses have very little of it to misallocate.

The third is learning. An experiment without a stated hypothesis produces no conclusion. When the outcome is ambiguous, everyone keeps their prior belief, and the organisation ends the quarter knowing exactly what it knew before.

A fourth cost appears later and is harder to reverse. Processes get rebuilt around a tool that was never chosen against a purpose, and unwinding them costs more than the original adoption did.

Reversal is also politically difficult once people have been trained and roles have shifted. The sunk investment argues for continuation long after the evidence stops supporting it.

These costs are avoidable at almost no expense. Writing down the intended outcome before the purchase takes an afternoon and converts an experiment into a test.

A test can fail usefully. An initiative can only fail embarrassingly, which is why so many of them are quietly kept alive well past the point of usefulness.

The uncomfortable part of this argument is that it removes the excuse. When strategy is treated as something a tool might supply, the absence of strategy can be blamed on the tool. Once the sequence is stated plainly, the responsibility returns to where it always sat. Someone has to decide what the business is for, what it will refuse, and what it intends to be better at. No purchase substitutes for that decision, and no amount of capability compensates for skipping it.

Frequently Asked Questions

Does every business need an AI strategy?
Every business needs a strategy, and the technology question is a subsection of it rather than a separate document. Producing a standalone plan for one category of tool usually signals that the underlying strategy is unclear. The better sequence starts with the commitments the business has already made and asks which of them can now be pursued differently. A separate plan tends to create a separate set of goals that compete with the real ones.

How do you know whether a tool is worth buying?
Write down what the business expects to be measurably better at, and by when, before seeing any demonstration. If that statement cannot be written, the purchase is premature regardless of how impressive the capability looks. Vendors answer questions the buyer arrives with, so arriving without questions guarantees a persuasive but unhelpful meeting. The written statement also gives the review an honest benchmark later.

What if competitors are moving faster?
Speed matters only when the direction is right, and most early movement in a new category is undirected. Watching competitors spend attention on experiments is genuinely informative and costs nothing. The risk of waiting is real but smaller than the risk of committing the limited change capacity of a business to the wrong target. Late and correct outperforms early and scattered.

Who should own this decision in a smaller company?
The person accountable for the commercial result, which in most smaller companies is the owner. Delegating it to whoever is most technically curious produces tool selection rather than strategy. Technical input is valuable at the evaluation stage and misplaced at the direction stage. The two roles should be kept separate even when the team is small enough that they overlap in practice.

Can strategy be developed after adopting a tool?
It can, and it usually is, which is why so many adoptions disappoint. Working backward from a purchase means the strategy gets shaped to justify the spend rather than the spend being shaped by the strategy. The recovery path is to stop, write the missing statement of intent, and evaluate the existing tool against it honestly. Some tools survive that review and many do not.

What is the most common sign a strategy is missing?
Activity rises while nothing about the competitive position changes. More output, more reports, more experiments, and no clearer answer to why a customer should choose this business over the alternative. The absence also shows in how decisions get made, with each new option evaluated on its own merits rather than against a stated commitment. Where everything looks worth doing, nothing has been decided.

Thursday, May 7, 2026

Small Businesses Are Not Small Enterprises

Small Businesses Are Not Small Enterprises. Enterprise AI advice assumes a technical function and a change budget.

AI for a local business should begin with one repetitive task, one named owner, and a manual fallback. Most published guidance assumes a technology function, a change budget, and a committee to run pilots. A business with a couple of dozen staff has none of those, so the correct opening move is smaller and far more specific.

The Advice Was Written for a Different Animal

Almost every widely circulated framework for adopting AI was designed inside large organisations. The vocabulary gives it away. Governance boards, centres of excellence, data readiness assessments, phased rollouts across business units.

Each of those terms describes a solution to a problem that large organisations genuinely have. Thousands of employees cannot be coordinated informally. Regulated data cannot be handled by whoever happens to be curious.

A local business has none of those problems. It has a different set, and applying the enterprise remedy to the small-business condition produces delay rather than safety.

The mismatch is not one of degree. A small business is not a large business with the numbers reduced. It is a structurally different organism, and the differences change what the first correct action is.

Three Things a Local Business Does Not Have

The first missing element is a technical function. There is no internal team whose job is to evaluate tools, integrate systems, and maintain what gets deployed. Whoever adopts a tool also owns it forever.

That single fact rules out most enterprise starting points. Advice that begins with connecting a model to internal data assumes someone will maintain the connection. In a small business, that someone is already fully occupied doing something else.

The second missing element is a change budget. Large organisations fund adoption separately from operations, which means the cost of learning is absorbed somewhere other than the profit and loss of the team doing the learning.

Small businesses have no such buffer. Every hour spent experimenting is an hour not spent serving customers, and that trade is felt immediately by the person making it.

The third missing element is tolerance for a failed pilot. Enterprises run many small bets and expect most to die quietly. A local business that wastes a quarter on a tool that goes nowhere will not try again for a long time.

These three absences are not weaknesses to be corrected. They are permanent conditions, and any approach worth following has to work inside them rather than around them.

What a Local Business Has Instead

The comparison is not one-sided. Small businesses hold advantages that large organisations spend enormous sums trying to simulate, and those advantages should shape the approach.

Decisions require one conversation. There is no procurement cycle, no security review, no committee that meets monthly. An owner who decides on Tuesday can have something running on Wednesday.

The distance between the decision maker and the work is short enough to observe directly. An owner can watch the task being done, notice where the time goes, and judge whether the output is acceptable without commissioning a study.

Feedback arrives immediately. When something breaks, the customer calls the same day and the person who changed the process hears about it. Large organisations pay consultants to reconstruct that signal.

These advantages favour a specific method. Choose narrowly, start immediately, watch closely, and keep the ability to revert. A practical view of where a smaller company should actually apply these tools first looks nothing like a phased enterprise programme.

The Correct First Move

Start with a task, not a technology. The task should be repetitive, frequent, low in consequence, and already understood by whoever performs it.

Frequency matters more than difficulty. A task performed many times a week returns the learning investment quickly, even when each instance is small. A hard task performed twice a month never repays the effort of automating it.

Low consequence matters because the first attempt will be imperfect. Drafting a first version of a customer reply is a safe place to be wrong. Calculating a payroll figure is not.

Existing understanding matters because the tool has to be judged against something. If nobody can say what a good output looks like, nobody can tell whether the tool is helping.

Candidate tasks in most local businesses look similar. Turning voice notes into written records, drafting routine replies, summarising a week of messages, rewriting the same quote for different customers, preparing a first pass at a job description.

None of those are impressive. Every one of them is performed weekly by someone whose time is worth more elsewhere, which is exactly the point.

Why the Cheap Tool Is Usually the Right One

Small businesses are frequently steered toward specialised software built for their industry. The pitch is that a general tool cannot understand the specifics of the trade.

That claim is often true and usually irrelevant at the start. Specialised software requires configuration, data migration, and staff training before it produces anything. A general assistant produces something in the first hour.

The purpose of the first attempt is not to solve the biggest problem. It is to find out whether the business can absorb a change in how work gets done, which is the real constraint.

Starting with a general tool also keeps the exit cheap. If the experiment fails, the loss is a subscription and some hours. If a specialised platform fails after migration, the loss includes the data structure the business built around it.

A grounded account of how a general assistant fits into everyday small-business work is more useful early than any vertical product comparison. It tests the organisation rather than the software.

Vertical products become the right answer later, once the business knows which task it wants handled and what acceptable output looks like. Buying one before that point means paying for configuration against requirements nobody has established.

What to Put in Place, and What to Ignore

Three conditions separate the small businesses that make this work from the ones that abandon it after a month.

A named owner comes first. Not a committee, not the whole team, one person who is responsible for whether the tool is used and whether it is producing acceptable output. Shared responsibility here reliably becomes no responsibility.

A manual fallback comes second. Every automated step needs a documented way to do the job by hand, and someone who still remembers how. Tools change pricing, change terms, and occasionally disappear.

The fallback also protects against a subtler risk. When the only people who understood a process have stopped performing it, the business loses the ability to judge whether the automated version is still correct.

Fallbacks also change the negotiating position at renewal time. A business that can still perform the task by hand treats a price increase as a choice rather than an emergency.

A review date comes third. Set a point in the calendar to decide whether the task is genuinely better handled this way. Without that date, the default answer is continuation, and unused subscriptions accumulate quietly.

These three conditions cost nothing. They also make the difference between a business that adds one working capability per quarter and one that collects tools it does not use.

Some enterprise practice transfers cleanly and some does not, and confusing the two wastes considerable time.

What transfers is the discipline of naming the outcome before choosing the tool. That habit is scale independent and it is the single most useful thing to borrow.

What transfers is the insistence on a human check before anything reaches a customer. Large organisations enforce it through policy. Small ones enforce it by keeping the volume small enough that checking is realistic.

What does not transfer is the phased programme. Phases exist to coordinate people who cannot all be in one conversation. A small team can simply have the conversation.

What does not transfer is the readiness assessment. Assessing readiness for months is a luxury purchased with a change budget that does not exist here. The faster test is to try one task and observe what happens.

What does not transfer is the vendor selection matrix. Scoring many products against weighted criteria makes sense when a decision binds thousands of users for years. It makes little sense when the commitment is a monthly subscription one person can cancel.

What does not transfer is the expectation of a dedicated specialist. Nobody is arriving to own this. The owner is whoever is already in the building, which is why the first task has to be small enough to fit alongside real work.

The pattern holds across every borrowed practice. Anything that exists to coordinate strangers can be discarded. Anything that exists to protect judgment should be kept, because judgment is scarce at every size.

The most damaging idea in circulation is that small businesses are behind. Being late to a technology is not a strategic problem, and the businesses that waited have watched other people pay for the early mistakes. The real risk is different. It is adopting an enterprise method inside a company that cannot support it, spending the limited appetite for change on ceremony, and concluding that the tools do not work. They work. The method has to match the animal.

Frequently Asked Questions

Does a small business need someone technical before starting?
No, and waiting for that person is usually the reason nothing begins. The first useful applications require no integration, no code, and no data preparation. What they require is someone willing to test outputs against a known standard and decide whether the result is good enough. Technical capability becomes relevant much later, when tools are being connected to systems of record.

How much should a local business expect to spend at the start?
The subscription is rarely the meaningful cost. The real expenditure is internal attention, which is the scarcest resource in a small company. A realistic plan budgets a few hours a week of one person for the first stretch, then reassesses. Businesses that budget only for software are budgeting for the smallest line.

Which task should be automated first?
Something repetitive, frequent, low in consequence, and already well understood by whoever performs it. Frequency determines whether the effort repays itself. Low consequence means an imperfect first attempt causes no damage. Existing understanding means someone can judge whether the output is acceptable, which is impossible when the task was always vague.

Is industry-specific software better than a general tool?
Eventually, often yes. At the start, usually no. Specialised platforms require configuration and migration before producing value, and that delay consumes the appetite for change before any benefit appears. A general tool tests whether the organisation can absorb a new way of working, which is the question that actually needs answering first.

What happens if staff resist?
Resistance is usually accurate rather than obstructive. Staff who resist are often protecting a step that exists for a reason nobody wrote down. The productive response is to ask what breaks if the step disappears, and to treat the answer as information rather than objection. Adoption that skips this conversation produces quiet workarounds instead of open refusal.

How long before a small business sees a return?
On a well-chosen first task, relief usually shows within a few weeks, because the task was picked for frequency. Broader returns take longer and depend on whether a second and third task follow. Businesses that stop after one application get one improvement. The compounding comes from repeating the selection method, not from the first tool.

Tuesday, April 28, 2026

Nobody Sells You the Thing You Actually Need

Nobody Sells You the Thing You Actually Need. Advice, implementation and software are three different purchases.

AI consulting is three different purchases wearing one name. Advice tells a business what to do. Implementation builds the thing. Software licenses a capability. Buyers usually describe a need in one of those categories and receive a quote from another, which is why so many engagements end with both sides feeling misled.

Three Purchases That Sound Like One

Advice is the purchase of judgment. The deliverable is a decision, a priority order, or a recommendation not to proceed. Nothing is built and the value lies in the choices that get avoided.

Implementation is the purchase of construction. Something specific gets configured, connected, tested, and handed over. The deliverable is a working system and the specification has to exist before the work starts.

Software is the purchase of capability. A vendor provides a tool with defined functions and the buyer supplies everything else, including the decision about what the tool is for.

These three are sold by different organisations with different economics. They are also frequently sold by the same organisation, which is where the confusion becomes profitable rather than merely unfortunate.

A buyer who needs advice and purchases implementation ends up with a well-built system solving a problem that was never the constraint. A buyer who needs implementation and purchases advice ends up with a document and no change. Both outcomes are common and neither is caused by incompetence.

Why the Mismatch Happens So Reliably

Buyers describe symptoms rather than categories. The opening statement is usually some version of wanting to use AI to become more efficient. That sentence does not indicate which of the three purchases is required.

Sellers interpret ambiguous requirements through the lens of what they sell. This is not cynicism. A firm that builds systems genuinely sees a building problem, and a firm that advises genuinely sees a decision problem.

The result is that the same enquiry produces three different diagnoses depending on who receives it. Each diagnosis is internally coherent and each is partly right, which makes comparison across quotes almost meaningless.

Buyers then compare on price, because price is the only dimension that appears comparable. Comparing the price of advice against the price of construction is a category error that produces predictably poor decisions.

Understanding how consultants, agencies and vendors differ in what they are actually selling resolves most of this before any conversation begins. The distinction is structural rather than a matter of quality.

The mismatch also survives because both parties want the meeting to go well. A buyer who cannot articulate the need adopts the framing offered by the seller, since that framing at least sounds definite. Agreement is reached on a description neither side wrote.

That agreement holds until delivery, at which point the buyer discovers the scope does not cover the part they cared about. The scope covered exactly what was written, and what was written came from a conversation where nobody named the category.

What Each Seller Is Positioned to Recommend

Incentives shape recommendations even among honest sellers, and pretending otherwise leads to bad buying.

A software vendor is compensated for seats and renewals. Its natural recommendation is that the tool be deployed widely and quickly, because usage drives retention. It has no economic reason to advise that a process be redesigned first, and often no capability to do so.

An implementation agency is compensated for built work. Its natural recommendation is a project with defined scope, because that is what can be priced and delivered. It has little incentive to conclude that the correct answer is a change of policy costing nothing.

An advisory firm is compensated for judgment, which creates its own distortion. Advice is easier to sell repeatedly than to make consequential, and the failure mode is a stream of recommendations that never quite reach implementation.

None of these incentives make a seller untrustworthy. They do mean that a buyer should know which one they are speaking to and should discount the recommendation in the direction set by the economics of that seller.

The practical protection is to obtain the diagnosis from a party with no stake in the remedy. That separation costs something and it costs far less than building the wrong thing well.

Diagnosing Which One the Business Actually Needs

Three questions usually settle the category, and they can be answered before any seller is contacted.

The first asks whether the business knows which specific process it wants changed. If the answer is no, the need is advice, and any implementation quote received at this stage is premature regardless of how attractive it looks.

The second asks whether the process is documented well enough that a stranger could describe the desired end state. If the process is known but undefined, the need is process work, which sits between advice and implementation and is rarely what anyone quotes for.

The third asks whether the specification exists and only the building remains. If so, the need is implementation, and the buyer should be comparing builders on capability and price rather than on strategic insight.

Most businesses that believe they need implementation are actually at the second question. The process is understood by the people who run it and has never been written down in a form anyone could build from.

Discovering that early saves the most money. A build against an undefined process produces something that works exactly as specified and does not fit how the work happens. That is the most expensive form of technically successful project.

A fourth question is worth asking privately. What would change in the business if the project succeeded completely? If the honest answer is a modest saving on a task nobody minds doing, the spending belongs elsewhere.

That question filters more proposals than any technical review. Projects survive scrutiny on feasibility and fail on relevance, and relevance is the cheaper of the two to test.

Buying in the Right Sequence

The sequence that works runs from diagnosis to definition to construction to licensing. Most buying runs in the reverse order, starting with a tool that impressed someone and working backwards to justify it.

Reverse-order buying is not irrational. The tool is the most visible element, it demonstrates easily, and it comes with a price that can be approved without a business case. Everything upstream of it is abstract by comparison.

The cost of that order appears later. The tool is in place, the process it should support has not been defined, and the definition work now has to happen with a sunk commitment shaping the answers.

Running the sequence properly means accepting a slower start. The first phase produces no visible technology and its output is a specification, a priority order, and a list of things not to do. That deliverable is unsatisfying and it determines whether everything after it works.

A serious view of what an implementation roadmap should contain before any building starts gives a buyer a way to judge proposals. The test is whether the thinking was done or skipped. The presence of a sequence, rather than a list of features, is the signal.

Slowing the start also gives the business time to notice its own capacity. Every one of these engagements consumes internal attention, and a company running three at once will deliver none of them properly.

What a Good Proposal Looks Like

A proposal worth accepting states which of the three purchases it is. That sounds elementary and very few proposals do it.

It defines the deliverable in terms the buyer could verify without technical help. A written recommendation, a documented process, a configured system passing stated tests. Anything described only as support or partnership is unbounded and should be rewritten before signature.

It names what is excluded. Exclusions are more informative than inclusions, because they reveal whether the seller has understood where the work ends and who picks it up afterwards.

It states who owns the result and what happens when the engagement ends. Ownership of configurations, instructions, and accounts should transfer to the buyer as a matter of course, and a proposal silent on this point is telling the buyer something.

It also states what the seller will need from the business. Engagements fail as often through unavailable internal time as through poor delivery, and a proposal that asks for nothing has not thought about how the work actually happens.

Any proposal that answers all five points can be compared fairly against another that does the same. Comparison between a proposal that answers them and one that does not is not a comparison at all, and price is the least useful place to start.

The market for this work is young enough that the categories are still blurred, and the blurring benefits sellers more than buyers. Clarity is available at no cost to anyone willing to decide, before the first conversation, which of the three things they are buying. The businesses that get value from this spending are usually not the ones that found the best supplier. They are the ones that knew what they were shopping for.

Frequently Asked Questions

What is the difference between an AI consultant and an AI agency?
A consultant sells judgment and the deliverable is a decision or a recommendation. An agency sells built work and the deliverable is a configured, working system. The two require different inputs, because an agency needs a specification that a consultant is often hired to produce. Firms offering both should be asked which capability is leading the engagement.

How can a buyer tell which category they need?
Ask whether the specific process to be changed is known, and whether it is documented well enough to build from. An unknown process indicates a need for advice. A known but undocumented process indicates process work. A documented process with an agreed end state indicates implementation.

Is it a problem if one firm offers all three?
Not necessarily, provided the phases are separately scoped and separately priced. The risk is that the diagnosis phase is conducted by a party whose revenue depends on the build that follows. Splitting the diagnosis from the delivery removes that pressure. Where splitting is impractical, the diagnosis deliverable should be usable by another supplier.

What should a diagnosis phase cost relative to the build?
It is normally a small fraction of the total and it determines whether the remainder is spent well. Buyers frequently resist paying for a phase that produces no visible technology. That resistance is the single most reliable predictor of an implementation that solves the wrong problem.

Who should own the tools and configurations after an engagement?
The buyer, without exception. Accounts, instructions, prompts, integrations, and documentation should sit in the name of the business rather than the supplier. Arrangements where the supplier retains ownership create a dependency that is expensive to unwind later. This should be settled in writing before work begins.

What is the most common wasted spend in this category?
Building a well-specified system for a process that was never the constraint. The work is delivered competently, the tool functions as promised, and nothing measurable improves because the bottleneck was elsewhere. That outcome is a diagnosis failure rather than a delivery failure. It is also almost entirely preventable by asking what would change if the project succeeded perfectly.

Friday, April 24, 2026

Consultant, Agency or Vendor: What You Are Actually Buying

Consultant, agency or vendor. Three different things are sold under one word. The mismatch is what disappoints.

AI business consulting describes three different products sold under one word. A consultant sells judgment about which decisions to make. An agency sells execution capacity for decisions already made. A vendor sells software and the support that surrounds it. Buyers who confuse the three pay for one thing and expect another.

Three Products, One Word

The word consulting has stretched to cover almost any paid outside help involving AI. That stretch is convenient for sellers and expensive for buyers. Underneath the shared label sit three businesses with different economics, different staffing, and different definitions of a job done well.

The stretching happened because demand arrived faster than expertise did. Every firm with adjacent capability repositioned toward the term, and the term absorbed all of them without resistance. A word that describes everything eventually stops describing anything a buyer can compare.

A consultant is paid to reduce uncertainty on behalf of a client. The output is a decision the client can defend: what to automate, what to leave alone, what sequence to follow, and what to stop doing entirely. Hours form the cost structure, and judgment is the actual product.

An agency is paid to produce work at a defined standard. The output is deliverables at volume, including content, campaigns, configured workflows, integrations, and trained staff. Agencies are excellent at execution and structurally poor at telling a client the work should not happen.

A vendor is paid for a product that must keep functioning. The output is software that works, plus onboarding and support sized to keep the subscription active. A vendor recommends its own category by design, and pretending otherwise helps nobody involved.

Understanding what advisory work on AI actually covers makes the boundaries visible again. Advisory work ends where implementation begins, and that ending is a feature rather than a limitation. Blurring the boundary produces engagements where nobody can state clearly what was purchased.

The categories are easy to identify from the questions asked in a first conversation. A consultant asks about decision rights, constraints, and what has already been tried. An agency asks about volume, timelines, and approval workflow. A vendor asks about current systems and seat counts.

The practical question is which of the three a business needs first. That answer depends entirely on whether the core decision has already been made. Working through which of the three fits a given situation before writing a brief prevents most of the disappointment that follows.

The Mismatch Produces the Disappointment, Not the Price

Failed engagements get described afterward in financial terms almost without exception. The fee was too high, the retainer ran too long, the return never appeared. Those descriptions are accurate and completely beside the point.

The failure almost always started as a category error made early. A business bought execution while the underlying decision was still open, or bought advice after the decision had already hardened. Both errors feel like value problems and are actually fit problems.

Buying execution too early is the more expensive of the two mistakes. An agency handed an undefined problem will build something competent, because building is what agencies exist to do. The output arrives on schedule and solves a problem the business never confirmed was worth solving.

Deciding where AI belongs in the business before capacity is purchased is unglamorous work that prevents this outcome. It requires naming which processes carry real cost, which carry only visible friction, and which remain too unstable to automate at all. Very little software appears anywhere in that conversation.

Buying advice too late produces the opposite form of waste. When the decision is settled and the constraints are known, further analysis restates what the room already agreed on. The business needs hands at that point rather than a second opinion.

Scope documents provide an early warning that a mismatch is forming. A scope written entirely in activities, such as workshops held or assets produced, describes motion instead of outcome. A scope that names a decision to be reached, or a standard to be met, tends to survive contact with reality.

Packaged offerings occupy the middle ground and confuse buyers most reliably. Reviewing productized offerings that assume the problem is already defined is worthwhile once the problem genuinely has been defined. A packaged solution applied to an undefined problem simply hides the definition step inside a scope document nobody rereads.

What Each Should Be Held To

Accountability differs by category, and applying the wrong standard breaks otherwise sound relationships. Holding a consultant accountable for deliverable volume produces slide decks nobody needed. Holding an agency accountable for strategic insight produces confident recommendations from people paid to stay busy.

A consultant should be judged on decision quality and on how quickly ambiguity gets resolved. The right test is whether the client can now explain the reasoning without the consultant in the room. An advisor who becomes structurally necessary has failed at the assignment.

An agency should be judged on throughput, consistency, and adherence to a defined standard. The right test is whether the work would be recognized as correct by someone who has never met the agency. Volume produced without a standard is simply cost with a delivery schedule.

A vendor should be judged on uptime, support responsiveness, and the cost of leaving the platform. That last item receives the least attention during selection and matters the most across a multiple-year horizon.

Reference conversations become useful once the categories are separated. The question worth asking a consultant reference is what decision the work produced and whether it held. The question worth asking an agency reference is whether quality stayed consistent after the first month. The question worth asking a vendor reference is what the migration looked like when someone tried to leave.

Compensation structure explains most provider behavior better than stated intent does. Hourly work rewards depth and punishes speed, retained work rewards continuity, and subscription work rewards renewal above everything else. Reading the incentive before signing predicts the relationship more accurately than any reference call.

Smaller companies face this problem with fewer resources to absorb an error. Advisory work scaled to a company without a large internal team differs from enterprise engagements in duration, depth, and the amount of implementation folded into the same relationship. The three categories still exist, though one person may end up covering two of them.

Capability categories create a second layer of confusion for buyers. Drafting, summarizing, and content production get sold as a single competence, though buyers rarely separate the tool from the process surrounding it. Treating content and drafting capability as an operating capability rather than a product line clarifies who should be responsible for the result.

The Handoff Is Where Value Leaks

Every engagement of every category ends at a boundary of some kind. Advice ends at a recommendation that someone must accept. Execution ends at a delivered artifact that someone must adopt. Product ends at a working login that someone must use.

None of those endings is the same thing as a changed business. The gap between the last deliverable and a durable operating change is where most of the value quietly escapes. Nobody is contractually responsible for that gap, which is exactly why it persists across engagements.

Closing it requires deciding in advance who runs the new process after the outside party leaves. That person should be named in the engagement rather than discovered afterward. An engagement without a named internal owner is a purchase rather than a change.

Knowledge transfer deserves the same treatment as any other deliverable. Written decisions, rejected options, and the reasoning behind both should sit inside the business when the relationship ends. Engagements that leave behind only finished artifacts leave the company unable to adjust when conditions shift.

The operating layer is where all of this becomes concrete. Treating the daily management of automated work as its own discipline assigns responsibility for monitoring output quality, handling exceptions, and retraining staff as the process shifts. Without that layer, a well-built system degrades without anyone noticing.

Sequencing matters as much as ownership in determining whether value survives. A sound progression from decision to deployment stages the work so each phase produces something usable rather than something impressive. Sequencing failures create long projects with no interim value, which is precisely how sponsors lose patience.

The end state is worth describing plainly, because it looks nothing like the sales material. A company where the work has been genuinely embedded into how the business runs appears unremarkable from outside. Processes move faster, exceptions get handled deliberately, and no separate AI initiative exists because the initiative already finished.

Buyers rarely get burned by the wrong price. They get burned by asking a category of provider to do something that category was never built to do, then blaming the provider for behaving exactly as its economics require. Naming the three products before writing the brief is a short exercise. It costs an afternoon and prevents the kind of engagement that ends politely, on budget, and with nothing changed.

Frequently Asked Questions

What is the difference between an AI consultant and an AI agency?
A consultant is paid to reduce uncertainty and produce a decision the client can defend. An agency is paid to produce work at volume once that decision already exists. The consultant should work toward becoming unnecessary, while the agency should work toward becoming reliable. Confusing the two produces either expensive analysis of a settled question or fast execution of an unexamined one.

When should a business hire a vendor instead of a consultant?
A vendor makes sense once the problem, the owner, and the required category of software have all been settled. At that point the remaining questions are functional, and a product conversation answers them efficiently. Hiring a vendor before those elements are settled means the vendor defines the problem. Vendors define problems in terms their own product already solves.

Can one provider be all three?
Some can, and smaller engagements frequently require exactly that. The risk is that advice and execution sit inside the same fee, which removes any incentive to recommend doing less. Buyers accepting this arrangement should insist that the recommendation phase produce a written decision, including the options that were rejected. That record keeps the advisory work honest even when the same party performs the build.

How long should an AI advisory engagement run?
Advisory work should end once the client can explain the reasoning without assistance. That point arrives sooner than most retainers assume it will. Engagements continuing past it convert judgment into maintenance, which is handled better internally or by an execution partner at lower cost. A defined end date protects both parties from a slow drift into dependence.

What is the most common reason these engagements disappoint?
The category was wrong for the stage the business had actually reached. Execution capacity was purchased before the decision was made, or analysis was purchased after it had already been settled. Neither error appears as a fit problem in the postmortem, because both of them look like poor value. The fix happens before the contract rather than during it.

Who should own the work after the engagement ends?
A named person inside the business, identified before the engagement starts. That person should control the process being changed rather than the technology being installed. Without this, the new way of working has no defender once the outside party leaves. Most decay in delivered systems traces back to that single omission.

Outsourced Against Fractional, Decided On Cost Structure

Companies facing an operating leadership gap must choose between two models that sound similar but work differently. Outsourced coo service...