Implementation Q&A

How an AI implementation becomes a working operating model.

Twelve practical questions and answers about choosing a first use case, handling data and adoption problems, measuring value, and scaling safely.

You implemented AI across the company. Where did you begin?

We did not begin with a platform purchase. We began by finding work that was both important and observable: resolving supply-chain exceptions. It had clear owners, repeated steps, measurable delay and service costs, and enough complexity that experienced people were spending most of their time gathering information rather than making the decision. That gave us a useful baseline before we built anything.

What made that a better first use case than a broad company assistant?

A broad assistant produces broad expectations. A narrow use case gives you a defined user, a specific decision, a limited data boundary, and a way to measure whether the process improved. We could compare an AI-supported exception with the existing process: time to understand the situation, quality of the recommendation, number of escalations, and impact on customer commitments.

What was the hardest technical difficulty?

Context. The information existed, but it was inconsistent across systems. Part numbers were named differently, policies were stored in documents with unclear ownership, and the most useful operational knowledge lived in individual inboxes. The model was not the hard part. The hard part was defining which source was authoritative, who could see it, how fresh it had to be, and how to show the user where an answer came from.

How did you resolve that context problem?

We created a small context team with people from operations, data, security, and the process itself. They did not try to clean every dataset in the company. They mapped the minimum evidence needed for one decision, assigned an owner to each source, added timestamps and permissions, and removed documents that were obsolete or duplicate. That work improved the human process before the AI assistant was even released.

How did you choose between models and vendors?

We treated it as a portfolio decision. A language model created the brief, retrieval found policy and similar cases, a structured model estimated delay risk, and a rules layer enforced non-negotiable constraints. We tested vendors on our cases, with our records and our users. A polished demonstration did not count as evidence. We compared quality, source grounding, latency, integration effort, data controls, cost per successful case, and how easily we could switch providers later.

What did you learn about human adoption?

People do not adopt a system because leadership says it is innovative. They adopt it when it saves a real step and when they can see why the output is credible. The first version was too verbose and too confident. Planners either ignored it or trusted it too much. We made it shorter, showed the sources, highlighted uncertainty, and designed a clear escalation path. Trust improved when users could correct the system and see that their corrections mattered.

Did you allow the system to take actions?

Not at first. It could gather evidence, draft a recommendation, and route work to the right planner. It could not change a schedule, issue a commitment, contact a customer, or move money. We added action permissions only for low-risk, reversible steps after the system had demonstrated reliable performance. Autonomy is not a switch. It is a series of decisions about risk, permissions, review, and rollback.

How did you evaluate whether the pilot was genuinely successful?

We used a mixed scorecard. Quality meant that the recommendation was accurate, complete, and grounded in the right evidence. Operational value meant less handling time and fewer avoidable escalations. Economics included model usage, integration, monitoring, human review, and failure handling. Adoption meant that people used it appropriately, corrected it when needed, and did not create workarounds. We reviewed the results with the process owner every two weeks.

What setbacks changed your strategy?

Early on, we tried to automate a process before standardizing it. Different teams followed slightly different rules, so the system was asked to imitate inconsistency. We paused, made the policy decisions explicit, and documented the exceptions instead of hiding them in prompts. Another setback was measuring only speed. A fast answer that causes a planner to redo the work is not productivity. We changed the metric to successful resolutions per planner-hour.

How did you organize the company around AI after the pilot?

We kept process ownership with the business. Each use case had an accountable executive, an operational owner, and named partners in technology, data, risk, and security. A small central group built reusable patterns for evaluation, access controls, logging, and vendor review. That avoided two failures: every team rebuilding the same capability, or a central AI team making decisions without understanding the work.

What is the strategy for scaling from one use case to many?

Scale the learning loop before you scale the automation. Reuse context patterns, evaluation methods, approval rules, and integration standards. Prioritize the next use case by value, data readiness, risk, and willingness of the operating team to participate. We did not move to a new workflow until we could explain the result of the current one, including where it failed and what we changed because of those failures.

What advice would you give a business leader starting today?

Choose a real workflow, not an abstract ambition. Put an operator and a process owner in the room from the first day. Build an evaluation set before you pick a model. Make the first system assistive, evidence-based, and easy to reverse. Then use the results to decide where autonomy is justified. AI strategy becomes credible when it has an operating rhythm: owners, measures, controls, reviews, and a way to learn.