In many organisations, the first phase of AI adoption is described with a simple scorecard: how many accounts have been activated, how many people use it weekly, and whether cost is rising or falling. Those measures are necessary for budget control and for knowing whether a programme is being taken up. They do not, however, answer the question an executive team will eventually ask: which work has changed, has quality been maintained, and where has the time or cost capacity created actually gone?
On September 16, OpenAI announced analytics capabilities in the ChatGPT Admin Console intended to connect usage and cost data, task insights, and outcome indicators across ChatGPT Work and Codex. According to the company’s description, administrators can see which groups use the tools, which tasks consume credits, where support or training may be needed, and, for engineering, contributions to merged commits and lines of code alongside review activity. This signals a market shift in enterprise AI tools: from the story of “granting access” to the story of “observing work and assessing value.”
Separate the fact from the expectation
A provider announcement is information about a product and how the provider proposes that customers use it; it is not evidence that every organisation will earn a return from AI. OpenAI itself presents the sales ROI example in the article as illustrative and notes that results depend on review and correction time, setup cost, training, and ongoing support. The important point is not an assumed ROI percentage but the measurement logic: begin with a common task, establish a baseline, test the change over a defined period, and then track which business result uses the capacity released.
That logic differs from making total activity or prompt volume a success KPI. A team may use AI heavily to create more drafts while not reducing preparation time, not improving quality, or simply shifting review work to someone else. Conversely, a workflow used less often can be valuable if it reduces errors in a risk-control step, shortens customer response time, or helps a manager see a problem earlier. Usage is a signal for investigation; it is not yet a conclusion about effectiveness.
| Layer to observe | Management question | Example indicator |
|---|---|---|
| Adoption | Who is using it, where, and for which task? | Active users by group; task type; tools used. |
| Process quality | How does AI change time, errors, or rework? | Time to prepare a brief; rework rate; review time. |
| Work outcome | Is the task output better? | Account-plan quality; ticket resolution time; release time. |
| Business value | What meaningful result does that outcome affect? | Qualified-opportunity rate; customer retention; cost to serve. |
A new risk: measuring more while learning less
When a platform can classify tasks, track tools, and combine them with cost, an organisation can easily expand a dashboard before agreeing its management purpose. If data are used mainly to rank individuals or cut cost, employees may avoid difficult tasks, move work outside the system, or optimise the number of uses instead of output quality. Trust falls, and the data lose meaning. Any measurement programme needs to state how data will be used to learn from and improve workflows, who may see it at an individual level, and which boundaries should not be crossed.
Connecting usage data with operational data does not automatically prove causation. If a sales team uses AI more in a quarter when revenue rises, the change could also reflect seasonality, lead quality, pricing changes, or staffing. A company does not need to turn every pilot into an academic study, but it does need a fair comparison: the same type of task, a clear time window, quality standards before and after, and recorded time for correction and checking. For an important workflow, a small pilot group or phased rollout produces better evidence than a broad launch followed by inference from correlation.
Example: assessing AI for customer-meeting preparation
A sales team uses AI to summarise account history and draft talking points before meetings. Rather than reporting that “the team created 900 prompts,” the manager selects one workflow: preparing briefs for strategic accounts. They measure current preparation time, randomly check brief quality against agreed criteria, record revision rounds, and track what the saved time is used for. After six weeks, they examine whether the number of well-prepared conversations rises, whether more opportunities have a clear next step, or whether the sales cycle changes. If results are not as expected, the next question may concern CRM data quality, prompting skills, or how released time is allocated — not simply an increase in the AI usage quota.
Three actions managers can take now
First, select one task that occurs frequently enough and connects directly to a business priority instead of trying to measure all AI across the organisation. Second, name the owner of the work outcome. Technology or AI-enablement teams can help design the tools and interpret data, but the sales, operations, or service leader knows which result has real value. Third, agree the full cost from the start: tool fees, setup time, training time, checking time, and risks that require control. Without this, “time saved” can easily become an attractive number that cannot be converted into value.
The enterprise AI market is adding a new observation layer to day-to-day work. That can help managers avoid two extremes: believing AI creates value automatically because it is used often, or rejecting it because a single aggregate number is not immediately visible. A more useful question is specific: in which workflow, for whom, and to which quality standard is AI helping us do what better — and which result will use that new capacity?
Reference
OpenAI. (2026, September 16). How to connect AI usage to business value.
