Dr.Hani Tiếng Việt
Applied Knowledge Market Notes

Do Not Measure AI by Usage Alone; Measure Workflow Capability

4 min readAssoc. Prof. Nguyen Hai Ninh
Hình ảnh dữ liệu số trừu tượng đại diện cho vận hành AI trong doanh nghiệp

In an article dated 1 September 2026, OpenAI states that the top 10% of firms by AI usage generate 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January. This is provider-reported data, so it cannot by itself represent every market or demonstrate business effectiveness. Even so, the described pattern deserves management attention: the difference is not only whether employees use a tool more, but whether an organisation turns a successful workflow into repeatable practice.

Tokens, conversation volume, or provisioned accounts can signal access to a tool. They do not tell us whether a workflow produces better, faster, or lower-risk outputs. When access metrics stand in for work results, an AI programme can move to either extreme: expanding quickly because use rises, or stopping too early because users do not create visible value in their first attempt.

From individual activity to workflow capability

OpenAI’s article presents three cases: Basis places onboarding into a skill that can be updated from exceptions; Clay maintains context for each account; and Exa combines information sources, tests, and human review before release. These are provider-selected business examples, not a representative survey. They nevertheless point to a common operating logic: AI creates more durable value when work has a trigger, known inputs, permission boundaries, a testable output, and a way to learn from exceptions.

The required shift is in the unit of management. Instead of asking, “How much has the team used AI?”, a manager can ask, “Which workflow now has a part of the work completed faster while quality is still checked?” The latter question brings the discussion to ownership, input data, stopping points, and evaluation criteria. It also prevents an organisation from treating a fast draft as a result ready to send to a customer or use in a decision.

Level to monitor Easy but insufficient metric Evidence closer to value
Access Provisioned accounts, visits, and prompts. Which user groups complete a defined job through an approved flow.
Workflow Number of drafts or tokens produced. Cycle time, rework errors, review load, and exception rate.
Outcome A feeling of “working faster”. Output quality against a predetermined standard, cost, revenue, or related risk.
Scalability Number of new tools or integrations. A workflow with documented owner, guidance, access, evidence, and improvement rhythm.

Design an experiment from which the team can learn

A suitable experiment does not have to begin as a large project. Choose a frequently occurring job, with an output that the responsible person can assess, and enough friction for improvement to matter. For example, a consulting team could test an assistant that creates a client-brief summary before an internal meeting. Before beginning, the group needs to settle what information may be used, who reviews it, what makes the summary usable, and when content must not leave the organisation.

During the first few weeks, measure both efficiency and the cost of control. If preparation time falls but time spent revising drafts rises, the benefit is not yet clear. If some cases must stop because data or access are missing, that is not automatically a user error. It is information for adding context, refining guidance, or retaining a step for human decision.

Example: turning a sales assistant into a controlled workflow

A B2B business wants to use AI to prepare opportunity dossiers. Rather than giving the whole team one general prompt, it defines the trigger as an opportunity moving to solution evaluation. The assistant may use selected CRM data and public customer material; it must create a one-page output containing the customer problem, source evidence, outstanding questions, and commercial risks. The salesperson checks that page before contacting the customer. Each month, the team leader reviews preparation time, data accuracy, user feedback, and stopped cases. Only once the workflow is stable does the business decide whether to extend it to another account group.

Decision rights and evidence must grow with consequence

An assistant that summarises internal documents and an agent that proposes a price or sends an external message do not have the same consequences. The closer a job is to a commitment, sensitive data, or a customer-affecting decision, the more clearly the team needs to define data sources, access rights, evidence the system must show, and the person authorised to approve. This clarity does not slow work needlessly; it lets the team know what can be delegated, what must be reviewed, and why an exception must stop.

The 8.3-times figure should therefore not be read as a usage target to chase. It suggests a more practical question: does the organisation have a few workflows clear enough for AI to perform part of the work within limits that can be observed, measured, and improved? If not, the immediate goal is not more tools. It is to choose one value surface, describe the job, set an output standard, and establish a learning rhythm from real cases.

References

OpenAI. (2026, September 1). How AI-native companies turn workflows into operating capability. https://openai.com/index/ai-native-company-workflows/

Continue reading

Related insights