“People help AI become more useful by bringing the ambition, context, judgment and feedback that teach it how work gets done and what good looks like.”

Kathleen Hogan, EVP and Chief Strategy and Transformation Officer, Microsoft

Redeployed is a weekly newsletter that breaks down one important AI story at a time for product builders and engineering leaders. Every issue explores what the shift means for technology companies and how leaders can respond.

Every few months, technology teams face another round of model decisions. OpenAI releases something new. Anthropic improves Claude. Google pushes Gemini forward. Prices change, benchmarks move, and the model that looked like the obvious choice six months ago suddenly has serious competition.

Microsoft thinks companies should design for that instability.

Last week, the company released a new enterprise AI playbook based on lessons from hundreds of AI transformation efforts across Microsoft. It recommends redesigning workflows, creating shared data foundations, defining where people and agents make decisions, and building company-specific evaluations and feedback loops before scaling AI across the organization.

There is a bigger strategic idea underneath those recommendations. Microsoft argues that companies should avoid making any one foundation model the center of their architecture. The proprietary value sits higher in the stack, inside the context, evaluations, workflows, institutional knowledge, and controls companies build for themselves.

Models Keep Changing

Building an AI strategy around a particular model creates a structural problem. A company might choose one provider for coding, another for cost, or another for a specialized reasoning task, but those relative advantages can change quickly.

Microsoft's playbook reflects that reality. Its architecture places foundation models underneath a company-owned layer containing enterprise knowledge, memory, tools, evaluations, orchestration, security, and governance. The underlying models can then change without forcing the company to rebuild everything above them.

For engineering teams, this makes model flexibility a useful architectural property. Applications can route tasks differently as model capabilities and economics change while preserving the company-specific systems that determine how AI actually works inside the business.

Define What Good Looks Like

One of the most interesting parts of Microsoft's approach is its emphasis on private evaluations.

Public benchmarks can tell you whether a model is good at coding, reasoning, mathematics, or tool use. They cannot tell you whether an agent writes a customer response that sounds right for your company, follows your risk standards, handles an unusual internal process correctly, or makes the same tradeoffs an experienced employee would. Companies have to define those standards themselves.

Microsoft describes this proprietary intelligence as a combination of the company's market perspective, data, workflows, institutional knowledge, quality standards, and risk boundaries. Those become evaluation criteria that agents can be measured against repeatedly.

Over time, those evaluations can become a detailed definition of how the organization expects AI to behave. A new model can be plugged into the system and tested against the same standards before receiving more responsibility. That makes the company's definition of "good" an asset that survives model changes.

Redesign the Workflow First

Microsoft also found that simply adding AI to existing processes produced limited results.

Its cloud supply-chain team started by mapping and simplifying workflows, then created a shared source of data before deploying more than 100 agents across planning, sourcing, fulfillment, and logistics. In five monthly planning cycles measured between April and August 2026, Microsoft says average cycle time in selected workflows fell from roughly 10 business days to less than 2.5.

The sequence matters because AI can accelerate a poorly designed process without fixing the reasons that process is inefficient.

Companies have already spent years building workflows around human constraints. Those workflows contain approval chains, handoffs, reporting requirements, and information flows designed for people doing most of the work manually. Adding agents gives teams an opportunity to reconsider how the process should operate when software can handle parts of the work continuously.

Microsoft's results suggest that workflow design deserves as much attention as model selection.

The Proprietary Layer Compounds

The most durable part of an enterprise AI system may be everything the organization learns while operating it.

Every correction can improve an evaluation. Every failed workflow can expose missing context. Every human intervention can clarify where an agent needs approval. Every successful deployment can produce patterns that other teams reuse.

Microsoft describes this as a continuous learning loop where people provide the context, judgment, and feedback that make AI more useful inside the organization. Over time, those lessons accumulate into institutional knowledge that the company controls.

That creates an advantage a model provider cannot simply deliver through an API. Two companies can use the same foundation model while getting very different results because one has spent years refining how AI operates inside the business.

The model provides intelligence. The organization determines how effectively that intelligence gets applied.

This issue of Redeployed is brought to you by Tecla Labs: As foundation models become easier to swap, more of the long-term value sits in the systems companies build around them. Tecla Labs helps companies turn AI into working business systems by identifying high-value workflows, designing the system around them, and building the integrations, context, and infrastructure needed to make AI useful in practice. Start with a free AI Assessment to identify where AI can create the most leverage inside your business.

Model Independence Has a Cost

A model-agnostic architecture can become its own engineering project. Abstraction layers, routing systems, evaluation infrastructure, shared context, and internal platforms all require people to build and maintain them.

There is also a risk of abstracting away capabilities that make a particular model valuable. Providers increasingly differentiate through tool use, memory, computer use, context windows, agent infrastructure, and other features that do not always map neatly across models.

Companies need enough flexibility to take advantage of a changing market without building an elaborate internal platform simply to avoid choosing a provider.

Evaluations introduce another risk. A private benchmark is only useful when it measures outcomes the business actually cares about. Poorly designed evals can turn assumptions into permanent infrastructure and reward agents for optimizing the wrong behavior.

Turn Model Selection Into a Repeatable Test

Private evaluations change how engineering teams can approach model selection. Instead of reopening an abstract benchmark debate every time a provider releases an update, teams can test new models against the workflows and quality standards that matter inside their own company.

A cheaper model may be sufficient for one task while a more capable model earns its cost somewhere else. Another provider may outperform both on a specialized workflow. The point is not to avoid choosing models. It is to make those choices against a stable company-owned standard rather than rebuilding the decision process around every new leaderboard.

This also gives teams a more practical way to respond as the model market changes. New releases can be evaluated against the same internal standards, making it easier to determine whether switching actually improves the system rather than reacting to public benchmarks alone.

Build What Survives the Model Change

Foundation models will keep improving, and leadership in the model market will continue to move. Companies cannot control that cycle. They can control what accumulates inside their own systems.

Microsoft's experience points toward a useful division of responsibility. Providers keep improving the underlying intelligence while companies build the systems and learning loops that make that intelligence valuable to their business.

Those assets become more useful with every deployment because they capture how the company operates and what it considers a good outcome.

Companies that build that layer well should be better positioned to take advantage of new models without rebuilding their AI strategy every time the leaderboard changes.

Connect With Other Technology Leaders

If you want to exchange practical ideas with senior technology and product leaders navigating the AI era, check out the upcoming GILD Forums. They bring together experienced operators for peer discussions around the technology and business decisions they are working through right now.

More to come…

Gino Ferrand, Founder @ Tecla Labs