WRWriting

Building an Enterprise AI Control Plane: The Architecture Beyond the Model

A practical enterprise architecture for governing identity, policy, model routing, tools, evaluation, cost, telemetry, and audit across public, private-cloud, and local AI.

AI / Enterprise ArchitectureAugust 10, 202621 min read

Enterprise applications and agents enter one AI control plane that governs identity, policy, models, routing, tools, evaluation, telemetry, and audit across public, private-cloud, and local execution.

Multi-model AI creates a new enterprise architecture problem.

Once applications can choose among public APIs, private-cloud models, and local inference, enterprises need a control plane that governs models, data, tools, cost, and execution without forcing every application team to build those controls themselves.

For the first generation of enterprise generative AI applications, the architecture was relatively straightforward.

An application called a model API. The model returned a response. Authentication, logging, and perhaps some basic content filtering surrounded the interaction.

That architecture works when there is one provider, a small number of models, limited tool access, and relatively simple governance requirements.

It becomes increasingly difficult to defend once the enterprise AI environment includes:

  • multiple model providers and dozens of available models;
  • public, private-cloud, and locally hosted execution;
  • different data classifications and business-unit policies;
  • agents invoking tools and enterprise systems;
  • rapidly changing model prices;
  • different latency and quality requirements;
  • regulatory and geographic restrictions;
  • model evaluations that change over time.

At that point, model selection can no longer remain buried inside application code.

Neither can governance. Neither can cost management. Neither can tool authorization.

The enterprise needs an architectural layer between applications and AI execution that makes those decisions consistently.

That layer is the enterprise AI control plane.

The strategic layer in enterprise AI will not be the model. It will be the control plane that decides what intelligence can be used, what data it can see, what tools it can invoke, what it may cost, and what evidence it must leave behind.

From model integration to AI infrastructure

The first article in this series argued that enterprises are moving beyond the idea of choosing a single AI model.

The second introduced right-sizing intelligence: organizations should stop sending every workload to their most capable—and frequently most expensive—model.

The third addressed the governance problem. Before a router optimizes for cost, quality, or latency, policy must determine which models are eligible to receive the request at all.

Those three ideas lead naturally to an architectural conclusion.

If model choice is dynamic, routing needs to be centralized.

If cost matters, usage and outcomes need to be measured centrally.

If governance determines model eligibility, policies need to be enforced centrally.

And if AI agents can act on enterprise systems, tool authorization needs to be controlled centrally.

The application should not have to understand all of this.

An application should be able to say, in effect:

I have an authenticated finance user asking for a contract analysis. The documents are classified Confidential. I need high-quality reasoning, the response must remain within an approved geography, and this task has a defined cost ceiling.

The control plane should determine what happens next.

That is a very different architecture from hardcoding a model name into an SDK call.

The architecture

At a high level, the enterprise AI control plane sits between applications and the environments where AI execution occurs.

                     ENTERPRISE AI

          Applications              Agents
               |                      |
               +----------+-----------+
                          v
                     AI GATEWAY
                          |
                IDENTITY + CONTEXT
                          |
                     POLICY ENGINE
                          |
                     MODEL REGISTRY
                          |
                  INTELLIGENT ROUTER
                          |
           +--------------+--------------+
           v              v              v
        PUBLIC         PRIVATE          LOCAL
          AI              AI              AI
           |              |              |
           +--------------+--------------+
                          |
                    TOOLS / MCP / DATA
                          |
                          v
        +----------------------------------+
        | Evaluation  | Cost   | Security  |
        | Telemetry   | Audit  | Outcomes  |
        +----------------------------------+
                          |
                          v
                 ROUTING POLICY UPDATE

The important part is not any individual component.

It is the separation of concerns.

Applications express intent.

Policy establishes boundaries.

The registry describes available intelligence.

The router selects execution.

The tool layer controls actions.

Evaluation determines whether the result was acceptable.

Telemetry measures what happened.

The audit layer preserves why it happened.

Together, those components turn a collection of model APIs into governed enterprise infrastructure.

The AI gateway: one entrance to enterprise intelligence

The first component is deceptively simple: a gateway.

Rather than allowing every application to integrate independently with every AI provider, requests enter through a common enterprise interface.

This is analogous to what API gateways did for service-oriented architectures.

The gateway can standardize:

  • authentication;
  • request identifiers;
  • rate limits and quotas;
  • logging;
  • streaming behavior;
  • retries and provider failover;
  • request metadata;
  • basic safety controls.

More importantly, the gateway creates a point where enterprise context can be attached before the request reaches a model.

Consider two identical prompts:

Summarize this document.

One originates from a marketing employee summarizing a public press release.

The other originates from legal summarizing a confidential acquisition document.

The text instruction may be identical.

The governance decision should not be.

A useful gateway therefore needs more than the prompt. It needs context.

Identity and context: the request is more than the prompt

One mistake enterprises can make is attempting to determine policy from prompt content alone.

The prompt rarely contains enough information.

The control plane should know:

  • who made the request;
  • which application and business unit originated it;
  • what role the user holds;
  • what data classification applies;
  • which geography the user and data belong to;
  • whether the request is interactive or automated;
  • which tools the application is authorized to use;
  • what cost and latency objectives apply.

That context turns an anonymous inference call into an enterprise transaction.

Request ID:        req-847291
User:              security-analyst-184
Business Unit:     Cybersecurity
Application:       Incident Assistant
Task:              Security reasoning
Data Class:        Confidential
Region:            United States
Latency Target:    < 5 seconds
Max Cost:          $0.25
Tool Profile:      incident-readonly
Audit Level:       Full

Now the system has enough information to make a meaningful decision.

Identity is therefore not merely an authentication concern.

Identity becomes an input into AI routing.

The policy engine: decide what is allowed before deciding what is best

This is the principle established in Article 3:

Govern first. Optimize second.

Suppose the enterprise has access to 40 models.

The router should not simply score all 40 on price, latency, and benchmark performance.

Policy may immediately eliminate 32 of them.

Perhaps Confidential data can only be processed by providers with zero-retention agreements. Restricted data may not leave a private-cloud environment. European customer information may have to execute within an approved EU region. Source code may be limited to three approved coding models.

The policy engine converts those requirements into an eligible execution set.

40 available models
        |
        v Enterprise policy
18 enterprise-approved
        |
        v Business-unit policy
11 approved for this business unit
        |
        v Application policy
6 approved for this application
        |
        v Request context
3 eligible for this request
        |
        v Optimization
1 selected

That ordering matters enormously.

A model that costs 90% less is irrelevant if it is prohibited from processing the data.

A model with better benchmark performance is irrelevant if it violates residency requirements.

A provider experiencing an outage is irrelevant even if policy permits it.

Governance defines the solution space.

Optimization operates inside it.

The model registry: models need enterprise metadata

Dynamic routing requires more than a list of model names.

The control plane needs a registry describing the operational characteristics of every available model.

A registry might contain:

  • model, provider, and version;
  • capability class and context window;
  • input and output cost;
  • expected latency;
  • supported regions;
  • data-retention mode;
  • approved data classifications;
  • tool-use and structured-output support;
  • evaluation scores;
  • security approval and availability status;
  • deprecation date.

This turns models into governed enterprise resources.

The registry also addresses a problem that will become increasingly important: models change.

Providers release new versions. Pricing changes. Context windows expand. Models are deprecated. Evaluation results improve or deteriorate. New regions become available.

An application should not need to be redeployed because a better model became available Tuesday morning.

The registry and router should absorb that change.

That is one of the larger architectural advantages of abstraction.

Applications become less dependent on individual model lifecycles.

The intelligent router: choosing among eligible intelligence

Only after governance has produced an eligible model set should optimization begin.

The router can evaluate:

  • Capability: Can this model reliably perform the task?
  • Quality: Has it met the evaluation threshold for this workload?
  • Cost: What is the expected cost of completing the request?
  • Latency: Can it meet the application's response-time requirement?
  • Availability: Is the provider healthy right now?
  • Capacity: Is local GPU capacity available?
  • Context: Can the model accommodate the required input?
  • Historical performance: How has it performed on similar requests?

This is where right-sizing intelligence becomes operational.

A simple extraction task may route to an efficient model. A customer-support summary may route to a balanced model. A difficult security investigation may route to a premium reasoning model. A highly sensitive legal document may route locally regardless of whether a cheaper public model exists.

Current platforms already expose parts of this pattern. Microsoft Foundry Model Router selects among eligible models using modes that balance quality and cost, while OpenRouter workspaces combine routing defaults, guardrails, budgets, and observability.

Governance narrows the model universe to eligible choices before an intelligent router scores capability, quality, cost, latency, availability, and evaluation history.

Escalation routing

Routing does not have to be static.

A useful pattern is to start with an efficient model, evaluate the result, and escalate when confidence or quality falls below a defined threshold.

Efficient Model
      |
      v
Evaluation
      |
   Pass? -- Yes --> Return result
      |
      No
      v
Balanced Model
      |
      v
Evaluation
      |
   Pass? -- Yes --> Return result
      |
      No
      v
Premium Reasoning Model

Now model cost becomes proportional to problem difficulty rather than application identity.

That is much closer to how mature cloud infrastructure is operated.

Provider adapters: abstraction without pretending models are identical

There is an important trap in building a control plane: abstraction can go too far.

Models are not interchangeable compute instances.

Providers expose different capabilities, including reasoning controls, structured outputs, tool-calling formats, multimodal inputs, caching, batch inference, context management, safety controls, and streaming protocols.

The control plane therefore needs provider adapters.

The application can interact with a common enterprise interface while the adapter translates the request into the capabilities of the selected provider.

But the abstraction should not hide meaningful differences.

If an application requires vision, computer use, deterministic structured output, or a particular tool-calling capability, those requirements become part of the routing request.

The goal is not to pretend every model is identical.

The goal is to prevent every application from independently solving provider integration, governance, telemetry, and failover.

Public, private, and local become execution targets

Once model selection is abstracted, the distinction between public API, private cloud, and local inference becomes another routing dimension.

                    ELIGIBLE REQUEST
                           |
              +------------+------------+
              v            v            v
         PUBLIC AI     PRIVATE CLOUD   LOCAL AI
              |            |            |
        API providers   managed AI    enterprise
                                     GPU infrastructure

A public API may provide the best economics and capability for one request. A managed private-cloud service may satisfy stronger controls for another. A locally hosted model may be required for highly restricted data or predictable high-volume workloads.

This creates an interesting possibility:

Workloads can move without applications being rewritten.

A company may begin with public APIs while demand is uncertain. Later, a high-volume workload may become economical to run locally. Or regulation may require a workload to move into a controlled environment.

If applications communicate with the control plane rather than directly with the execution environment, those changes become infrastructure decisions instead of application rewrites.

Tool governance: the model is only half the risk

This is where enterprise AI architecture becomes more consequential.

A model generating text is one thing.

An agent capable of acting is another.

Modern AI applications may query databases, retrieve documents, search repositories, create tickets, inspect cloud infrastructure, execute code, modify records, send messages, and invoke APIs.

MCP provides a standard client-server architecture for exposing tools, resources, and prompts to AI applications. Its authorization guidance protects sensitive operations using standardized authorization flows and recommends separating access by tool or capability where possible.

That is powerful.

It also means the governance boundary cannot stop at model selection.

Consider an incident-response agent.

It may reasonably need permission to:

  • query security telemetry;
  • retrieve CloudTrail events;
  • inspect IAM configuration;
  • search threat intelligence.

But should it be allowed to disable an IAM user, delete an access key, terminate an instance, modify a firewall, or send an external email?

Those are different authorization decisions.

The control plane therefore needs a tool registry and tool-policy layer.

Tool:              iam-disable-user
System:            AWS IAM
Risk Level:        High
Action Type:       Write / destructive
Allowed Roles:     Incident Commander
Human Approval:    Required
Permitted Apps:    IR Assistant
Audit Level:       Full

This leads to a crucial architectural principle:

Model authorization and action authorization should be separate decisions.

A model may be authorized to reason about an incident without being authorized to remediate it.

An agent may be authorized to recommend an action without being authorized to execute it.

AWS's Agentic AI guidance makes this separation explicit: authorize each tool invocation externally, propagate identity and user context, block and log unauthorized calls, and place human checkpoints in front of high-risk mutations.

Model authorization determines whether AI may reason about a task; a separate tool-policy engine decides whether proposed read, write, or destructive actions may execute.

Human approval should be a policy primitive

“Human in the loop” is often discussed as though every AI action should require someone to click Approve.

That does not scale.

The better approach is risk-based approval.

Low-risk actions can execute automatically. Medium-risk actions may require approval depending on context. High-impact or irreversible actions require explicit authorization.

Read CloudWatch logs              -> Automatic
Query customer record             -> Automatic if authorized
Create Jira incident              -> Automatic
Restart development service       -> Policy dependent
Disable production IAM principal  -> Human approval
Delete production database        -> Prohibited

AWS similarly recommends risk-tiered human review, warning that reviewing every action creates fatigue while reviewing none creates unbounded autonomy.

The approval requirement belongs in policy, not in application-specific prompt instructions.

“Ask the user before doing something dangerous” is not an enterprise authorization architecture.

The system should technically prevent execution until the required approval exists.

Evaluation: successful execution does not mean successful work

Traditional infrastructure telemetry tells us whether a request succeeded technically.

HTTP 200
Latency: 1.4 seconds
Tokens: 8,240
Cost: $0.07

That tells us almost nothing about whether the AI did useful work.

An enterprise control plane therefore needs an evaluation layer.

Depending on the workload, evaluation might measure:

  • factual accuracy;
  • structured-output validity;
  • citation correctness;
  • task completion;
  • hallucination rate;
  • human acceptance;
  • escalation rate;
  • security-policy adherence;
  • business outcome.

This connects directly to the FinOps argument from Article 2.

If Model A costs $0.03 per request but succeeds 60% of the time, while Model B costs $0.08 but succeeds 96% of the time, token cost alone gives an incomplete answer.

The meaningful metric is closer to:

Cost per successful business outcome.

Evaluation feeds routing. Routing generates outcomes. Outcomes generate evaluation data. Evaluation data changes future routing.

The architecture becomes a feedback system.

Telemetry: AI needs its own observability model

For each request, the control plane should be capable of capturing appropriate telemetry such as:

  • request ID, application, and user or service identity;
  • task and data classification;
  • model candidates and selected provider;
  • policy decision and execution environment;
  • latency, tokens, cost, and cache utilization;
  • tool calls and fallback events;
  • evaluation score, errors, and outcome.

This enables questions that will matter increasingly to technology and finance leadership:

  • Which applications consume the most AI spend?
  • Which models produce the highest successful-outcome rate?
  • How frequently are requests escalated to premium models?
  • Which workloads could move to less expensive models?
  • Which providers are failing latency objectives?
  • How much demand is served locally versus externally?
  • Which tools are agents invoking most frequently?
  • How much AI spend produces results users reject?

Without a common telemetry layer, every application answers those questions differently—or cannot answer them at all.

The decision audit trail

Governance becomes especially important when someone asks months after execution:

Why was this data sent to that model?

A mature AI control plane should be able to reconstruct the decision.

Request:              req-847291
Timestamp:            2026-08-10T14:32:18Z
Identity:             security-analyst-184
Application:          Incident Assistant
Data Classification: Confidential

Enterprise Policy:    AI-POLICY-2026.08 v4
Business Unit Policy: SEC-AI-12 v7
Application Policy:   IR-ASSISTANT v11

Eligible Models:      model-b, model-f, private-r1
Selected Model:       private-r1
Selection Reason:     quality + confidentiality + latency

Execution Region:     us-east
Retention:            metadata-only
Tools Invoked:        cloudtrail-search, iam-read
Human Approval:       not required

Cost:                 $0.14
Evaluation:           passed
Outcome:              analyst accepted

That is dramatically more useful than a log saying an API request completed successfully.

It creates the foundation for security investigations, compliance evidence, cost analysis, model-performance analysis, and governance reviews.

The record should be tamper-resistant and retained according to policy.

The objective is simple:

If you cannot reconstruct the decision later, you do not have governance.

A request moving through the control plane

Consider a practical example.

A security analyst asks:

Analyze the last 24 hours of IAM activity for this user and identify whether the sequence is consistent with credential compromise.

The request enters through the organization's Incident Assistant.

Step 1: Identity

The gateway authenticates the analyst and determines that they belong to the Security Operations team.

Step 2: Context

The application classifies the request as Security investigation / Confidential / US execution required.

Step 3: Governance

Enterprise policy eliminates models that do not meet the organization's retention requirements. Security policy restricts the set to models approved for security telemetry. Application policy requires tool calls to remain read-only.

Step 4: Routing

Three models remain eligible. The router determines that the task requires advanced reasoning and selects the model with the best historical evaluation score within the application's cost and latency thresholds.

Step 5: Tool authorization

The model requests cloudtrail-search: approved.

It then requests iam-get-user: approved.

If it attempts iam-delete-access-key, the policy rejects the call because the Incident Assistant has read-only authorization.

The model can recommend disabling the key.

It cannot perform the action.

Step 6: Evaluation

The response is evaluated for evidence references and whether its conclusions are grounded in retrieved events.

Step 7: Outcome

The analyst accepts the finding and escalates the incident.

Step 8: Audit

The control plane records the identity, policies, model, region, tool calls, evaluation, cost, and outcome.

One request has now exercised nearly every part of the control plane.

Critically, the application developer did not have to implement those controls individually.

A closed-loop AI control plane governs, routes, executes, evaluates, observes, learns, and updates routing policy with human oversight.

The control plane should not become the new monolith

There is a danger in this architecture.

Once an enterprise realizes how much functionality belongs in the control plane, there is a temptation to build an enormous centralized platform before allowing anyone to deploy AI.

That would reproduce an old enterprise technology failure in a new domain.

The control plane should begin with a small number of high-value primitives:

Gateway. Policy. Registry. Routing. Telemetry.

Then add sophisticated evaluation, tool governance, automated optimization, and additional execution environments as requirements mature.

The goal is not centralization for its own sake.

The goal is to centralize decisions that must be consistent while leaving application teams free to innovate above them.

Build the paved road, not the roadblock

The best enterprise platforms win adoption because the governed path is easier than bypassing it.

An application team should prefer the enterprise AI gateway because it gives them:

  • immediate access to approved models;
  • provider abstraction and centralized credentials;
  • automatic telemetry and cost reporting;
  • retries and failover;
  • evaluation infrastructure;
  • approved tools;
  • compliance controls.

If using the control plane creates three months of architecture reviews while calling a model API directly takes 15 minutes, developers will find ways around the platform.

Governance works best when the secure path is also the fastest path.

This is the same lesson enterprises learned with cloud platforms.

Successful cloud platform teams did not merely prohibit developers from building infrastructure incorrectly.

They created paved roads that made building it correctly easier.

AI platforms need the same philosophy.

The control plane becomes the optimization surface

Once this architecture exists, something more interesting happens.

The enterprise can begin optimizing AI globally rather than application by application.

Suppose evaluation data shows that an efficient model now performs within 2% of the premium model on customer-ticket classification. The routing policy can shift that workload. Every application using the capability benefits.

Suppose a provider increases pricing. Routing can adjust.

Suppose a new model passes security review and dramatically outperforms the incumbent on code analysis. Add it to the registry.

Suppose local inference capacity is underutilized overnight. Batch workloads can route there.

Suppose a provider has a regional outage. Traffic can fail over to another eligible model.

Suppose evaluation reveals that a cheap model's lower quality causes excessive retries. The router can stop using it despite its attractive token price.

This is why the control plane becomes strategically important.

It creates an enterprise-wide optimization surface across capability, cost, security, latency, availability, and governance.

What enterprises should build versus buy

Not every organization should build an AI control plane from scratch.

In fact, most probably should not.

Cloud providers, model gateways, AI platforms, observability vendors, security products, and emerging infrastructure companies already provide portions of this architecture.

The enterprise architectural task is less about writing every component and more about defining where the control boundary belongs.

Organizations should retain control over the policies and evidence that matter to them:

  • identity;
  • data classification;
  • provider approval and model eligibility;
  • tool authorization;
  • cost policy;
  • evaluation criteria;
  • audit requirements.

The underlying gateway, router, inference platform, or telemetry system may be commercial, open source, cloud-native, internally built, or some combination.

The important thing is that enterprise governance does not disappear simply because routing has been outsourced.

From AI applications to AI infrastructure

This series began with a simple observation:

The question “Which AI model should we use?” is becoming less useful.

As the model ecosystem expands, enterprises will use multiple models.

Once they do, economics demands right-sizing.

Once routing becomes dynamic, governance must determine which choices are permissible.

And once applications, models, data, and tools begin interacting dynamically, enterprises need an architecture capable of making and recording those decisions.

That is the AI control plane.

2023–2024
APPLICATION → MODEL

2025–2026
APPLICATION → MULTIPLE MODEL PROVIDERS

NEXT
APPLICATION
     ↓
AI CONTROL PLANE
     ↓
POLICY + ROUTING + EVALUATION + TOOLS
     ↓
PUBLIC + PRIVATE + LOCAL INTELLIGENCE

We are moving from integrating models into applications to operating intelligence as enterprise infrastructure.

And infrastructure eventually demands the same disciplines we learned in cloud: identity, policy, observability, security, cost management, resilience, governance, and continuous optimization.

The models will keep changing.

The best model six months from now may not be the best model today. Providers will change. Prices will change. Local inference economics will change. New agent capabilities will appear. Regulatory requirements will evolve.

Enterprise applications should not have to be redesigned every time that happens.

They need a stable layer between business intent and a rapidly changing intelligence market.

The model is becoming an execution resource. The control plane is becoming the architecture.

Series conclusion

Across these four articles, we have moved through four layers of the emerging enterprise AI stack:

  1. Beyond One Model established why model portfolios are replacing one-model architectures.
  2. Right-Sizing Intelligence established why models should be selected according to workload economics rather than habit.
  3. The AI Governance Gap established why policy must constrain model selection before optimization begins.
  4. Building an Enterprise AI Control Plane brings those ideas together into an operating architecture.

The next challenge is implementation.

Not another chatbot.

Not another isolated proof of concept.

But the infrastructure that allows hundreds of AI-enabled applications and agents to operate across multiple models and execution environments while remaining observable, economical, secure, and governable.

That may turn out to be one of the most important enterprise architecture problems of the AI era.

References