Architecture Ho Bae

External LLM Approval Checklist for Sensitive Data

External LLM sensitive data approval shown as a dotted protected path between enterprise documents and an approved model service.

Teams rarely approve an AI model in the abstract. They approve a particular way of accessing it: a consumer chat application, an enterprise workspace, a developer API, or a managed cloud service. An external LLM approval checklist makes that distinction explicit before sensitive data enters the workflow. The model brand may be the same, but the data contract is not.

An external LLM sensitive data review must therefore name the exact product surface, account, endpoint, and enabled features. It must also separate what the provider controls from what the enterprise can keep inside its own environment. The goal is not to declare one vendor safe. It is to approve one bounded configuration for one real workflow.

Key takeaways

  • Approve the exact service surface and configuration, not only the model brand.
  • Treat training, retention, abuse review, application state, files, caches, and tools as separate questions.
  • Keep original values and the protected mapping inside the customer-controlled environment; send only an evaluated protected working version through the approved model path.
  • Verify provider-side state and customer-side logs, retries, errors, and destinations before approval.
  • Re-run the review whenever the account tier, endpoint, model, feature set, terms, observability stack, or output destination changes.

Same model brand, different data contract

The first review question is not “Which model are we using?” It is “Which service is receiving the request?” A consumer chat application, enterprise workspace, developer API, and managed cloud service can expose different controls for training, history, files, policy enforcement, retention, and deletion.

Provider documentation makes the distinction concrete. OpenAI’s API data controls separate abuse-monitoring logs from application state and document retention by endpoint and feature. Anthropic’s commercial-product policy distinguishes Claude for Work and the Anthropic API from consumer plans. Google’s Gemini API terms distinguish paid and unpaid services and describe different treatment for submitted content.

These documents also show why “not used for training by default” is an incomplete approval statement. A service may still create state for abuse monitoring, conversation history, files, caches, tools, grounding, or product operation. Some controls are account- or feature-dependent. Others require approval or make particular features unavailable.

Record the exact surface before assessing any promise:

  • product and service name;
  • organization, workspace, tenant, or cloud project;
  • consumer, business, enterprise, or developer account type;
  • billing state and region;
  • model and endpoint;
  • files, history, caching, search, grounding, tools, connectors, and background processing; and
  • retention, deletion, and zero-data-retention controls actually enabled.

A procurement record or vendor name cannot substitute for this inventory. Approval attaches to the configured path that handles the request.

Customer-controlled boundary Original values Protected mapping Only an evaluated protected working version leaves this boundary.
Exact approved service configuration Product surface, account, endpoint, model, region, and enabled features
History Files Caches Tools Review and logs
Authorized return path Reconstruction Approved caller and scope The reconstructed output reaches only the approved destination.

What LLM Capsule changes in the external-model path

LLM Capsule is a context-preserving data layer for AI. In its documented target architecture, original values and the protected mapping remain inside the customer-controlled environment. A protected working version follows the approved model path, and Reconstruction reconnects supported values inside the customer environment before the result reaches its authorized destination.

Consider a contract-review workflow that needs a model to compare renewal clauses. The original supplier names, internal project references, and restricted amounts can remain inside. The external service receives a protected working version that preserves the supported distinctions and relationships required for comparison. When the response returns, Reconstruction resolves authorized references for the originating workflow.

This changes what crosses the model route. It does not remove the need to approve that route. The protected working version may still contain confidential language, derived context, or combinations that the organization treats as sensitive. It remains subject to the provider’s service contract and the enterprise’s own data-handling rules.

The deployment review must prove, rather than assume, that:

  • the original record enters the protection step inside the approved environment;
  • the outbound request contains the expected protected representation;
  • the protected mapping is not copied to the model service;
  • Reconstruction is limited to the approved caller, scope, and destination; and
  • logs, traces, retries, and errors do not recreate the original-value path.

What LLM Capsule does not change

LLM Capsule does not control the external provider’s logs, product history, caches, policy review, connected tools, or deletion behavior. It also does not control copies created by the customer’s gateway, queue, tracing platform, support tooling, or downstream application.

OpenAI documents default abuse-monitoring retention as well as feature-specific application state and endpoint eligibility for Modified Abuse Monitoring or Zero Data Retention. Anthropic states that API inputs and outputs are normally deleted within 30 days, while naming exceptions such as persistent services, policy enforcement, legal requirements, and separate agreements. Google’s documentation identifies distinct treatment for paid and unpaid services and feature-specific state for caching, grounding, files, and other capabilities.

Those are descriptions of named services at a point in time, not universal properties of the companies or models. The exact terms and product behavior can change. Reviewers should retain the source URL, retrieval date, account configuration, and evidence captured from the approved environment.

Customer-side copies need the same attention. A protected request may be duplicated in an API gateway, retry queue, failure payload, distributed trace, or support bundle. The reconstructed output may be written to the wrong ticket, document, or user session even when the model call itself follows the intended route.

The boundary is complete only when the team can trace the request from the original record to the final stored result.

Review OpenAI, Anthropic, and Google at the service level

The following matrix is a review aid, not a vendor ranking. Populate it for the exact product and configuration under consideration. Recheck every linked policy immediately before approval because retention periods, feature behavior, and eligibility can change.

Provider scopeExact surface to nameDefaults and exceptions to verifyFeature-specific state to inspectEvidence to retain
OpenAIConsumer ChatGPT, business workspace, or specific API organization and projectTraining setting, abuse-monitoring treatment, ZDR or Modified Abuse Monitoring approval, legal and safety exceptionsResponses storage, conversations, threads, files, vector stores, prompt caching, background mode, hosted tools, and third-party servicesOfficial policy version, organization/project setting, endpoint list, request configuration, and deletion test
AnthropicConsumer Claude plan, Claude for Work or Enterprise, Anthropic Console, or API organizationCommercial versus consumer treatment, feedback or opt-in, standard API retention, contractual ZDR, policy-enforcement and legal exceptionsPersistent chats, coding sessions, Files API, covered-model rules, and any service with customer-controlled retentionOfficial policy version, product/account type, retention agreement, feature inventory, and deletion result
GoogleGemini Apps, Workspace offering, Gemini Developer API, or named managed cloud servicePaid versus unpaid classification, product-improvement treatment, human review, abuse logging, region-specific terms, and ZDR eligibilityAI Studio, caching, files, grounding, interactions, connected apps, request-response logging, and live-session stateOfficial terms version, billing/project status, enabled features, cache and logging configuration, and deletion test

The table deliberately avoids a single “retention” value for each provider. One number would hide the decision that matters: which state is created by the exact feature set the workflow uses.

Build the go/no-go evidence packet

The architecture owner should be able to hand a reviewer one compact packet for the proposed workflow. It should identify the source record, protected representation, provider surface, return mechanism, and operational copies without including unnecessary original values.

At minimum, retain:

  1. the approved product, account, tenant or project, billing state, region, model, and endpoint;
  2. the enabled files, history, cache, search, grounding, tool, connector, and background-processing features;
  3. the provider’s current training, feedback, abuse-review, retention, deletion, and ZDR terms for that surface;
  4. a sanitized outbound request showing that original values and the protected mapping are absent;
  5. the policy and mapping scope that authorize Reconstruction;
  6. the final destination and evidence that only the permitted result was written there; and
  7. an inventory of customer-side logs, traces, queues, retries, errors, caches, and support copies.

Run failure cases as part of the same review. Disable or add one optional feature and confirm that the approval record detects the change. Send an altered or unknown protected reference and verify that Reconstruction does not guess. Force a provider error, retry, and rejected downstream write. Confirm that each path produces a bounded, sanitized artifact and a visible failure state.

The packet should also record what triggers reassessment. A new endpoint, model family, account tier, caching option, grounding service, connected tool, provider term, observability platform, or output destination can change the approved boundary even when the application code appears unchanged.

Approve one configuration for one workflow

An external-model approval should be narrow enough to test. Name the workflow, data class, caller, provider surface, model route, enabled features, Reconstruction scope, and destination. State which changes return the workflow to review.

This approach lets a team continue using an approved external model without turning the vendor’s brand into a security conclusion. LLM Capsule governs what follows the model path in its documented architecture. The provider review governs what that service may store or observe. Deployment evidence connects the two.

Approve the workflow only when all three are aligned: the original-value boundary, the exact provider configuration, and the authorized return path.

Request an LLM Capsule deployment review

FAQ

Can sensitive data be used with an external LLM?

A workflow can use an approved external LLM while keeping selected restricted original values off the model route, provided the deployed architecture sends only an evaluated protected working version and the organization separately approves the provider service, configuration, remaining context, and return path.

Does “not used for training” mean the provider retains no data?

No. Training treatment, abuse or safety review, application state, history, files, caching, tools, and deletion are separate controls. Verify each one for the exact product surface and feature set.

Does LLM Capsule provide zero data retention at the model provider?

No. LLM Capsule changes what follows the approved model path in its documented target architecture. Provider-side retention and feature state remain governed by the selected service, account, configuration, and terms.

Should ChatGPT, Claude, and Gemini be approved at the brand level?

No. Approval should name the exact consumer app, enterprise workspace, developer API, or managed service, along with the account, endpoint, enabled features, retention controls, and workflow destination.