The problem this solves

Revenue-ranked reporting tells a flattering story that profit-ranked reporting often contradicts. A store can post strong revenue and still be a net drag on the business once it carries its fair share of corporate overhead โ€” and the reverse is just as common: a modest-revenue, thinly-staffed digital channel can be quietly one of the most profitable things in the portfolio. The "whale curve" โ€” profit ranked and plotted as cumulative % โ€” is the standard way profitability analysts surface this: it shows, in one chart, how concentrated real profit actually is, and how much of it a long tail of thin or loss-making entities is eroding.

This is the payoff use case for the whole PCMCS module โ€” it takes the exact overhead allocation UC13 computes for one entity at a time and runs it for all 42 at once, then asks the question none of the other use cases ask directly: given everyone's fair share of overhead, who's actually making the company money?

Audience

Finance leadership and portfolio strategy teams deciding where to invest, close, or renegotiate. FP&A teams building the board's profitability narrative. Store and channel controllers who want to see where their entity sits in the full ranking, not just its own number in isolation.

Every entity, the same fully-loaded calculation

StepSource
NLQ query: "Show the whale curve for FY26" โ€” DeepSeek resolves year and scenario; the ranking recomputes for all 42 entities
Pre-allocation net incomegenIS(entityId, year, scenario).FY._NI โ€” the FP&A module's existing P&L generator
Allocated overheadcomputeAllocationWaterfall(entityId, year, scenario).totalAllocated โ€” the same UC13 waterfall, called once per entity
Fully-loaded profitpreAllocationNI โˆ’ allocatedOverhead

There is no separate profitability data model โ€” this use case is entirely a ranking and presentation layer over the other two modules' existing calculations.

A ranked curve and the full entity ledger behind it

  • KPI row: entities ranked, count profitable, total company fully-loaded profit, peak cumulative %
  • Whale curve chart: cumulative % of total profit against entity rank, with a 100% reference line โ€” the peak-then-decline-to-100% shape is the headline visual
  • Detail table: every entity's revenue, pre-allocation NI, allocated overhead, and fully-loaded NI, ranked and colored by sign

The last step in the PCMCS sequence

Allocation Waterfall (UC13) answers "how much overhead does one entity carry." Reciprocal Cost Allocation (UC14) handles the harder circular-cost case. Profitability Curve takes UC13's output, runs it across the entire portfolio, and turns it into a decision-support ranking โ€” which is the point the other two use cases build toward but don't answer directly.

Recommended integration points

  • Portfolio review / store rationalization: the bottom of the curve is exactly the entity list a rationalization exercise starts from โ€” but always paired with the "why" from UC13's trace, not the ranking alone
  • Investment prioritization: entities near the top of the curve are candidates for expansion capital; the curve makes "where does the next dollar of growth investment earn the most" a visual answer, not a spreadsheet sort
  • Board / investor profitability narrative: "our top 16 entities generate 116% of company profit" is a much sharper story than an average margin percentage

The numbers

42
Entities ranked
16
Profitable after full allocation
116%
Peak cumulative profit %
0
Separate data model needed
$0.0002
Cost per NLQ query
100%
Curve always ends here, by construction

The honest checklist

  • โœ“You already allocate overhead down to entities (or products, customers, channels) and want to see the profitability picture that emerges after that allocation, not before it
  • โœ“Leadership makes portfolio decisions (expand, hold, close, renegotiate) and needs a profit-concentration view, not just a revenue ranking
  • โœ“You want the allocation methodology (UC13/UC14) and its downstream profitability consequence (this) visibly connected, not computed in separate disconnected tools
  • โœ—You haven't yet built a defensible overhead allocation โ€” start with UC13; a profitability ranking is only as trustworthy as the allocation feeding it
  • โœ—You need product- or customer-level profitability rather than entity-level โ€” this demo ranks the 42 entities in the shared cube; the same technique applies to any cost object with a revenue and an allocated-cost figure

Try it now

The live demo ranks all 42 entities and redraws the curve in real time.

Launch Profitability Curve โ†’ Start a Lab engagement

Things to try

  • Type "Which stores are unprofitable after allocation?" and scroll the detail table to the bottom rows โ€” compare their revenue to their allocated overhead
  • Switch scenario to Actual and watch both the curve shape and the peak cumulative % shift as every entity's underlying P&L recomputes
  • Cross-reference a top-of-curve entity here against its Allocation Waterfall trace (UC13) โ€” the two numbers should tie out exactly
HOW WE BUILT IT

A ranking is only as honest as its inputs

It would be easy to build a "profitability curve" that ranks entities by raw, pre-allocation net income โ€” that's just a revenue-adjacent sort and tells you nothing a simple P&L report doesn't already show. The entire value of this use case is that it ranks by fully-loaded profit, meaning every entity has already absorbed its real, traceable, auditable share of corporate overhead before the ranking happens. That's why this use case has no data model of its own โ€” building one would risk the ranking disagreeing with the allocation the rest of the module already computes.

Reuse, sort, accumulate

1
computeEntityProfitability()
entityId, year, scenario
  • preAllocationNI = genIS(entityId,...).FY._NI
  • waterfall = computeAllocationWaterfall(entityId,...) — UC13, reused as-is
  • fullyLoadedNI = preAllocationNI − waterfall.totalAllocated
2
computeWhaleCurve()
year, scenario
  • data = allocationUniverse().map(computeEntityProfitability)
  • Sort descending by fullyLoadedNI
  • totalProfit = sum(fullyLoadedNI); running cumulative → cumulativePct
3
Renderer
profitability.html
  • KPI row + line chart (cumulative % vs rank, 100% reference line)
  • Detail table

App UI โ€” Component breakdown

ComponentBehaviour
Year/Scenario filtersOnly two filters โ€” like Ownership (UC9), there's no entity selector because all 42 entities always render together; that's the entire point of a ranking
KPI rowEntity count, profitable count, total fully-loaded profit (colored green/red by sign), peak cumulative %
Whale curve chartLine chart, cumulative % of total profit (Y) against entity rank (X), plus a dashed 100% reference line so the "erosion back to 100%" is visually obvious
Detail tableEvery entity, ranked, with revenue/pre-allocation NI/allocated overhead/fully-loaded NI/cumulative % โ€” green for profitable, red for loss-making rows

One subtraction, correctly sourced

fullyLoadedNI = genIS(entityId, year, scenario).FY._NI
              - computeAllocationWaterfall(entityId, year, scenario).totalAllocated

The calculation itself is a single subtraction โ€” the engineering value is entirely in calling the same waterfall function UC13 exposes to users, rather than re-deriving a parallel allocation inline. If UC13's allocation methodology ever changes (a new pool, a different driver, a resized pool), this use case's ranking updates automatically and stays consistent with it โ€” there's no second copy of the allocation logic that could silently drift out of sync.

A property of the math, not a coincidence

cumulativePct is defined as runningSum / totalProfit ร— 100. By the time every entity has been included in the running sum, runningSum === totalProfit by construction โ€” so the curve's final point is always exactly 100%, regardless of how profit is distributed across entities. What varies is everything before that final point: how high the curve peaks (concentration among top performers) and how far it dips before recovering (how much of a drag the bottom performers are). This was verified directly: a unit test asserts the last point's cumulativePct equals 100 to within 0.01, for any year/scenario combination.

Getting a real whale shape, not a cliff

The first pass at sizing UC13's allocation pools (7.5% of revenue combined) produced a company-wide fully-loaded loss โ€” only 6 of 42 entities stayed profitable, and the "curve" was really a near-monotonic climb with no meaningful peak, because almost the entire portfolio had gone negative together rather than showing genuine dispersion. A second pass at 3% combined over-corrected the other way: 40 of 42 entities profitable, peak cumulative % of only 100.1% โ€” barely a whale shape at all, since almost nothing in the tail was actually eroding the total. The pools were re-tuned to 4.7% combined (1.9% G&A + 1.3% IT + 1.5% Marketing), which produces 16 of 42 entities profitable and a peak of 116% โ€” a shape with a real long tail, verified by rerunning the same Node unit-test harness after each adjustment rather than eyeballing the chart.

The smallest schema, matching Ownership's pattern

Profitability's NLQ schema โ€” {intent, year, scenario, confidence} โ€” is deliberately the same shape as Ownership's (UC9): no entity or method to resolve, because the output is always the full ranking of all 42 entities at once. The few-shot examples cover the phrasing variety users would naturally reach for โ€” "whale curve," "profitability curve," "which stores are unprofitable," "fully loaded profit," "cumulative profit contribution" โ€” rather than any lookup logic, since year/scenario extraction is the only real parsing work this use case's NLQ layer has to do.

Cost model โ€” At-scale projections

LLM cost is flat per query at roughly $0.0002 โ€” year/scenario only, the same minimal footprint as Ownership. The ranking calculation is O(n log n) in entity count (dominated by the sort), and each entity's profitability calculation is itself O(regions) via the shared waterfall โ€” so a company with 500 entities instead of 42 would still rank the full portfolio in well under a second of client-side compute.

Key files

FileRole
epm-nlq-src/assets/epm-pcmcs-data.jscomputeEntityProfitability(), computeWhaleCurve() โ€” both call UC13's computeAllocationWaterfall() directly
epm-nlq-src/pages/profitability.htmlUI: NLQ bar, year/scenario filters, KPI row, whale-curve line chart, ranked detail table
epm-nlq-src/assets/epm-data.jsShared genIS() โ€” pre-allocation net income comes straight from the FP&A module
functions/api/nlq-query.jsShared NLQ endpoint; Profitability uses useCase: "profitability" โ€” the same minimal schema shape as Ownership

Why this use case has no data model of its own

Every number this page shows already exists somewhere else in the codebase โ€” revenue and pre-allocation NI in epm-data.js, allocated overhead in this same file's computeAllocationWaterfall(). This use case's entire contribution is the sort-and-accumulate step that turns 42 independent entity calculations into one ranked, cumulative story โ€” a deliberately thin layer, by design.

Tech stack โ€” Every tool in this build

LayerToolWhy
Data layerepm-pcmcs-data.js (vanilla JS)Thin sort/accumulate layer over UC13's allocation and the FP&A module's P&L โ€” no independent data model
LLMDeepSeek V3Shared NLQ endpoint; minimal year/scenario-only schema, same shape as Ownership UC9
ChartsChart.js 4.xLine chart with a dashed reference line โ€” shared library, same pattern as Capital UC6's depreciation curve
Edge hostingCloudflare PagesStatic file, no server compute needed for this use case
BuildEleventy v3.1.5Copies epm-nlq-src/pages/ and epm-nlq-src/assets/ to _site/ verbatim

Known attack surfaces

ThreatMitigation in this build
Ranking silently drifting from the allocation it's supposed to reflectNo separate allocation logic exists in this file โ€” computeEntityProfitability() calls UC13's computeAllocationWaterfall() directly, so the two can never disagree
Curve not actually ending at 100% (a sign of a calculation bug)Verified directly in testing on every code change: the last point's cumulativePct must equal 100 to within 0.01, or the test suite fails
A degenerate all-positive or all-negative portfolio producing a meaningless "curve"Pool sizing was explicitly tuned (see "Tuning the pool sizes") and unit-tested to guarantee a genuine mix of profitable and unprofitable entities, not a monotonic line

Guardrails โ€” What prevents a misleading curve

  • Zero-division guard on cumulativePct: if totalProfit is ever exactly zero, the calculation falls back to 0 rather than dividing by zero and propagating NaN or Infinity into the chart
  • Sort order is asserted, not assumed: unit-tested that every entity's fullyLoadedNI is greater than or equal to the next entity's โ€” a sort-comparator regression would fail the test before ever reaching the UI
  • Universe matches UC13's exactly: both use the same allocationUniverse() function (non-corporate entities only), so the entity count shown here always reconciles with the entity count UC13's dropdown offers

Who is asking, and what are they allowed to see?

The demo answers neither question — it has a cookie gate and no notion of a user. In production these are the two questions everything else rests on, and they have different answers: authentication is who you are, authorization is what you may see. Corporate SSO settles the first. Only Oracle EPM Cloud can settle the second, and the single most important rule in this section is that this application must never become the place where that decision is made.

The rule that governs every choice below: a user must see exactly what they would see by logging into Oracle EPM Cloud directly — no more, and no less. If this tool can surface a number the user could not retrieve themselves, it has become a privilege-escalation path, and it will be found in the first access review.

9.1 · The identity chain, end to end

Finance user opens the tool in a browser — no local account, no password held here
1
Corporate identity provider
Entra ID · Okta · OCI IAM
OIDC Authorization Code + PKCE  (or SAML 2.0)
  • MFA and Conditional Access are enforced here — device compliance, location, risk signals
  • Returns an ID token (who the user is) and an access token (what they may call)
  • Group membership arrives as a claim; the application never handles a password
2
Application session
validate, never trust
  • Verify signature, issuer, audience and expiry against the IdP’s published keys
  • Read the group claims — there is no local user table and no local role table
  • Short-lived access token with refresh-token rotation; session timeout set to the data classification
Pattern A — identity propagation
OAuth 2.0 token exchange (on-behalf-of)
  • The API is called as the user
  • EPM enforces its own security natively
  • Audit trail names the real user
  • Preferred where the API supports it
Pattern B — service account + filtering
one read-only integration account
  • The application becomes the enforcement point
  • Entitlements fetched separately, applied in one audited place
  • Simpler and cacheable — and a filtering bug is a data breach
3
EPM identity domain
roles + dimension security
  • Roles: Service Administrator, Power User, User, Viewer — assigned to groups, never to individuals
  • Data level: PCMCS POV access; the ranking is portfolio-wide and should not be assembled from a partial entitlement
  • Group → role mapping lives in the platform, not in this application
Result, filtered to this user
profitability.html
  • The user sees exactly what they would see logging into the source system directly — no more
  • Every query logged against the real end user, never a shared account

9.2 · Federating the corporate identity provider

Oracle EPM Cloud does not replace your directory — it trusts it. The EPM Cloud identity domain is federated with the corporate IdP so authentication happens where it already happens, under policies security has already written.

Identity providerProtocolNotes
Microsoft Entra ID (formerly Azure AD)SAML 2.0 or OIDCThe common case. Conditional Access, MFA and device compliance are enforced at Entra and inherited automatically. On-premises Active Directory federates through Entra Connect rather than being integrated directly.
OktaSAML 2.0 or OIDCSame pattern; Okta groups drive EPM roles through SCIM provisioning.
OCI IAM (identity domains)NativeAlready present with Oracle EPM Cloud. Can be the primary IdP for a smaller estate, or a federated spoke of Entra/Okta for a larger one.
AD FSSAML 2.0Still seen where the estate is not yet cloud-first. Works, but you inherit the on-premises availability of the token service — if AD FS is down, nobody logs in.

For the browser application itself, use OIDC Authorization Code flow with PKCE. Not the implicit flow, which is deprecated and leaks tokens through the URL, and never a resource-owner password grant — a finance tool should not be capable of handling a password at all.

9.3 · From group membership to EPM roles

Roles are granted to groups, never to individuals, and the groups come from the directory. That one discipline is what makes joiner/mover/leaver work without anyone having to remember this application exists.

Entra ID group                    โ†’  EPM role / entitlement
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
FIN-EPM-Analysts                  โ†’  Planning User
FIN-EPM-Controllers-EMEA          โ†’  Power User + EMEA data scope
FIN-EPM-Admins                    โ†’  Service Administrator
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
Provisioned by SCIM. Remove the user from the group and the
entitlement disappears on the next sync โ€” including here.

For this use case the relevant native entitlement is: PCMCS User, with the ranked portfolio view typically limited to leadership and FP&A.

9.4 · The architectural decision: who enforces?

This is the choice that determines whether the deployment is defensible. Both patterns appear in the diagram above; the difference is where the security boundary actually sits.

Pattern A — identity propagationPattern B — service account + filtering
HowThe user’s token is exchanged (OAuth 2.0 on-behalf-of) for one scoped to the EPM API; calls are made as the userA single read-only integration account calls the API; the application filters the results
Enforcement pointOracle EPM CloudThis application
Audit trail showsThe real end userThe service account — you must log the real user separately
Failure modeToken plumbing is more complex; per-user rate limits applyA filtering bug is a data breach, and the entitlement copy drifts from reality
VerdictPrefer this wherever the API supports user-token authenticationAcceptable with discipline: narrowest possible service account, filtering centralised in one tested place, real user in every log line

The shortcut to refuse. Pattern B built with a Service Administrator account and no filtering at all is the most common way this gets delivered, because it works perfectly in UAT — testers are usually over-entitled, so nobody notices that everyone can see everything. It fails at the first access review, and by then it is in production with real users depending on it.

9.5 · Data-level security is the part that matters

Role membership decides whether a user can open the application. It does not decide which rows they get back, and confusing the two is the most expensive mistake available here.

  • For this use case: PCMCS POV access; the ranking is portfolio-wide and should not be assembled from a partial entitlement.
  • Apply it before aggregation, not after. Filtering a total that has already been computed across entities the user cannot see still leaks the total.
  • The NLQ layer needs its own check. Layer 4 already validates that the resolved point of view uses approved members; production adds a second test — that the resolved POV sits inside this user’s scope — and it runs before the data call, not after. A natural-language interface is very good at asking for things politely; the authorization check must not care how the question was phrased.
  • Fail closed. If entitlements cannot be resolved, return nothing and say so. An empty result is a support ticket; a permissive default is an incident.

9.6 · Provisioning, sessions and the leaver problem

  • SCIM provisioning from Entra or Okta into the EPM Cloud identity domain (OCI IAM, formerly IDCS), covering joiner, mover and leaver. The mover is the case people forget — somebody changing region should lose the old scope, not accumulate both.
  • No local user store. If this application keeps its own copy of who may do what, a leaver keeps access until somebody remembers to update it. Nobody ever does.
  • Short-lived access tokens with refresh-token rotation; align session timeout with the data classification rather than with convenience.
  • MFA and Conditional Access at the IdP — not reimplemented here. Device compliance and location policy come free with federation.
  • Quarterly recertification of both the groups that grant access and the service account’s own entitlements, evidenced and signed.
  • Break-glass access is a named, monitored, time-boxed account — never a shared credential in a password manager.

9.7 · What this means for Profitability Curve

ConcernAnswer for this use case
Native entitlement requiredPCMCS User, with the ranked portfolio view typically limited to leadership and FP&A
Data-level controlPCMCS POV access; the ranking is portfolio-wide and should not be assembled from a partial entitlement
Use-case-specific sensitivityA ranked list of loss-making stores is commercially sensitive and, in some jurisdictions, employment-sensitive — it reads as a closure shortlist whether or not it is one. Limit the audience to leadership and FP&A, log every view, and treat an export as a restricted document.

9.8 · Security configuration checklist

  • ✓Oracle EPM Cloud federated with the corporate IdP over SAML 2.0 or OIDC; the cookie gate removed entirely
  • ✓Browser app uses OIDC Authorization Code + PKCE — no implicit flow, no password grant
  • ✓MFA and Conditional Access enforced at the IdP, not reimplemented in the application
  • ✓Roles granted to directory groups, never to individuals; SCIM covers joiner, mover and leaver
  • ✓Enforcement pattern chosen deliberately — Pattern A where the API supports it, or Pattern B with filtering centralised and tested
  • ✓Data-level security applied before aggregation, and the resolved POV checked against the user’s scope before the data call
  • ✓No local user table and no local role table anywhere in the application
  • ✓Every query logged against the real end user, even when a service account makes the call
  • ✓Authorization failures fail closed and are logged as security events rather than swallowed
  • ✓Quarterly recertification of access groups and of the service account’s own entitlements
Talk through your identity model → Back to the demo

From demo to a governed enterprise deployment

Everything above runs on synthetic data, a public LLM API key, a cookie gate, and no audit trail — deliberately, so the mechanics are inspectable. Taking Profitability Curve to production is not a rewrite; the 4-layer pipeline and the data-layer contract survive intact. It is a controlled-change program across six workstreams: architecture, LLM platform, security, SOX/audit, environment promotion, and operations. This section is the checklist we run with clients.

The one rule that matters most for this use case: in the demo the browser computes the result; in production Oracle computes and this layer retrieves and explains. Never ship a second calculation engine that can disagree with the system of record — the moment two numbers exist, the audit question becomes “which one is right,” and the answer must always be the EPM module.

10.1 · Production reference architecture

Finance user · corporate SSO (OIDC/SAML + MFA) · EPM role claims
1
Edge / API Gateway
WAF · rate limit · identity
  • Terminates SSO, validates the session, attaches the user’s EPM groups to the request
  • Rate limits per user, blocks anonymous access, scrubs PII patterns before anything reaches the orchestrator
2
NLQ Orchestrator
the 4-layer pipeline, hardened
L1 guardrails → L2 grounding → L3 LLM adapter → L4 eval + fallback
  • L2 grounding reads dimension metadata from EPM on a schedule — not a hardcoded schema
  • Only the schema + user query go to the model; financial values never leave the data layer
  • L4 rejects anything outside the approved member lists and falls back to the deterministic parser
Oracle OCI Generative AI
same tenancy as EPM Cloud
  • Data stays inside the OCI boundary
  • Natural fit when EPM is already in OCI
Azure OpenAI / AWS Bedrock / Vertex AI
private endpoint, zero retention
  • Use the hyperscaler the org already governs
  • Enterprise DPA, no training on prompts
Self-hosted open weights
VPC / air-gapped
  • For regulated or sovereign data
  • Highest control, highest run cost
3
EPM Data Layer
Oracle EPM REST API
  • Oracle PCMCS analysis views / profit curves — fully-loaded profitability from the PCMCS model via REST, using the same allocation results UC13 surfaces
  • Least-privilege service account (read-only role, one app, one pod) with the token in a vault and rotated
  • Results filtered to the requesting user’s EPM security before rendering
4
Audit & Observability
append-only
  • Every query logged: user, timestamp, raw query, parsed intent JSON, model + prompt version, POV returned, latency, cost
  • Exported to the SIEM; retained per the SOX evidence schedule
  • Dashboards for fallback rate, eval pass rate, guardrail hits, p95 latency, spend
Rendered result + evidence trail
profitability.html
  • The parsed JSON is shown to the user as the explanation (“AI: entity · year · scenario — 93% confident”) — the same line the demo prints today
  • Every number on screen traces to an EPM cell intersection an auditor can reproduce

10.2 · Choosing the LLM platform

The demo’s DeepSeek call is a placeholder for a single adapter, callLLM(system, user), behind Layer 3. Swapping the provider changes one function and zero business logic. Pick the platform the organisation already governs — the security and procurement review is the long pole, not the integration.

OptionChoose whenData posture
Oracle OCI Generative AI (Cohere Command, Llama)EPM Cloud already lives in OCI; you want one cloud boundary and one contractPrompts stay in the OCI tenancy; no training on customer data; dedicated AI clusters available for isolation
Azure OpenAI ServiceMicrosoft-first finance estate (Entra ID, Purview, Sentinel already in place)Private endpoint, regional deployment, zero-retention by default under the enterprise agreement
AWS Bedrock (Claude, Titan) / Google Vertex AI (Gemini)The org’s landing zone is AWS or GCP; VPC endpoints and IAM already auditedVPC/PSC private access, no data used for training, CloudTrail/Cloud Audit Logs integration
Direct enterprise API (Anthropic, OpenAI)Fastest model access; acceptable when a zero-data-retention agreement and DPA are signedZDR endpoint, SSO-managed keys, SOC 2 report on file
Self-hosted open weights (Llama, Mistral, Qwen via vLLM)Sovereign or air-gapped requirements; regulated data classification forbids any external inferenceFull control; you own patching, eval, and capacity — budget for an MLOps owner

Put a model gateway in front of whichever you choose (Azure API Management, OCI API Gateway, Kong AI Gateway, LiteLLM, or Portkey): it owns key custody, per-team spend caps, routing and fallback between models, prompt/response logging, and lets you retire a deprecated model without touching the application.

10.3 · Security controls

ControlImplementation
Identity & accessCovered in full in section 09 — corporate SSO, group-to-role mapping, and the decision about who enforces data-level security. Listed here because it is a production gate, not because it is optional.
Service accountOne read-only EPM service account per application per pod, least-privilege role, no interactive login, credential in a vault (OCI Vault, Azure Key Vault, HashiCorp Vault), rotated on a schedule and on staff change.
Secrets & configNo secrets in code or build artifacts; environment-specific config injected at deploy; .dev.vars-style files never leave a developer machine.
NetworkPrivate endpoints to the LLM provider and to EPM where the platform supports them; egress allow-list so the orchestrator can reach exactly two hosts; TLS 1.2+ everywhere.
Prompt-injection & input guardrailsLayer 1 (already in the demo) blocks instruction-override patterns, enforces length and scope; extend with a classifier on the gateway and log every rejection.
Output guardrailsLayer 4 (already in the demo) validates every returned member against the approved lists and strips unexpected keys; production adds a policy check that the resolved POV is inside the user’s security scope before the data call.
Data minimisationPrompts contain metadata and the user’s query only. No cell values, no employee names, no free-text comments from EPM. Logged prompts are classified and retained accordingly.
EncryptionIn transit (TLS) and at rest (provider-managed KMS); audit logs on immutable storage with customer-managed keys where policy requires.

10.4 · SOX, audit, and model-risk controls

A read-only NLQ layer does not change a financial-reporting control, but it is an interface to a SOX-relevant system and lands squarely in ITGC scope. Treat prompts, schemas, and eval sets as code — that single decision satisfies most of what an auditor will ask for.

RequirementHow it is satisfied
Complete, immutable audit trailAppend-only log of user, timestamp, raw query, parsed JSON, model and prompt version hash, POV returned, and row count — WORM storage, retained for the evidence period (typically 7 years), exported to the SIEM.
Change managementPrompt templates, few-shot examples, approved-member schema, and code are version-controlled; every change follows ticket → peer review → test evidence → CAB approval → deploy. A prompt edit is a code change.
Segregation of dutiesDevelopers cannot deploy to production; the service-account owner is not a developer; production secrets are held by platform operations.
Access recertificationQuarterly review of who can use the tool and of the service account’s EPM roles, evidenced and signed.
Testing evidenceA golden-query regression suite (the few-shot examples plus a larger labelled set) runs in CI before every release; pass rate and diffs are archived as release evidence.
Model risk managementAn inventory entry (intended use, limitations, owner, validation date) in the model-risk register — the SR 11-7 pattern for financial services; periodic re-validation when the model or prompt changes.
Reproducibility & lineageEvery displayed number traces to an EPM POV and a consolidation/calculation timestamp; an auditor can re-query the same intersection in EPM and match it.
ExplainabilityThe parsed intent JSON is the explanation and is shown to the user on every response — no hidden reasoning between the query and the data call.

10.5 · Dev → Test → Prod promotion

EnvironmentEPM targetDataGate to leave
DevEPM Test pod (developer slice)Synthetic or maskedUnit tests on the data layer; lint; eval suite ≥ threshold against the Test LLM deployment
Test / UATEPM Test pod (full refresh)Masked copy of productionBusiness UAT sign-off on the golden queries; security scan; performance run (p95 latency, fallback rate)
ProdEPM Production podLiveChange ticket approved; deploy in window; smoke test; hypercare with rollback ready
  • Promoted artifacts: application build, prompt templates (versioned), approved-member schema snapshot, eval set, infrastructure config (IaC) — all from the same Git tag.
  • Pipeline: branch → PR review → CI (tests + evals) → deploy to Test → UAT sign-off → CAB → deploy to Prod → smoke test. Hosting can stay on Cloudflare Pages/Workers or move to OCI Functions + API Gateway or the org’s standard platform — the code does not care.
  • Configuration: per-environment secrets and endpoints injected at deploy; the same build runs in every environment.
  • Metadata sync: a scheduled job refreshes dimension metadata into Layer 2 grounding with change detection, so a new entity or account appears in the approved lists without a code release.
  • Rollback: previous build and previous prompt version retained; rollback is a redeploy, and because prompts are versioned it also reverts a prompt regression.

10.6 · Operating it

  • SLOs: p95 latency, availability of the read path (the deterministic fallback keeps it alive when the LLM is down — already built), fallback rate as a quality signal, eval pass rate per release.
  • Cost governance: per-user and per-team token budgets at the gateway; alert on anomalies; the unit-cost model earlier in this kit is the baseline.
  • Model lifecycle: providers retire models on a schedule — re-run the eval suite on the successor before switching, and record the switch as a change.
  • Incident runbook: LLM outage → fallback parser; EPM API outage → cached metadata with a stale banner; guardrail spike → review logs for injection attempts.

10.7 · What changes for Profitability Curve

ConcernProduction answer
System of recordOracle PCMCS analysis views / profit curves — fully-loaded profitability from the PCMCS model via REST, using the same allocation results UC13 surfaces
Read/write postureRead-only. Ranking is decision support; closure or investment decisions are recorded elsewhere.
Use-case-specific controlA ranked list of loss-making stores is commercially sensitive — limit the audience to leadership/FP&A roles and log every view.

10.8 · Production readiness checklist

  • ✓LLM platform selected from the governed list, DPA / zero-retention terms on file, gateway in front of it
  • ✓SSO integrated; authorisation derived from EPM security groups; cookie gate removed
  • ✓Read-only EPM service account per pod, credential in a vault, rotation scheduled
  • ✓Prompts, schema, and eval set version-controlled and under change management
  • ✓Append-only audit log wired to the SIEM with the agreed retention
  • ✓Golden-query eval suite passing in CI; results archived as release evidence
  • ✓Dev / Test / Prod pipeline with gates, IaC, and a rehearsed rollback
  • ✓Model-risk register entry and owner named; first re-validation date set
  • ✓Metadata refresh job scheduled with change detection
  • ✓SLOs, cost caps, and the incident runbook agreed with platform operations
Plan a production rollout with us → Back to the demo