The problem this solves
Corporate overhead doesn't sell anything itself, but every store still has to absorb a fair share of it before anyone can say whether that store is actually profitable. The question "how much of Corporate G&A does the NYC Flagship really carry?" has no single obvious answer โ it depends on which driver you pick (revenue? headcount? square footage?) and how the allocation cascades through the org (does cost hit stores directly, or pass through a regional layer first?).
Get this wrong and every downstream profitability number is wrong with it: a small, high-headcount store can look artificially unprofitable if IT cost is allocated purely by revenue, while a huge, thin-staffed digital channel can look artificially cheap to run. PCMCS exists specifically to make this allocation auditable โ not just "here's the final number" but "here's exactly which pool, at which stage, produced this number."
Audience
FP&A and cost accounting teams who own the corporate allocation methodology. Store and regional controllers who need to understand โ and be able to challenge โ the overhead they're being charged. Finance leadership deciding whether a driver (revenue vs headcount vs another basis) is still the right allocation basis as the business mix shifts.
Three cost pools, two different drivers
| Pool | Size | Driver | Why this driver |
|---|---|---|---|
| NLQ query: "Trace NYC Flagship's allocated overhead" โ DeepSeek resolves the entity, year, and scenario; the waterfall recomputes for that entity in real time | |||
| Corporate G&A | 1.9% of total revenue | Revenue | Executive, legal, and finance overhead scales roughly with the business each entity generates |
| IT & Digital Infrastructure | 1.3% of total revenue | Headcount | Seats, devices, and support tickets track people, not revenue โ a lean digital channel needs far less IT support per dollar of sales than a large staffed store |
| Marketing & Brand | 1.5% of total revenue | Revenue | Brand spend is sized to protect and grow the revenue base it supports |
Pool sizes are computed as a percentage of total company revenue at runtime, so they stay proportional no matter which year or scenario is selected โ there's no hardcoded dollar figure to go stale.
A fully traced allocation, pool by pool
- KPI row: total allocated overhead for the selected entity, as a % of that entity's revenue, its region, and how many pools were traced
- Bar chart: the three pools' allocated amounts side by side for the selected entity
- Plain-language trace: a sentence per pool explaining exactly why the entity was charged that amount โ its share of the driver within its region, and that region's share of the driver company-wide
- Detail table: every number in the two-stage cascade โ pool total, region share %, stage-2 amount, entity share within region %, stage-3 amount โ nothing is a black box
The first step of the PCMCS module
Allocation Waterfall answers "how much overhead does this entity carry, and why." Reciprocal Cost Allocation (UC14) is the same idea applied to a harder case โ cost centers that serve each other, not just downstream entities. Profitability Curve (UC15) is the payoff: it takes this exact waterfall, runs it for every entity at once, and ranks the results.
Recommended integration points
- Store-level P&L review: when a store controller disputes their overhead charge, this trace is the answer โ every dollar has an explicit path back to a pool and a driver
- Driver methodology review: before switching a pool's driver from revenue to headcount (or vice versa), model the impact on every entity first
- New entity onboarding: a newly opened store immediately needs an overhead allocation โ this shows exactly what it will be charged and why, driver by driver
The numbers
The honest checklist
- โYou allocate shared/corporate cost to stores, products, or business units and need the allocation to be auditable, not just a spreadsheet formula nobody remembers
- โDifferent cost pools genuinely should use different drivers โ you're not comfortable allocating everything by revenue just because it's simplest
- โEntities push back on their overhead charge and you need a trace, not just a total, to answer them
- โYour organization has a single flat overhead pool with no meaningful driver choice to make
- โYou need cost centers that allocate to each other, not just downstream to final entities โ that's UC14, not this one
Try it now
The live demo recomputes the full two-stage waterfall for any of 42 entities in real time.
Things to try
- Type "Trace NYC Flagship's allocated overhead" into the NLQ bar, then switch to a small store like Santiago โ notice the IT pool's headcount-driven share behaves very differently from the revenue-driven G&A and Marketing pools
- Compare a Digital entity (very high revenue, very low headcount) against a similarly-sized physical store โ the driver choice visibly changes which pool dominates its total charge
Traceability is the actual product
The allocated dollar amount itself is the easy part โ a spreadsheet SUMPRODUCT can do that. What's hard, and what production PCMCS earns its keep on, is being able to answer "why" for any single dollar without re-deriving the calculation by hand. This demo is built around that constraint: every number the UI shows is stored and rendered at every stage, not just the final total, so the trace is a first-class output rather than something reconstructed after the fact.
Pool total โ region share โ entity share
- Per-entity revenue (genIS) and headcount (genWorkforce)
- Summed to company totals and region totals
- For each of 3 ALLOCATION_POOLS:
- poolTotal = pctOfRevenue × totalRevenue
- regionShare% = region driver ÷ company driver → stage2 = poolTotal × regionShare%
- entityShareInRegion% = entity driver ÷ region driver → stage3 = stage2 × entityShareInRegion%
- KPI row + bar chart
- Plain-language trace
- Full detail table
App UI โ Component breakdown
| Component | Behaviour |
|---|---|
| Entity selector | All 42 non-corporate entities โ corporate entities are cost sources, never allocation destinations |
| KPI row | Total allocated overhead, % of that entity's revenue, its region, pool count |
| Bar chart | The three pools' Stage-3 (entity-level) amounts, side by side |
| Trace panel | One auto-generated sentence per pool stating the entity's share of the driver within its region, and the region's share company-wide |
| Detail table | Every intermediate number in the cascade โ nothing is computed and then discarded before rendering |
Why two different drivers, not one
A single driver for every pool is the most common allocation shortcut in practice, and it's also the most common source of allocation disputes. Revenue-based allocation systematically under-charges high-revenue/low-headcount entities (a busy digital channel) and over-charges low-revenue/high-headcount ones (a large-format store with heavy staffing but modest per-store sales) for anything that actually tracks people rather than sales โ IT seats, help-desk tickets, device refreshes. This demo deliberately splits the three pools across two drivers (revenue for G&A and Marketing, headcount for IT) specifically so that contrast is visible when comparing entities with different revenue-to-headcount ratios.
The driver assignment is a property of the pool definition (ALLOCATION_POOLS[i].driver), not a per-query choice โ in a real deployment, changing a pool's driver is a methodology decision that should be deliberate and reviewed, not something an end user toggles per query.
Why region-then-entity, not straight to entity
// Stage 1 โ 2: pool total splits across 4 regions by driver share
regionSharePct = regionDriverTotal / companyDriverTotal
stage2Amount = poolTotal ร regionSharePct
// Stage 2 โ 3: that region's allocation splits across its entities
entityShareInRegionPct = entityDriver / regionDriverTotal
stage3Amount = stage2Amount ร entityShareInRegionPct
Routing through a regional layer instead of going straight from pool to entity mirrors how real allocation methodologies are usually built โ regional management overhead genuinely is organized regionally, and even for the two pools that aren't regional in nature, cascading through region first makes the calculation auditable one layer at a time: "is this region's total allocation right?" is a much smaller question than "are all 42 entities' allocations right?", and the region check has to pass before the entity-level checks are even meaningful.
Every pool sums back to exactly its total
Because regionSharePct and entityShareInRegionPct are both computed as a share of driver, not as independently-set percentages, summing stage3Amount across every one of the 42 entities for a given pool reproduces that pool's poolTotal exactly (within floating-point rounding) โ there is no leakage and no double-counting built into the math. This was verified directly: a unit test sums every entity's Stage-3 allocation for each of the three pools and asserts the sum equals poolTotal to within half a dollar.
A four-field schema built for one job: pick the entity
Allocation's NLQ schema is {intent, entity, year, scenario, confidence} โ deliberately narrow. There's no pool selector, no stage selector: the demo always shows the full three-pool waterfall for whichever entity the query resolves to, because a partial trace (one pool, one stage) isn't actually useful on its own โ the value is in seeing all three pools' contributions side by side. The few-shot examples cover the entity-name variety already established across the site's other entity-bearing use cases (bare names, "Trace X's...", possessive phrasing, region-adjacent phrasing) plus year/scenario extraction, reusing the same ENTITY_ALIASES table the other eight entity-driven use cases already share.
Cost model โ At-scale projections
LLM cost is flat per query at roughly $0.0002 regardless of company size โ the query only resolves an entity name plus year/scenario. The allocation calculation itself is O(entities) per pool for the driver totals (computed once per query, not once per entity), so it scales linearly and stays trivial client-side compute even at real-world scale โ a group with 500 entities instead of 42 would still resolve a single entity's trace in well under a millisecond.
Key files
| File | Role |
|---|---|
epm-nlq-src/assets/epm-pcmcs-data.js | ALLOCATION_POOLS, computeDriverBases(), computeAllocationWaterfall(), allocationUniverse() |
epm-nlq-src/pages/allocation.html | UI: NLQ bar, entity/year/scenario filters, KPI row, bar chart, trace panel, detail table |
epm-nlq-src/assets/epm-data.js | Shared ENTITIES, REGIONS, genIS() โ revenue driver comes straight from the FP&A module |
epm-nlq-src/assets/epm-workforce-capex-data.js | genWorkforce() โ headcount driver comes straight from the Workforce module |
functions/api/nlq-query.js | Shared NLQ endpoint; Allocation uses useCase: "allocation" |
Why this module reuses two other modules' generators instead of its own
PCMCS doesn't own revenue or headcount data โ it allocates cost using whatever the entity's actual revenue and headcount already are. Calling genIS() and genWorkforce() directly (rather than duplicating a parallel set of numbers inside the PCMCS data file) keeps the allocation grounded in the same P&L and staffing the FP&A and Workforce modules already show, so a query in one module and a query in another about the same entity never disagree.
Tech stack โ Every tool in this build
| Layer | Tool | Why |
|---|---|---|
| Data layer | epm-pcmcs-data.js (vanilla JS) | Deterministic allocation math โ the LLM only resolves entity/year/scenario |
| LLM | DeepSeek V3 | Shared NLQ endpoint; four-field schema, same pattern as Workforce and Translation |
| Charts | Chart.js 4.x | Shared with Analytics UC4, Capital UC6, Translation UC7; simple bar chart for the three pools |
| Edge hosting | Cloudflare Pages | Static file, no server compute needed for this use case |
| Build | Eleventy v3.1.5 | Copies epm-nlq-src/pages/ and epm-nlq-src/assets/ to _site/ verbatim |
Known attack surfaces
| Threat | Mitigation in this build |
|---|---|
| Corporate entities receiving their own allocated overhead back (circular allocation) | allocationUniverse() explicitly excludes type === 'corporate' entities from being allocation destinations โ corporate is always a source, never a recipient, by construction |
| Pool totals silently drifting from their intended % of revenue | Pool size is computed at runtime as pctOfRevenue ร totalRevenue, never hardcoded โ it can't go stale independent of the underlying P&L |
| Allocation percentages not summing to 100% of a pool (leakage) | Verified directly in testing: summing every entity's Stage-3 allocation for a pool reproduces that pool's total exactly |
Guardrails โ What prevents bad allocation math
- Driver is a property of the pool, not the query: a user's NLQ phrasing can never change which driver a pool uses โ that's a methodology decision fixed in
ALLOCATION_POOLS - Share percentages are always driver-ratios, never independently set:
regionSharePctandentityShareInRegionPctare computed from the same driver totals used to size the pool, so the cascade can't produce a total that diverges frompoolTotal - Missing driver data fails to zero, not to a crash:
regionRev[region] ? ... : 0guards every division, so an entity or region with zero of a driver gets zero allocation rather than aNaNpropagating through the UI
Who is asking, and what are they allowed to see?
The demo answers neither question — it has a cookie gate and no notion of a user. In production these are the two questions everything else rests on, and they have different answers: authentication is who you are, authorization is what you may see. Corporate SSO settles the first. Only Oracle EPM Cloud can settle the second, and the single most important rule in this section is that this application must never become the place where that decision is made.
9.1 · The identity chain, end to end
- MFA and Conditional Access are enforced here — device compliance, location, risk signals
- Returns an ID token (who the user is) and an access token (what they may call)
- Group membership arrives as a claim; the application never handles a password
- Verify signature, issuer, audience and expiry against the IdP’s published keys
- Read the group claims — there is no local user table and no local role table
- Short-lived access token with refresh-token rotation; session timeout set to the data classification
- The API is called as the user
- EPM enforces its own security natively
- Audit trail names the real user
- Preferred where the API supports it
- The application becomes the enforcement point
- Entitlements fetched separately, applied in one audited place
- Simpler and cacheable — and a filtering bug is a data breach
- Roles: Service Administrator, Power User, User, Viewer — assigned to groups, never to individuals
- Data level: PCMCS application and POV access; model and rule-set editing is a separate, much narrower entitlement
- Group → role mapping lives in the platform, not in this application
- The user sees exactly what they would see logging into the source system directly — no more
- Every query logged against the real end user, never a shared account
9.2 · Federating the corporate identity provider
Oracle EPM Cloud does not replace your directory — it trusts it. The EPM Cloud identity domain is federated with the corporate IdP so authentication happens where it already happens, under policies security has already written.
| Identity provider | Protocol | Notes |
|---|---|---|
| Microsoft Entra ID (formerly Azure AD) | SAML 2.0 or OIDC | The common case. Conditional Access, MFA and device compliance are enforced at Entra and inherited automatically. On-premises Active Directory federates through Entra Connect rather than being integrated directly. |
| Okta | SAML 2.0 or OIDC | Same pattern; Okta groups drive EPM roles through SCIM provisioning. |
| OCI IAM (identity domains) | Native | Already present with Oracle EPM Cloud. Can be the primary IdP for a smaller estate, or a federated spoke of Entra/Okta for a larger one. |
| AD FS | SAML 2.0 | Still seen where the estate is not yet cloud-first. Works, but you inherit the on-premises availability of the token service — if AD FS is down, nobody logs in. |
For the browser application itself, use OIDC Authorization Code flow with PKCE. Not the implicit flow, which is deprecated and leaks tokens through the URL, and never a resource-owner password grant — a finance tool should not be capable of handling a password at all.
9.3 · From group membership to EPM roles
Roles are granted to groups, never to individuals, and the groups come from the directory. That one discipline is what makes joiner/mover/leaver work without anyone having to remember this application exists.
Entra ID group โ EPM role / entitlement
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
FIN-EPM-Analysts โ Planning User
FIN-EPM-Controllers-EMEA โ Power User + EMEA data scope
FIN-EPM-Admins โ Service Administrator
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Provisioned by SCIM. Remove the user from the group and the
entitlement disappears on the next sync โ including here.
For this use case the relevant native entitlement is: PCMCS User with access to the profitability application and its POV.
9.4 · The architectural decision: who enforces?
This is the choice that determines whether the deployment is defensible. Both patterns appear in the diagram above; the difference is where the security boundary actually sits.
| Pattern A — identity propagation | Pattern B — service account + filtering | |
|---|---|---|
| How | The user’s token is exchanged (OAuth 2.0 on-behalf-of) for one scoped to the EPM API; calls are made as the user | A single read-only integration account calls the API; the application filters the results |
| Enforcement point | Oracle EPM Cloud | This application |
| Audit trail shows | The real end user | The service account — you must log the real user separately |
| Failure mode | Token plumbing is more complex; per-user rate limits apply | A filtering bug is a data breach, and the entitlement copy drifts from reality |
| Verdict | Prefer this wherever the API supports user-token authentication | Acceptable with discipline: narrowest possible service account, filtering centralised in one tested place, real user in every log line |
The shortcut to refuse. Pattern B built with a Service Administrator account and no filtering at all is the most common way this gets delivered, because it works perfectly in UAT — testers are usually over-entitled, so nobody notices that everyone can see everything. It fails at the first access review, and by then it is in production with real users depending on it.
9.5 · Data-level security is the part that matters
Role membership decides whether a user can open the application. It does not decide which rows they get back, and confusing the two is the most expensive mistake available here.
- For this use case: PCMCS application and POV access; model and rule-set editing is a separate, much narrower entitlement.
- Apply it before aggregation, not after. Filtering a total that has already been computed across entities the user cannot see still leaks the total.
- The NLQ layer needs its own check. Layer 4 already validates that the resolved point of view uses approved members; production adds a second test — that the resolved POV sits inside this user’s scope — and it runs before the data call, not after. A natural-language interface is very good at asking for things politely; the authorization check must not care how the question was phrased.
- Fail closed. If entitlements cannot be resolved, return nothing and say so. An empty result is a support ticket; a permissive default is an incident.
9.6 · Provisioning, sessions and the leaver problem
- SCIM provisioning from Entra or Okta into the EPM Cloud identity domain (OCI IAM, formerly IDCS), covering joiner, mover and leaver. The mover is the case people forget — somebody changing region should lose the old scope, not accumulate both.
- No local user store. If this application keeps its own copy of who may do what, a leaver keeps access until somebody remembers to update it. Nobody ever does.
- Short-lived access tokens with refresh-token rotation; align session timeout with the data classification rather than with convenience.
- MFA and Conditional Access at the IdP — not reimplemented here. Device compliance and location policy come free with federation.
- Quarterly recertification of both the groups that grant access and the service account’s own entitlements, evidenced and signed.
- Break-glass access is a named, monitored, time-boxed account — never a shared credential in a password manager.
9.7 · What this means for Allocation Waterfall
| Concern | Answer for this use case |
|---|---|
| Native entitlement required | PCMCS User with access to the profitability application and its POV |
| Data-level control | PCMCS application and POV access; model and rule-set editing is a separate, much narrower entitlement |
| Use-case-specific sensitivity | Allocation drivers cross module boundaries — headcount comes from Workforce, revenue from Planning. A user entitled to see allocated cost is not automatically entitled to see the headcount behind it; show the driver share without exposing underlying employee counts where that separation matters. |
9.8 · Security configuration checklist
- ✓Oracle EPM Cloud federated with the corporate IdP over SAML 2.0 or OIDC; the cookie gate removed entirely
- ✓Browser app uses OIDC Authorization Code + PKCE — no implicit flow, no password grant
- ✓MFA and Conditional Access enforced at the IdP, not reimplemented in the application
- ✓Roles granted to directory groups, never to individuals; SCIM covers joiner, mover and leaver
- ✓Enforcement pattern chosen deliberately — Pattern A where the API supports it, or Pattern B with filtering centralised and tested
- ✓Data-level security applied before aggregation, and the resolved POV checked against the user’s scope before the data call
- ✓No local user table and no local role table anywhere in the application
- ✓Every query logged against the real end user, even when a service account makes the call
- ✓Authorization failures fail closed and are logged as security events rather than swallowed
- ✓Quarterly recertification of access groups and of the service account’s own entitlements
From demo to a governed enterprise deployment
Everything above runs on synthetic data, a public LLM API key, a cookie gate, and no audit trail — deliberately, so the mechanics are inspectable. Taking Allocation Waterfall to production is not a rewrite; the 4-layer pipeline and the data-layer contract survive intact. It is a controlled-change program across six workstreams: architecture, LLM platform, security, SOX/audit, environment promotion, and operations. This section is the checklist we run with clients.
10.1 · Production reference architecture
- Terminates SSO, validates the session, attaches the user’s EPM groups to the request
- Rate limits per user, blocks anonymous access, scrubs PII patterns before anything reaches the orchestrator
- L2 grounding reads dimension metadata from EPM on a schedule — not a hardcoded schema
- Only the schema + user query go to the model; financial values never leave the data layer
- L4 rejects anything outside the approved member lists and falls back to the deterministic parser
- Data stays inside the OCI boundary
- Natural fit when EPM is already in OCI
- Use the hyperscaler the org already governs
- Enterprise DPA, no training on prompts
- For regulated or sovereign data
- Highest control, highest run cost
- Oracle PCMCS — allocation results and the trace from the PCMCS calculation engine (Trace Allocations REST/UI); pools, drivers, and rule sets are PCMCS model artifacts
- Least-privilege service account (read-only role, one app, one pod) with the token in a vault and rotated
- Results filtered to the requesting user’s EPM security before rendering
- Every query logged: user, timestamp, raw query, parsed intent JSON, model + prompt version, POV returned, latency, cost
- Exported to the SIEM; retained per the SOX evidence schedule
- Dashboards for fallback rate, eval pass rate, guardrail hits, p95 latency, spend
- The parsed JSON is shown to the user as the explanation (“AI: entity · year · scenario — 93% confident”) — the same line the demo prints today
- Every number on screen traces to an EPM cell intersection an auditor can reproduce
10.2 · Choosing the LLM platform
The demo’s DeepSeek call is a placeholder for a single adapter, callLLM(system, user), behind Layer 3. Swapping the provider changes one function and zero business logic. Pick the platform the organisation already governs — the security and procurement review is the long pole, not the integration.
| Option | Choose when | Data posture |
|---|---|---|
| Oracle OCI Generative AI (Cohere Command, Llama) | EPM Cloud already lives in OCI; you want one cloud boundary and one contract | Prompts stay in the OCI tenancy; no training on customer data; dedicated AI clusters available for isolation |
| Azure OpenAI Service | Microsoft-first finance estate (Entra ID, Purview, Sentinel already in place) | Private endpoint, regional deployment, zero-retention by default under the enterprise agreement |
| AWS Bedrock (Claude, Titan) / Google Vertex AI (Gemini) | The org’s landing zone is AWS or GCP; VPC endpoints and IAM already audited | VPC/PSC private access, no data used for training, CloudTrail/Cloud Audit Logs integration |
| Direct enterprise API (Anthropic, OpenAI) | Fastest model access; acceptable when a zero-data-retention agreement and DPA are signed | ZDR endpoint, SSO-managed keys, SOC 2 report on file |
| Self-hosted open weights (Llama, Mistral, Qwen via vLLM) | Sovereign or air-gapped requirements; regulated data classification forbids any external inference | Full control; you own patching, eval, and capacity — budget for an MLOps owner |
Put a model gateway in front of whichever you choose (Azure API Management, OCI API Gateway, Kong AI Gateway, LiteLLM, or Portkey): it owns key custody, per-team spend caps, routing and fallback between models, prompt/response logging, and lets you retire a deprecated model without touching the application.
10.3 · Security controls
| Control | Implementation |
|---|---|
| Identity & access | Covered in full in section 09 — corporate SSO, group-to-role mapping, and the decision about who enforces data-level security. Listed here because it is a production gate, not because it is optional. |
| Service account | One read-only EPM service account per application per pod, least-privilege role, no interactive login, credential in a vault (OCI Vault, Azure Key Vault, HashiCorp Vault), rotated on a schedule and on staff change. |
| Secrets & config | No secrets in code or build artifacts; environment-specific config injected at deploy; .dev.vars-style files never leave a developer machine. |
| Network | Private endpoints to the LLM provider and to EPM where the platform supports them; egress allow-list so the orchestrator can reach exactly two hosts; TLS 1.2+ everywhere. |
| Prompt-injection & input guardrails | Layer 1 (already in the demo) blocks instruction-override patterns, enforces length and scope; extend with a classifier on the gateway and log every rejection. |
| Output guardrails | Layer 4 (already in the demo) validates every returned member against the approved lists and strips unexpected keys; production adds a policy check that the resolved POV is inside the user’s security scope before the data call. |
| Data minimisation | Prompts contain metadata and the user’s query only. No cell values, no employee names, no free-text comments from EPM. Logged prompts are classified and retained accordingly. |
| Encryption | In transit (TLS) and at rest (provider-managed KMS); audit logs on immutable storage with customer-managed keys where policy requires. |
10.4 · SOX, audit, and model-risk controls
A read-only NLQ layer does not change a financial-reporting control, but it is an interface to a SOX-relevant system and lands squarely in ITGC scope. Treat prompts, schemas, and eval sets as code — that single decision satisfies most of what an auditor will ask for.
| Requirement | How it is satisfied |
|---|---|
| Complete, immutable audit trail | Append-only log of user, timestamp, raw query, parsed JSON, model and prompt version hash, POV returned, and row count — WORM storage, retained for the evidence period (typically 7 years), exported to the SIEM. |
| Change management | Prompt templates, few-shot examples, approved-member schema, and code are version-controlled; every change follows ticket → peer review → test evidence → CAB approval → deploy. A prompt edit is a code change. |
| Segregation of duties | Developers cannot deploy to production; the service-account owner is not a developer; production secrets are held by platform operations. |
| Access recertification | Quarterly review of who can use the tool and of the service account’s EPM roles, evidenced and signed. |
| Testing evidence | A golden-query regression suite (the few-shot examples plus a larger labelled set) runs in CI before every release; pass rate and diffs are archived as release evidence. |
| Model risk management | An inventory entry (intended use, limitations, owner, validation date) in the model-risk register — the SR 11-7 pattern for financial services; periodic re-validation when the model or prompt changes. |
| Reproducibility & lineage | Every displayed number traces to an EPM POV and a consolidation/calculation timestamp; an auditor can re-query the same intersection in EPM and match it. |
| Explainability | The parsed intent JSON is the explanation and is shown to the user on every response — no hidden reasoning between the query and the data call. |
10.5 · Dev → Test → Prod promotion
| Environment | EPM target | Data | Gate to leave |
|---|---|---|---|
| Dev | EPM Test pod (developer slice) | Synthetic or masked | Unit tests on the data layer; lint; eval suite ≥ threshold against the Test LLM deployment |
| Test / UAT | EPM Test pod (full refresh) | Masked copy of production | Business UAT sign-off on the golden queries; security scan; performance run (p95 latency, fallback rate) |
| Prod | EPM Production pod | Live | Change ticket approved; deploy in window; smoke test; hypercare with rollback ready |
- Promoted artifacts: application build, prompt templates (versioned), approved-member schema snapshot, eval set, infrastructure config (IaC) — all from the same Git tag.
- Pipeline: branch → PR review → CI (tests + evals) → deploy to Test → UAT sign-off → CAB → deploy to Prod → smoke test. Hosting can stay on Cloudflare Pages/Workers or move to OCI Functions + API Gateway or the org’s standard platform — the code does not care.
- Configuration: per-environment secrets and endpoints injected at deploy; the same build runs in every environment.
- Metadata sync: a scheduled job refreshes dimension metadata into Layer 2 grounding with change detection, so a new entity or account appears in the approved lists without a code release.
- Rollback: previous build and previous prompt version retained; rollback is a redeploy, and because prompts are versioned it also reverts a prompt regression.
10.6 · Operating it
- SLOs: p95 latency, availability of the read path (the deterministic fallback keeps it alive when the LLM is down — already built), fallback rate as a quality signal, eval pass rate per release.
- Cost governance: per-user and per-team token budgets at the gateway; alert on anomalies; the unit-cost model earlier in this kit is the baseline.
- Model lifecycle: providers retire models on a schedule — re-run the eval suite on the successor before switching, and record the switch as a change.
- Incident runbook: LLM outage → fallback parser; EPM API outage → cached metadata with a stale banner; guardrail spike → review logs for injection attempts.
10.7 · What changes for Allocation Waterfall
| Concern | Production answer |
|---|---|
| System of record | Oracle PCMCS — allocation results and the trace from the PCMCS calculation engine (Trace Allocations REST/UI); pools, drivers, and rule sets are PCMCS model artifacts |
| Read/write posture | Read-only. Driver assignment and pool sizing are PCMCS model changes under change control. |
| Use-case-specific control | Driver data lineage matters to auditors: show the driver source (e.g. headcount from Workforce, revenue from Planning) and its snapshot date alongside every allocated amount. |
10.8 · Production readiness checklist
- ✓LLM platform selected from the governed list, DPA / zero-retention terms on file, gateway in front of it
- ✓SSO integrated; authorisation derived from EPM security groups; cookie gate removed
- ✓Read-only EPM service account per pod, credential in a vault, rotation scheduled
- ✓Prompts, schema, and eval set version-controlled and under change management
- ✓Append-only audit log wired to the SIEM with the agreed retention
- ✓Golden-query eval suite passing in CI; results archived as release evidence
- ✓Dev / Test / Prod pipeline with gates, IaC, and a rehearsed rollback
- ✓Model-risk register entry and owner named; first re-validation date set
- ✓Metadata refresh job scheduled with change detection
- ✓SLOs, cost caps, and the incident runbook agreed with platform operations