The problem this solves

A client reporting 99% on-time completion may still have a serious control failure. Completion measures that the process ran, not that the balances are right. What an auditor tests instead: aged reconciling items (a reconciliation can complete while carrying six-month-old unexplained differences), the denominator (99% of what? accounts never set up as profiles are invisible), rubber-stamp review (140 approvals in eleven minutes with no comments), segregation of duties, and supporting evidence (marked complete with no bank statement attached — the assertion exists, the proof does not).

Any mature reconciliation function is measured on aging, not on completion percentage. Build the aging report in phase one — it is the report that changes client behaviour. This use case is that report, plus the four others that reveal actual health.

Audience

Close managers who need portfolio status, bottlenecks, and completeness. Controllers who asked to review everything and should be given the dashboard instead of the queue. Internal audit and external auditors who want the evidence package and the exceptions, not the completion chart.

Every reconciliation in the period, aggregated

InputSource
NLQ query: “Show aging of reconciling items for August” or “High risk reconciliations this period” — DeepSeek resolves year, month, and an optional risk tier
The portfoliogenPortfolio(year, period) builds every scheduled reconciliation via the same genReconciliation() UC10 uses one at a time — 436 in August, 516 at quarter end
Reconciling items and agesFrom each reconciliation’s item list, bucketed 0–30 / 31–60 / 61–90 / 91–180 / 180+
Workflow historyWho approved or rejected, and how many seconds they spent — the timestamps ARCS records on every action
Loaded vs true balancesFor the completeness check against the trial balance

The compliance dashboard an auditor would actually use

  • Six KPIs, completion first and labelled a vanity metric: items aged over 90 days, aggregate unexplained difference, auto-reconciled share, late count, closed-without-evidence count
  • Aging chart: open reconciling items by bucket — the single most important view in ARCS
  • Status by risk tier: where the close actually stands, for the accounts that matter
  • Reviewer behaviour: approvals, rejections, average time, rejection rate, and the rubber-stamp flag
  • Completeness check: loaded total vs trial balance, the gap, and the profiles that loaded $0 and auto-passed
  • Threshold trap: aggregate permitted difference for a flat $10K rule versus percentage-plus-cap by tier, against group materiality

The portfolio view on top of UC10 and UC11

Reconciliation Compliance shows one reconciliation with its evidence; this page shows all of them and the patterns only the portfolio reveals — the reviewer who never rejects, the cluster of $0 auto-passes, the flat threshold that adds up to twice materiality. Transaction Matching’s unmatched balance arrives here through the cash reconciliations it feeds. Together the three are the ARCS module: assurance at transaction level, at balance level, and across the portfolio.

Recommended integration points

  • Close-manager stand-up: aging and late counts by tier, every morning of the close
  • Quarterly control review: reviewer table and completeness check as standing agenda items
  • Audit fieldwork: the evidence package is a by-product of the work — two to three weeks of manual assembly a year that this replaces often funds the project

The numbers

436
Reconciliations in August
5
Aging buckets
1
Rubber-stamp reviewer flagged
12
Silent $0 auto-passes caught
$0.0002
Cost per NLQ query
2.2×
Flat $10K threshold vs materiality

The honest checklist

  • ✓Your reconciliation KPI is completion percentage and audit findings keep arriving anyway
  • ✓Nobody can say what the current aging of reconciling items is — which is itself the finding
  • ✓A controller wants to approve everything, and you need to offer oversight instead of a queue
  • ✗Reconciliations are not yet in a system that timestamps preparer and reviewer actions — there is nothing to measure until UC10 exists
  • ✗The organisation is not ready for individual performance to become visible; this dashboard needs an executive sponsor, not just training

Try it now

The live demo aggregates every reconciliation in the chosen month and recomputes all six panels.

Launch Aging & Audit Evidence → Start a Lab engagement

Things to try

  • Type “Which reviewer is rubber-stamping?” and read the reviewer table: 38 approvals, 8 seconds each, 0% rejections. Then compare the Group Controller’s 89 approvals at nine minutes each
  • Change the flat threshold to $5,000 and watch the aggregate permitted difference still exceed group materiality — the arithmetic, not the number, is the problem
  • Filter to High risk and look at the aging chart — the 180+ bucket is where the planted suspense items live
HOW WE BUILT IT

Three audiences, three surfaces

AudienceWantsWhere
Preparers and reviewersMy work, what is due, what is lateWorklist
Close managersPortfolio status, bottlenecks, completenessCompliance dashboard
Auditors and controllersEvidence, aging, exceptions, segregationReports and BI

This use case serves the second and third. ARCS ships seeded reports and dashboards, supports custom report definitions, and can burst on a schedule; anything more sophisticated pushes data to Oracle BI or a warehouse — scope that explicitly rather than discovering it in UAT. Design guidance mirrors FCCS: a small set of parameterised reports rather than forty static ones, and the aging report built early.

Every reconciliation → aggregate → the five signals

Year + month + optional risk tier (from NLQ or the filters)
1
Build the portfolio
genPortfolio(year, period)
genProfiles().map(p ⇒ genReconciliation(p.id, year, period)).filter(scheduled)
  • Reuses UC10’s reconciliation function unchanged — the dashboard can never disagree with the drill-down
  • 436 reconciliations in a normal month; 516 at quarter end when Low-risk Variance profiles fall due
Aging
items bucketed by ageDays
  • 0–30 · 31–60 · 61–90 · 91–180 · 180+
  • Counted regardless of reconciliation status
Reviewer statistics
history where Approved | Rejected
  • Approvals, rejections, average seconds
  • Flag: ≥30 decisions, <20s, 0 rejections
Completeness
Σ loaded vs Σ true balance
  • Gap = the key-mismatch profiles
  • Listed by entity and account type
2
Threshold exposure
computeThresholdExposure(portfolio, policy)
flat: n × $10K  |  tiered: min(0.5% × balance, cap by tier)
  • Compared to group materiality (0.5% of the reconciled balance base)
  • The flat policy exceeds materiality; the tiered one does not — unit-tested
Renderer
aging.html
  • Six KPIs led by aging, two charts, reviewer table, completeness panel, editable threshold panel, workload and audit package
  • Risk-tier filter narrows the reconciliation set; reviewer and completeness views stay portfolio-wide

App UI — Component breakdown

ComponentBehaviour
KPI rowCompletion first, greyed and labelled vanity; then aged > 90, unexplained aggregate, auto-reconciled %, late, closed without evidence
Aging chartOpen reconciling items by bucket, green to red
Status by tierStacked Closed / Open with Reviewer / Open with Preparer / Rejected per risk tier
Reviewer tableSorted by approvals; sub-20-second averages and 0% rejection rates in red; rubber-stamp flag; the “oversight vs approval” note
Completeness panelLoaded total, trial balance total, gap, and the clustered list of $0 auto-passes with the “99% of what?” explanation
Threshold panelEditable flat threshold vs the tiered policy, each against materiality, with the better design stated

Completion percentage is the vanity metric

  • Aging of reconciling items — buckets at 30/60/90/180+ days. The single most important view in ARCS. An old unresolved item means an error nobody chased, a fraud, or a process that quietly broke.
  • Unexplained difference by account and in aggregate — catches the threshold-aggregation problem.
  • Late completion trend — is the process improving or drifting.
  • Auto-reconciliation rate — efficiency, and a check that it is not creeping into material accounts.
  • Rejection rate by reviewer — a reviewer with a 0% rejection rate over a year is not reviewing.
  • Preparer workload distribution — finds the one person carrying 300 reconciliations.

The build orders the KPI row in that spirit: completion is shown, because clients will ask for it, but it is greyed and labelled, and the five signals that reveal actual health sit beside it. Aging is computed from item ages rather than reconciliation status on purpose — a reconciliation can be Closed and still carry a 200-day item, and that item counts.

A silent failure wearing the costume of success

Every loaded balance must resolve to a profile, and profiles are keyed on Account ID. If the GL extract’s segment combination does not match exactly, the balance does not land — and the reconciliation sits at zero, which presents as a clean reconciliation rather than a failure. A zero balance with a nil difference auto-passes. It does not error, reject, or appear in an exception queue. This is the ARCS equivalent of unmapped intercompany in FCCS.

// The two defences, both in this build:
// 1. Build profiles from the GL extract, not a hand-maintained list  → genProfiles()
// 2. Every period, compare Σ loaded balance to the trial balance total:
completeness = {
  loadedTotal:       Σ reconciliation.glBalance,       // what landed
  trialBalanceTotal: Σ reconciliation.trueBalance,     // what the GL holds
  gap:               trialBalanceTotal − loadedTotal,  // ≠ 0 ⇒ accounts missing from the universe
  missing:           reconciliations where keyMismatch  // 3 entities × 4 account types
}

Where the failure clusters — here 12 profiles across three EMEA entities — it is structural rather than random: delimiter format, leading zeros stripped by Excel, segment count, or recently changed entity codes. The panel names the entities and account types so the first check is obvious. It also answers the denominator question directly: those 12 are not in the 1% either. They are invisible, which is why completeness of the account universe is tested early and manual processes routinely fail it.

Three findings the portfolio view exists to surface

“Auto-reconcile anything under $10,000”

Four problems. Aggregation: 436 accounts each permitted $10,000 of unexplained difference is $4.36M in total — 2.2× group materiality — and “each was under threshold” is not a defence. Flat thresholds ignore account size: $10,000 is rounding on $2Bn and 83% of a $12,000 petty cash float; use the lesser of a percentage and an absolute cap, tuned by risk tier, which the panel shows landing at 0.35× materiality. Systematic error hides inside “immaterial”: 1,400 accounts all understated in the same direction is one broken interface, not noise. Auto-closing is the wrong response even where the threshold is right: a $9,000 difference persisting twelve months is a problem whatever its size. Better design: auto-close only on a nil difference, and use thresholds to route small differences to lighter review rather than none.

The reviewer who never rejects

ARCS timestamps everything, which is exactly why rubber-stamping becomes visible. The reviewer table computes, from the Approved and Rejected history entries, each reviewer’s decision count, average seconds, and rejection rate; one reviewer in this build approves in 4–12 seconds and rejects nothing, and is flagged when those three facts coincide over thirty or more decisions. The remedy is not a reprimand but a redesign: risk-tier the review so the controller takes only High-risk profiles, use summary profiles so related accounts approve as one, and give them the dashboard rather than the queue. What they usually want is assurance nothing is missed — a completeness and aging view. Oversight and approval are different needs.

Segregation of duties and the audit package

ARCS prevents preparer and reviewer being the same user; it cannot detect that the reviewer reports to the preparer. That assessment is design-time work, done against the org chart, and it belongs on the same quarterly agenda as the reviewer table. The audit evidence package — the reconciliation with balances and items, attached documentation, the full timestamped workflow history, questions and checklist responses, and the period lock — is a by-product of doing the work, not a separate exercise. Clients typically spend two to three weeks a year assembling it manually; that saving alone often funds the project.

The portfolio schema: risk tier, year, month

Aging’s NLQ schema is {intent, risk, year, period, confidence} — the same minimal shape as Ownership and Profitability, because the output is always the whole portfolio and the only things worth resolving are the period and an optional risk-tier filter. The few-shot examples deliberately include questions that name a metric rather than a filter — “which reviewer is rubber-stamping?”, “completeness check against the trial balance” — and map them to the dashboard for the right period, since every metric renders together and the user’s eye goes to the panel they asked about.

Cost model — Cost per query

ComponentDetailCost
LLM input tokens~600 tokens (schema block + few-shot examples + query)$0.000084
LLM output tokens~45 tokens$0.0000126
Portfolio build (436 reconciliations) + metricsClient-side, deterministic, once the query resolves$0
Total per query~$0.0001
With 2× safety margin~$0.0002

Year, month, and a 3-value risk enum — the same footprint as Ownership. The portfolio itself is 516 profile evaluations per query, client-side, in tens of milliseconds; production ARCS computes these views on its own dashboards.

Key files

FileRole
epm-nlq-src/assets/epm-arcs-data.jsgenPortfolio() (every profile’s reconciliation for the period, aggregated), ARCS_AGING_BUCKETS, computeThresholdExposure(), the reviewer statistics with the rubber-stamp flag, and the completeness check
epm-nlq-src/pages/aging.htmlUI: NLQ bar, year/month/risk filters, six KPIs led by aging not completion, aging and status-by-tier charts, reviewer table, completeness panel, threshold trap with an editable flat threshold, workload and audit package
epm-nlq-src/assets/epm-fccs-data.jsgenConsolidationBS() — GL balances come from the same balance sheet FCCS consolidates, so ARCS proves the balances FCCS combines
epm-nlq-src/assets/epm-data.jsShared ENTITIES, PERIODS, ENTITY_SEEDS, SEASON
functions/api/nlq-query.jsShared NLQ endpoint; this use case uses useCase: "aging"

Why ARCS is a separate data file from FCCS

They answer different questions and are separately licensed products. ARCS proves each balance is real; FCCS combines them into group results — assurance, then aggregation. Keeping epm-arcs-data.js separate mirrors that: it consumes the FCCS balance sheet for its GL balances but adds its own objects (profiles, formats, reconciliations, match types, rules) and never writes back. The deterministic FNV-1a hash (arcsRand) means every reconciliation and every bank row is reproducible on any machine, which is what makes the trap scenarios teachable.

Tech stack — Every tool in this build

LayerToolWhy
Data layerepm-arcs-data.js (vanilla JS)Portfolio aggregation over the same deterministic reconciliations UC10 shows one at a time — the LLM only resolves period and tier
LLMDeepSeek V3Shared NLQ endpoint; year, month, and an optional risk tier — the minimal portfolio schema
ChartsChart.js 4.xAging buckets and status stacked by risk tier; shared library with the other kits
Edge hostingCloudflare PagesStatic file, no server compute needed for this use case
BuildEleventy v3.1.5Copies epm-nlq-src/pages/ and epm-nlq-src/assets/ to _site/ verbatim

Known attack surfaces

ThreatMitigation in this build
Reporting completion % as control effectivenessCompletion is shown first and labelled “vanity metric”; aged items, unexplained aggregate, missing evidence, and late counts sit beside it so the contrast is unavoidable
Accounts missing from the profile universe are invisible in every metricThe completeness check compares total loaded balance to the trial balance total every period and lists the profiles that loaded $0 — the denominator is tested, not assumed
Auto-reconciliation creeping into material accounts under a flat thresholdThe threshold panel sums permitted differences across the portfolio against group materiality for a flat policy and for a percentage-plus-cap-by-tier policy, with the flat threshold editable
A reviewer approving in bulk without opening anythingApproval durations and rejection rates are computed from the timestamped history; a reviewer with ≥30 decisions, <20s average, and zero rejections is flagged

Guardrails — What prevents bad control math

  • Aging is computed from item ages, not from reconciliation status: a closed reconciliation can still carry a 200-day item, and it counts
  • The completeness gap reconciles exactly to the planted key mismatches — unit-tested, so the panel cannot show a gap that the profile list does not explain
  • Threshold exposure is additive by construction: permitted difference is summed per reconciliation, which is the arithmetic “each was under threshold” forgets
  • Reviewer statistics come only from Approved/Rejected history entries — a reviewer with open work is not penalised for the queue they have not reached

Who is asking, and what are they allowed to see?

The demo answers neither question — it has a cookie gate and no notion of a user. In production these are the two questions everything else rests on, and they have different answers: authentication is who you are, authorization is what you may see. Corporate SSO settles the first. Only Oracle EPM Cloud can settle the second, and the single most important rule in this section is that this application must never become the place where that decision is made.

The rule that governs every choice below: a user must see exactly what they would see by logging into Oracle EPM Cloud directly — no more, and no less. If this tool can surface a number the user could not retrieve themselves, it has become a privilege-escalation path, and it will be found in the first access review.

9.1 · The identity chain, end to end

Finance user opens the tool in a browser — no local account, no password held here
1
Corporate identity provider
Entra ID · Okta · OCI IAM
OIDC Authorization Code + PKCE  (or SAML 2.0)
  • MFA and Conditional Access are enforced here — device compliance, location, risk signals
  • Returns an ID token (who the user is) and an access token (what they may call)
  • Group membership arrives as a claim; the application never handles a password
2
Application session
validate, never trust
  • Verify signature, issuer, audience and expiry against the IdP’s published keys
  • Read the group claims — there is no local user table and no local role table
  • Short-lived access token with refresh-token rotation; session timeout set to the data classification
Pattern A — identity propagation
OAuth 2.0 token exchange (on-behalf-of)
  • The API is called as the user
  • EPM enforces its own security natively
  • Audit trail names the real user
  • Preferred where the API supports it
Pattern B — service account + filtering
one read-only integration account
  • The application becomes the enforcement point
  • Entitlements fetched separately, applied in one audited place
  • Simpler and cacheable — and a filtering bug is a data breach
3
EPM identity domain
roles + dimension security
  • Roles: Service Administrator, Power User, User, Viewer — assigned to groups, never to individuals
  • Data level: Portfolio-wide by nature, which is precisely why it needs its own entitlement rather than inheriting a preparer’s
  • Group → role mapping lives in the platform, not in this application
Result, filtered to this user
aging.html
  • The user sees exactly what they would see logging into the source system directly — no more
  • Every query logged against the real end user, never a shared account

9.2 · Federating the corporate identity provider

Oracle EPM Cloud does not replace your directory — it trusts it. The EPM Cloud identity domain is federated with the corporate IdP so authentication happens where it already happens, under policies security has already written.

Identity providerProtocolNotes
Microsoft Entra ID (formerly Azure AD)SAML 2.0 or OIDCThe common case. Conditional Access, MFA and device compliance are enforced at Entra and inherited automatically. On-premises Active Directory federates through Entra Connect rather than being integrated directly.
OktaSAML 2.0 or OIDCSame pattern; Okta groups drive EPM roles through SCIM provisioning.
OCI IAM (identity domains)NativeAlready present with Oracle EPM Cloud. Can be the primary IdP for a smaller estate, or a federated spoke of Entra/Okta for a larger one.
AD FSSAML 2.0Still seen where the estate is not yet cloud-first. Works, but you inherit the on-premises availability of the token service — if AD FS is down, nobody logs in.

For the browser application itself, use OIDC Authorization Code flow with PKCE. Not the implicit flow, which is deprecated and leaks tokens through the URL, and never a resource-owner password grant — a finance tool should not be capable of handling a password at all.

9.3 · From group membership to EPM roles

Roles are granted to groups, never to individuals, and the groups come from the directory. That one discipline is what makes joiner/mover/leaver work without anyone having to remember this application exists.

Entra ID group                    →  EPM role / entitlement
──────────────────────────────────────────────────────────────────
FIN-EPM-Analysts                  →  Planning User
FIN-EPM-Controllers-EMEA          →  Power User + EMEA data scope
FIN-EPM-Admins                    →  Service Administrator
──────────────────────────────────────────────────────────────────
Provisioned by SCIM. Remove the user from the group and the
entitlement disappears on the next sync — including here.

For this use case the relevant native entitlement is: ARCS Power User or the compliance-dashboard entitlement — wider than a preparer, narrower than everyone.

9.4 · The architectural decision: who enforces?

This is the choice that determines whether the deployment is defensible. Both patterns appear in the diagram above; the difference is where the security boundary actually sits.

Pattern A — identity propagationPattern B — service account + filtering
HowThe user’s token is exchanged (OAuth 2.0 on-behalf-of) for one scoped to the EPM API; calls are made as the userA single read-only integration account calls the API; the application filters the results
Enforcement pointOracle EPM CloudThis application
Audit trail showsThe real end userThe service account — you must log the real user separately
Failure modeToken plumbing is more complex; per-user rate limits applyA filtering bug is a data breach, and the entitlement copy drifts from reality
VerdictPrefer this wherever the API supports user-token authenticationAcceptable with discipline: narrowest possible service account, filtering centralised in one tested place, real user in every log line

The shortcut to refuse. Pattern B built with a Service Administrator account and no filtering at all is the most common way this gets delivered, because it works perfectly in UAT — testers are usually over-entitled, so nobody notices that everyone can see everything. It fails at the first access review, and by then it is in production with real users depending on it.

9.5 · Data-level security is the part that matters

Role membership decides whether a user can open the application. It does not decide which rows they get back, and confusing the two is the most expensive mistake available here.

  • For this use case: Portfolio-wide by nature, which is precisely why it needs its own entitlement rather than inheriting a preparer’s.
  • Apply it before aggregation, not after. Filtering a total that has already been computed across entities the user cannot see still leaks the total.
  • The NLQ layer needs its own check. Layer 4 already validates that the resolved point of view uses approved members; production adds a second test — that the resolved POV sits inside this user’s scope — and it runs before the data call, not after. A natural-language interface is very good at asking for things politely; the authorization check must not care how the question was phrased.
  • Fail closed. If entitlements cannot be resolved, return nothing and say so. An empty result is a support ticket; a permissive default is an incident.

9.6 · Provisioning, sessions and the leaver problem

  • SCIM provisioning from Entra or Okta into the EPM Cloud identity domain (OCI IAM, formerly IDCS), covering joiner, mover and leaver. The mover is the case people forget — somebody changing region should lose the old scope, not accumulate both.
  • No local user store. If this application keeps its own copy of who may do what, a leaver keeps access until somebody remembers to update it. Nobody ever does.
  • Short-lived access tokens with refresh-token rotation; align session timeout with the data classification rather than with convenience.
  • MFA and Conditional Access at the IdP — not reimplemented here. Device compliance and location policy come free with federation.
  • Quarterly recertification of both the groups that grant access and the service account’s own entitlements, evidenced and signed.
  • Break-glass access is a named, monitored, time-boxed account — never a shared credential in a password manager.

9.7 · What this means for Aging & Audit Evidence

ConcernAnswer for this use case
Native entitlement requiredARCS Power User or the compliance-dashboard entitlement — wider than a preparer, narrower than everyone
Data-level controlPortfolio-wide by nature, which is precisely why it needs its own entitlement rather than inheriting a preparer’s
Use-case-specific sensitivityThe reviewer-behaviour panel makes individual performance visible — approval counts, average seconds, rejection rates. That is a people-data question as much as a security one: agree the audience with HR, consult the works council where required, log every view of it, and give it an executive sponsor before it goes live.

9.8 · Security configuration checklist

  • ✓Oracle EPM Cloud federated with the corporate IdP over SAML 2.0 or OIDC; the cookie gate removed entirely
  • ✓Browser app uses OIDC Authorization Code + PKCE — no implicit flow, no password grant
  • ✓MFA and Conditional Access enforced at the IdP, not reimplemented in the application
  • ✓Roles granted to directory groups, never to individuals; SCIM covers joiner, mover and leaver
  • ✓Enforcement pattern chosen deliberately — Pattern A where the API supports it, or Pattern B with filtering centralised and tested
  • ✓Data-level security applied before aggregation, and the resolved POV checked against the user’s scope before the data call
  • ✓No local user table and no local role table anywhere in the application
  • ✓Every query logged against the real end user, even when a service account makes the call
  • ✓Authorization failures fail closed and are logged as security events rather than swallowed
  • ✓Quarterly recertification of access groups and of the service account’s own entitlements
Talk through your identity model → Back to the demo

From demo to a governed enterprise deployment

Everything above runs on synthetic data, a public LLM API key, a cookie gate, and no audit trail — deliberately, so the mechanics are inspectable. Taking Aging & Audit Evidence to production is not a rewrite; the 4-layer pipeline and the data-layer contract survive intact. It is a controlled-change program across six workstreams: architecture, LLM platform, security, SOX/audit, environment promotion, and operations. This section is the checklist we run with clients.

The one rule that matters most for this use case: in the demo the browser computes the result; in production Oracle computes and this layer retrieves and explains. Never ship a second calculation engine that can disagree with the system of record — the moment two numbers exist, the audit question becomes “which one is right,” and the answer must always be the EPM module.

10.1 · Production reference architecture

Finance user · corporate SSO (OIDC/SAML + MFA) · EPM role claims
1
Edge / API Gateway
WAF · rate limit · identity
  • Terminates SSO, validates the session, attaches the user’s EPM groups to the request
  • Rate limits per user, blocks anonymous access, scrubs PII patterns before anything reaches the orchestrator
2
NLQ Orchestrator
the 4-layer pipeline, hardened
L1 guardrails → L2 grounding → L3 LLM adapter → L4 eval + fallback
  • L2 grounding reads dimension metadata from EPM on a schedule — not a hardcoded schema
  • Only the schema + user query go to the model; financial values never leave the data layer
  • L4 rejects anything outside the approved member lists and falls back to the deterministic parser
Oracle OCI Generative AI
same tenancy as EPM Cloud
  • Data stays inside the OCI boundary
  • Natural fit when EPM is already in OCI
Azure OpenAI / AWS Bedrock / Vertex AI
private endpoint, zero retention
  • Use the hyperscaler the org already governs
  • Enterprise DPA, no training on prompts
Self-hosted open weights
VPC / air-gapped
  • For regulated or sovereign data
  • Highest control, highest run cost
3
EPM Data Layer
Oracle EPM REST API
  • Oracle ARCS compliance dashboards and report queries — aging, late completion, auto-reconciliation rate, reviewer history — via the ARCS REST API, or a governed extract to Oracle BI / the warehouse for anything more sophisticated
  • Least-privilege service account (read-only role, one app, one pod) with the token in a vault and rotated
  • Results filtered to the requesting user’s EPM security before rendering
4
Audit & Observability
append-only
  • Every query logged: user, timestamp, raw query, parsed intent JSON, model + prompt version, POV returned, latency, cost
  • Exported to the SIEM; retained per the SOX evidence schedule
  • Dashboards for fallback rate, eval pass rate, guardrail hits, p95 latency, spend
Rendered result + evidence trail
aging.html
  • The parsed JSON is shown to the user as the explanation (“AI: entity · year · scenario — 93% confident”) — the same line the demo prints today
  • Every number on screen traces to an EPM cell intersection an auditor can reproduce

10.2 · Choosing the LLM platform

The demo’s DeepSeek call is a placeholder for a single adapter, callLLM(system, user), behind Layer 3. Swapping the provider changes one function and zero business logic. Pick the platform the organisation already governs — the security and procurement review is the long pole, not the integration.

OptionChoose whenData posture
Oracle OCI Generative AI (Cohere Command, Llama)EPM Cloud already lives in OCI; you want one cloud boundary and one contractPrompts stay in the OCI tenancy; no training on customer data; dedicated AI clusters available for isolation
Azure OpenAI ServiceMicrosoft-first finance estate (Entra ID, Purview, Sentinel already in place)Private endpoint, regional deployment, zero-retention by default under the enterprise agreement
AWS Bedrock (Claude, Titan) / Google Vertex AI (Gemini)The org’s landing zone is AWS or GCP; VPC endpoints and IAM already auditedVPC/PSC private access, no data used for training, CloudTrail/Cloud Audit Logs integration
Direct enterprise API (Anthropic, OpenAI)Fastest model access; acceptable when a zero-data-retention agreement and DPA are signedZDR endpoint, SSO-managed keys, SOC 2 report on file
Self-hosted open weights (Llama, Mistral, Qwen via vLLM)Sovereign or air-gapped requirements; regulated data classification forbids any external inferenceFull control; you own patching, eval, and capacity — budget for an MLOps owner

Put a model gateway in front of whichever you choose (Azure API Management, OCI API Gateway, Kong AI Gateway, LiteLLM, or Portkey): it owns key custody, per-team spend caps, routing and fallback between models, prompt/response logging, and lets you retire a deprecated model without touching the application.

10.3 · Security controls

ControlImplementation
Identity & accessCovered in full in section 09 — corporate SSO, group-to-role mapping, and the decision about who enforces data-level security. Listed here because it is a production gate, not because it is optional.
Service accountOne read-only EPM service account per application per pod, least-privilege role, no interactive login, credential in a vault (OCI Vault, Azure Key Vault, HashiCorp Vault), rotated on a schedule and on staff change.
Secrets & configNo secrets in code or build artifacts; environment-specific config injected at deploy; .dev.vars-style files never leave a developer machine.
NetworkPrivate endpoints to the LLM provider and to EPM where the platform supports them; egress allow-list so the orchestrator can reach exactly two hosts; TLS 1.2+ everywhere.
Prompt-injection & input guardrailsLayer 1 (already in the demo) blocks instruction-override patterns, enforces length and scope; extend with a classifier on the gateway and log every rejection.
Output guardrailsLayer 4 (already in the demo) validates every returned member against the approved lists and strips unexpected keys; production adds a policy check that the resolved POV is inside the user’s security scope before the data call.
Data minimisationPrompts contain metadata and the user’s query only. No cell values, no employee names, no free-text comments from EPM. Logged prompts are classified and retained accordingly.
EncryptionIn transit (TLS) and at rest (provider-managed KMS); audit logs on immutable storage with customer-managed keys where policy requires.

10.4 · SOX, audit, and model-risk controls

A read-only NLQ layer does not change a financial-reporting control, but it is an interface to a SOX-relevant system and lands squarely in ITGC scope. Treat prompts, schemas, and eval sets as code — that single decision satisfies most of what an auditor will ask for.

RequirementHow it is satisfied
Complete, immutable audit trailAppend-only log of user, timestamp, raw query, parsed JSON, model and prompt version hash, POV returned, and row count — WORM storage, retained for the evidence period (typically 7 years), exported to the SIEM.
Change managementPrompt templates, few-shot examples, approved-member schema, and code are version-controlled; every change follows ticket → peer review → test evidence → CAB approval → deploy. A prompt edit is a code change.
Segregation of dutiesDevelopers cannot deploy to production; the service-account owner is not a developer; production secrets are held by platform operations.
Access recertificationQuarterly review of who can use the tool and of the service account’s EPM roles, evidenced and signed.
Testing evidenceA golden-query regression suite (the few-shot examples plus a larger labelled set) runs in CI before every release; pass rate and diffs are archived as release evidence.
Model risk managementAn inventory entry (intended use, limitations, owner, validation date) in the model-risk register — the SR 11-7 pattern for financial services; periodic re-validation when the model or prompt changes.
Reproducibility & lineageEvery displayed number traces to an EPM POV and a consolidation/calculation timestamp; an auditor can re-query the same intersection in EPM and match it.
ExplainabilityThe parsed intent JSON is the explanation and is shown to the user on every response — no hidden reasoning between the query and the data call.

10.5 · Dev → Test → Prod promotion

EnvironmentEPM targetDataGate to leave
DevEPM Test pod (developer slice)Synthetic or maskedUnit tests on the data layer; lint; eval suite ≥ threshold against the Test LLM deployment
Test / UATEPM Test pod (full refresh)Masked copy of productionBusiness UAT sign-off on the golden queries; security scan; performance run (p95 latency, fallback rate)
ProdEPM Production podLiveChange ticket approved; deploy in window; smoke test; hypercare with rollback ready
  • Promoted artifacts: application build, prompt templates (versioned), approved-member schema snapshot, eval set, infrastructure config (IaC) — all from the same Git tag.
  • Pipeline: branch → PR review → CI (tests + evals) → deploy to Test → UAT sign-off → CAB → deploy to Prod → smoke test. Hosting can stay on Cloudflare Pages/Workers or move to OCI Functions + API Gateway or the org’s standard platform — the code does not care.
  • Configuration: per-environment secrets and endpoints injected at deploy; the same build runs in every environment.
  • Metadata sync: a scheduled job refreshes dimension metadata into Layer 2 grounding with change detection, so a new entity or account appears in the approved lists without a code release.
  • Rollback: previous build and previous prompt version retained; rollback is a redeploy, and because prompts are versioned it also reverts a prompt regression.

10.6 · Operating it

  • SLOs: p95 latency, availability of the read path (the deterministic fallback keeps it alive when the LLM is down — already built), fallback rate as a quality signal, eval pass rate per release.
  • Cost governance: per-user and per-team token budgets at the gateway; alert on anomalies; the unit-cost model earlier in this kit is the baseline.
  • Model lifecycle: providers retire models on a schedule — re-run the eval suite on the successor before switching, and record the switch as a change.
  • Incident runbook: LLM outage → fallback parser; EPM API outage → cached metadata with a stale banner; guardrail spike → review logs for injection attempts.

10.7 · What changes for Aging & Audit Evidence

ConcernProduction answer
System of recordOracle ARCS compliance dashboards and report queries — aging, late completion, auto-reconciliation rate, reviewer history — via the ARCS REST API, or a governed extract to Oracle BI / the warehouse for anything more sophisticated
Read/write postureRead-only. Metrics are computed by ARCS on the locked period; this layer selects, filters, and narrates them.
Use-case-specific controlReviewer-behaviour analytics make individual performance visible for the first time — agree the audience with HR and, where required, the works council; log every view of the reviewer table; and pair the dashboard with an executive sponsor, because this is an organisational reaction, not a training issue.

10.8 · Production readiness checklist

  • ✓LLM platform selected from the governed list, DPA / zero-retention terms on file, gateway in front of it
  • ✓SSO integrated; authorisation derived from EPM security groups; cookie gate removed
  • ✓Read-only EPM service account per pod, credential in a vault, rotation scheduled
  • ✓Prompts, schema, and eval set version-controlled and under change management
  • ✓Append-only audit log wired to the SIEM with the agreed retention
  • ✓Golden-query eval suite passing in CI; results archived as release evidence
  • ✓Dev / Test / Prod pipeline with gates, IaC, and a rehearsed rollback
  • ✓Model-risk register entry and owner named; first re-validation date set
  • ✓Metadata refresh job scheduled with change detection
  • ✓SLOs, cost caps, and the incident runbook agreed with platform operations
Plan a production rollout with us → Back to the demo