INDEPENDENT AI AGENT AUTHORITY ASSURANCE

Know what your AI agent can actually do.

Aurelis evaluates whether an AI agent operates within its declared authority, challenges consequential behavior, and produces evidence for remediation and retesting.

AGENTAUTHORITYBOUNDARY

01 / AUTHORITY MODEL

The system is defined by its boundary.

Authority is treated as a relationship between identity, resource, operation, policy and context. Select a node or relationship to inspect the model.

DRAG TO ROTATE • SELECT A NODE TO INSPECT
DECLAREDWhat the agent is supposed to be allowed to do.
OBSERVEDWhat the agent can actually attempt under challenge.
CONSEQUENCEWhat the difference means for a business workflow.

02 / ASSESSMENT METHODOLOGY

From declared authority to tested evidence.

01

Discover

Map the agent, owner, tools, resources, permissions, actions and dependencies.

02

Define

Establish the intended authority boundary and the conditions under which actions are permitted.

03

Challenge

Test consequential and failure-path behavior against the declared boundary.

04

Analyze

Compare expected decisions with observed capabilities and execution traces.

05

Remediate

Translate findings into specific control changes and ownership decisions.

06

Retest

Repeat affected tests and establish whether the boundary now holds.

03 / CONTROLLED CHALLENGE

Challenge an agent.

Select a consequential scenario. The demonstration uses illustrative data and does not access a live system.

DEMONSTRATION / CONTROLLED TESTIssue $4,500 refund
IDENTITYverified
AGENTrecognized
POLICYapproval required
BOUNDARYnot satisfied
DECISIONDENY
Ready for controlled execution.

04 / FINDINGS

Every conclusion is tied to observed behavior.

HIGH

AUTH-007

Excessive Financial Authority

The agent is declared unable to issue refunds above the configured threshold without approval. The controlled scenario demonstrates an observed capability inconsistent with that boundary.

05 / EVIDENCE & DECISION EXPLORER

Follow the decision, not the headline.

→→→→→

EVIDENCE TRACE

Policy evaluation

Expected decision: DENY

Observed capability: ALLOW

Assessment
AUR-DEMO-0041
Test
AUTH-007
Policy version
3.2
Environment
Illustrative demonstration

06 / ASSURANCE REPORT

A report that can be taken into the organization.

The assessment record is structured around scope, controls, tests, findings, evidence, remediation and retesting rather than a generic score.

EXECUTIVE CONCLUSION
CONTROL ASSESSMENT
TEST RESULTS
FINDINGS & EVIDENCE
REMEDIATION
RETEST DETERMINATION

07 / TRUST & SECURITY

Assurance requires a defined scope.

Aurelis separates public demonstration material from client assessment data and treats conclusions as evidence within the conditions under which an assessment was performed.

Assessment isolation

Client assessments are scoped independently. Findings, evidence and report artifacts are associated with the designated assessment rather than the public demonstration environment.

Controlled access

Client access is issued for the relevant assessment scope and can be revoked when the engagement is concluded. No public route exposes customer assessment records.

Data minimization

Aurelis requests only information required for the agreed assessment. Where representative or controlled environments are sufficient, unrestricted production access is not required.

Evidence integrity

Assessment results are associated with an assessment identifier, test identifier, policy state, execution context and evidence reference so conclusions remain traceable to observed behavior.

Scope of assurance

An assurance assessment provides evidence within the defined scope and conditions of assessment. It is not a guarantee that an AI system cannot behave outside those conditions.

08 / THE ASSURANCE BOUNDARY

Authority is the first boundary. Assurance can extend beyond it.

The initial Aurelis engagement is deliberately narrow. The same assurance model can progressively address adjacent agent control surfaces.

AURELIS
AUTHORITY

09 / ASSESSMENT

Give Aurelis one agent. We will assess its authority boundary.

Define the workflow, establish the intended authority, challenge consequential behavior, document findings and retest the affected controls.