In brief
- AI now arrives inside every tool the portfolio office runs, and each agent governs itself.
- An AI north star names the decisions AI must improve and the measures that prove it, against a baseline.
- The high-value uses run in order of maturity, from articulating risk to forecasting actuals. A forecast is relied on only after back-testing on your own closed programmes.
- Set the bar centrally, once, and soon. Otherwise every team builds its own, and AI ends up competing with AI.
Why is AI in the portfolio office fragmenting?
Because AI now arrives inside every tool the office already runs. The portfolio platform, the finance system and the delivery partner's portal each bring an agent, and each works from its own data and governs only itself.
The first wave was personal. The second is agentic: it updates records, chases owners and drafts gate papers, faster than the governance around it can follow.
APRA saw the same pattern in the large banks, insurers and superannuation trustees it examined in late 2025. Its letter of 30 April 2026 reports boards relying on vendor presentations, governance lagging adoption, and little continuous monitoring for model drift. A portfolio office is exposed in the same way.
What is an AI north star for a transformation office?
It is a one-page statement of the decisions AI must make better and the measures that will prove it. A proposed use that improves none of those decisions waits.
- The decisions. Which investments proceed or stop, when to re-baseline or release contingency, how to share scarce people, and when a benefit counts as realised.
- The measures. Forecast error against actuals, how many weeks earlier a warning arrives, time from first signal to decision, hours per reporting cycle, and how often a person corrects an AI output.
- The baseline. Each measure taken before any tool goes live.
What does good look like in AI standards, value and quality?
The published frameworks agree on the core: know every AI use, give each an accountable owner, assess its impact, test it against written acceptance criteria before release, and monitor it afterwards.
| Source | What it asks of the office |
|---|---|
| Guidance for AI adoption, National AI Centre | Six essential practices: decide who is accountable, understand impacts and plan accordingly, measure and manage risks, share essential information, test and monitor, and maintain human control. |
| Policy for the responsible use of AI in government, version 2.0 | In effect for non-corporate Commonwealth entities, with some exceptions, from 15 December 2025. Each in-scope use case is registered with a risk rating and an accountable owner, assessed for impact from the design stage, and reviewed at least every 12 months once deployed. |
| APRA letter to industry on AI, 30 April 2026 | An inventory of AI tooling and use cases, lifecycle ownership, human involvement in high-risk decisions, and proportionate monitoring. |
| AS ISO/IEC 42001:2023 | The AI management system standard, adopted by Standards Australia in February 2024. |
| NIST AI Risk Management Framework | A voluntary framework built on four functions: govern, map, measure and manage. |
Value is measured against the north star. Quality comes down to three numbers: how often an output traces to its source record, how often a person corrects it, and, for a forecast, the error measured on closed work.
Which uses of AI add the most value in a transformation office?
Seven, in rough order of the maturity each needs. The early ones need only well-kept documents; the later ones need years of consistent portfolio history.
- Articulating risk. Turning meeting notes and RAID entries into risks stated as cause, event and consequence, each with a proposed owner for a person to accept or reject.
- Reviewing business cases against written criteria. Every claim traced to its evidence and every case scored the same way, so divisions can be compared. A person decides every rating.
- Reading emerging indicators. Comparing the status a programme reports with what its schedule, cost and RAID data show. Milestones that move while the finish date holds are a typical signal.
- Predictive constraint analysis. Finding where programmes will collide over the same people, environments, suppliers or decisions, months before it shows in any single plan.
- Long-range forecasting of actuals. Burn, cost at completion and finish dates as ranges with a stated confidence, relied on only once the method has been back-tested on your own closed programmes and its error measured.
- Anticipating decisions. Telling the board which decisions the portfolio will need from it next quarter, so the papers arrive before the decision is forced.
- Tracking benefits realisation. Following each benefit from the business case to the operational data that proves it, with forecast realisation stated as a range and its error reported.
The honest claim for forecasting is a stated confidence and a measured error. A promise of accuracy with no back-test on your own history is an opinion.
What does an AI maturity blueprint for a portfolio office look like?
Our Blueprint has five levels, from Individual to Orchestrated, assessed across six dimensions. Each level adds a capability and the governance that makes it safe to rely on.
| Level | What is true | Uses it supports |
|---|---|---|
| Level 1: Individual | People use AI tools on their own. The office cannot say which, or with what data. | Ad hoc drafting |
| Level 2: Sanctioned | Approved tools in approved environments. Every use registered, every output reviewed by a person. | Status packs and risk articulation |
| Level 3: Consolidated | One portfolio data spine and one published bar. Every output traces to its sources. | Business case review and emerging indicators |
| Level 4: Predictive | Forecasts carry ranges and a published track record. | Constraint analysis, forecasts of actuals, benefits tracking |
| Level 5: Orchestrated | Bounded agents run defined workflows inside a decision rights matrix, and the office's AI is independently reviewed each year. | Anticipated decisions, and agents that validate data and draft gate papers |
The six dimensions are value and use cases, data foundations, platform and security, governance and the bar, people and operating model, and measurement. The step that opens level 4 later is cheap and usually missed: keep monthly snapshots of portfolio data from now on.
Precision perspective
What happens when AI competes with AI?
The committee receives several confident answers to one question and has no way to choose between them. Time meant for the decision goes on reconciling the tools.
- Two ratings for one programme. The platform's risk agent rates a programme amber. Finance's model, on different data, rates it red. Neither can show its working against a common standard.
- Two forecasts for one finish date. The delivery partner's agent forecasts on-time completion from its own task data. The office's model, reading dependency history, forecasts a slip of two quarters. The board paper and the contract report now cite different dates.
- Cases written to pass the screen. Proponents draft business cases with AI and tune them against the AI that screens them, until the screen measures how well a case was written for it.
- Each tool marking its own work. Every agent arrives with its vendor's guardrails and reports on its own performance, and nobody holds the evidence to compare them.
None of this requires the technology to fail. It requires only that nobody set the bar first.
Why build the bar centrally, once, and now?
Because a standard is cheap to set before teams configure their own tools and expensive to retrofit afterwards. Set it centrally and let teams build inside it.
- One register of every AI tool and use, each with an accountable owner and a risk rating.
- One data spine: shared identifiers across schedule, cost, RAID, resources and benefits, so every tool reads the same numbers.
- One set of criteria for business cases, risk ratings and status, applied the same way by every tool and every person.
- One evaluation standard: acceptance tests every model passes before release and after each change.
- One decision rights matrix saying what AI may draft, flag, recommend or carry out, and what only a person decides.
The platforms are being configured now, team by team, and every configuration made before the bar exists is one to unpick later.
What comes after agentic AI in the portfolio office?
Coordination: agents that work under one register, are tested in one evaluation harness, and explain their forecasts to the people who rely on them.
- Many agents, one register. Each agent holds its own identity and permissions, logged like a staff member's. APRA found access controls had not yet adjusted to AI agents.
- Evaluation harnesses. Closed programmes with known outcomes, run against every model before release and after every update, so drift shows up as a number.
- Forecasts that explain themselves. A range, the drivers behind it, the closed programmes it resembles and the error the method showed on them. A board can challenge that.
- Continuous assurance. APRA observes that point-in-time, sample-based assurance is ill suited to models that learn, adapt and degrade. Monitoring has to run as long as the AI does.
What should an executive ask on Monday?
Five questions show where the office stands. Any answer of "we are not sure" marks the place to start.
- How many AI tools touch portfolio data today, including those embedded in platforms we already license?
- Which decisions is each meant to improve, and what was the baseline before it went live?
- If two tools disagree about a programme, whose number goes to the committee?
- Has any forecast we rely on been back-tested on our own closed programmes, and what was its error?
- Who is accountable for each AI use, and what may it do without a person deciding?
Sources and further reading
Statements about government frameworks were checked against these sources on 1 October 2026. Frameworks are revised, so read the current guidance for your jurisdiction.
- National AI Centre, Guidance for AI adoption: implementation guidance
- Digital Transformation Agency, Policy for the responsible use of AI in government, version 2.0
- Digital Transformation Agency, AI use case impact assessment (policy version 2.0)
- APRA, Letter to industry on artificial intelligence, 30 April 2026
- Standards Australia, adoption of AS ISO/IEC 42001:2023, AI management system
- NIST, AI Risk Management Framework
- NIST, AI RMF 1.0, section 5, the Core: govern, map, measure and manage