Abstract
Climate forecasts describe possible physical outcomes, whereas adaptation finance requires a defensible choice among interventions, beneficiaries and budget constraints. This paper sets out a developed agentic artificial intelligence governance framework for Caribbean climate decisions. Specialised agents collect dated scientific evidence, examine hazards and vulnerability, assess adaptation options, check the conditions attached to finance, and identify conflicts in the analysis. A separately tasked reviewer challenges weak sources, disparities among communities and recommendations outside policy authority. A recorded approval process reserves final decisions for named human officials. The system is represented as an evidence-and-decision graph rather than as an unconstrained collection of chatbots. A Jamaican flood-adaptation funding case illustrates its intended workflow. The proposed empirical assessment compares multi-agent output with a conventional analyst group and a single-model assistant, using source traceability, contradiction detection, time to complete the review, reproducibility and equity of reasoning. The framework is designed, but it has not been deployed in a public agency and has no independent effectiveness results. Its practical claim is narrower than algorithmic policymaking: automation may help organise a difficult review provided that each material assertion and proposed expenditure remains open to human examination.
1. The decision that a forecast does not make
A precipitation projection may show that one catchment faces a serious flooding hazard. It does not specify which intervention a government should finance, how landowners will be affected, whether an allocation satisfies the conditions of a donor or which evidence is strong enough for a spending decision. Those additional questions depend on vulnerability, exposure, statutory authority, procurement, costs, equity and the institution's tolerance of uncertainty. Climate modelling and public decision-making are therefore linked, but they do not solve the same problem.
The framework presented here is designed for Caribbean public agencies and regional institutions that must compare adaptation measures with incomplete data and limited analyst capacity. Its organising principle is to keep the physical forecast, the policy interpretation and the funding recommendation as distinct objects. An agent may retrieve an assessment or test a calculation. It may not invent missing evidence, silently change a policy threshold or treat agreement between AI agents as independent corroboration.
The completed contribution is a governance architecture and decision protocol. It should not be confused with an installed product at any Caribbean ministry. The paper specifies a reproducible test using historical scenarios, but no government deployment, cost saving or increase in allocation quality is yet measured.
2. Institutional and technical basis
Disaster risk can be examined in terms of hazard, exposure, vulnerability and capacity, concepts defined in the Sendai Framework terminology maintained by the UN Office for Disaster Risk Reduction [1]. A policy decision adds institutional powers, eligibility criteria and distributive consequences to that physical assessment. The definition of the inputs therefore needs to be explicit before a conversational model is asked to recommend an intervention.
The US National Institute of Standards and Technology's AI Risk Management Framework describes governance, mapping, measurement and management of AI risk, and emphasises trustworthiness throughout the system lifecycle [2]. The UNESCO Recommendation on the Ethics of Artificial Intelligence supplies a broader normative reference for human rights, accountability and human oversight [3]. Neither document constitutes proof that a particular chatbot or multi-agent workflow is safe. They supply criteria against which its design and operation can be audited.
The model uses agents because different reviews require different source collections and prompts, not because more agents automatically produce greater accuracy. Common language-model weights or shared retrieval corpora may create correlated errors. A challenge agent only provides independent scrutiny if the protocol imposes different evidentiary tests and preserves dissent in the record.
3. Architecture and decision rights
3.1 The evidence ledger
Every material assertion enters an append-only ledger with a source URL or document identifier, publisher, observation date, retrieval date, geographic extent, unit, confidence qualifier and the step that consumed the assertion. A claim must identify whether it comes from a measured observation, reanalysis, scenario, expert judgment, administrative rule or model-generated inference. The ledger lets a reviewer distinguish a flood map from a legal eligibility requirement and a projected loss from a recorded loss.
Five functional roles then work against that ledger: an evidence collector; a hazard-and-exposure analyst; a vulnerability and distribution reviewer; a finance-and-policy reviewer; and a challenge reviewer. These are software roles, not public offices. Each produces an output with citations and assumptions. The final decision record belongs to a human official and includes a signed rationale, alternatives considered, objections unresolved and the version of each underlying source.
3.2 A decision graph rather than an autonomous consensus
The structure can be represented as a directed graph G = (V, E), in which V includes source claims, proposed interventions, required approvals and contradictory observations. Edges indicate support, dependence, contradiction or referral. A recommendation is valid for human review only if its material dependencies resolve to accepted sources or explicitly disclosed assumptions.
Some conditions must be represented as hard gates rather than compensated by a high score. For example, absent procurement authority, a prohibited expenditure or an unresolved conflict of interest cannot be cancelled out by a large predicted physical benefit. The system should return “insufficient evidence” when relevant information is missing.
4. Interactive evidence-flow figure
5. Worked policy scenario
Consider a municipality evaluating three flood-adaptation measures: relocating repeatedly exposed households, improving drainage in an urban catchment, and protecting an essential road. The physical evidence agent retrieves a dated hazard forecast and historical event record. The exposure agent determines which roads, homes and services lie in the mapped hazard zone, identifying the data's resolution and any unmapped settlements. The policy agent retrieves the relevant planning, land and funding provisions.
The challenge reviewer then asks whether a high predicted reduction in damage is sensitive to an unverified asset inventory, whether relocation has consent and compensation requirements, and whether the least visible informal settlements have been omitted from the spatial data. The output is a comparative memorandum containing competing priorities, rather than a single apparently definitive ranking.
The designated official can authorise further study, reject the options or select one with recorded reasons. Before release of funds, the workflow requires a review of statutory approvals, safeguards and budget eligibility. A later independent reviewer must be able to reproduce the evidence trail without access to a hidden chat history.
A key failure case is conflicting source age. An agent might retrieve a recent funding rule and an outdated planning document that both appear authoritative. The protocol treats inconsistent publication dates as a contestable finding rather than resolving the contradiction by fluency or majority vote.
6. Proposed evaluation and falsification
A retrospective evaluation will provide the same bounded evidence packet to three independent workflows: a conventional group of human analysts, one AI assistant and the governed multi-agent system. The team will specify questions and rating rubrics before reviewing responses. Analysts will be blinded to origin where practical; independent reviewers will verify factual statements against original source documents.
| Measure | Operational definition | Undesirable outcome |
|---|---|---|
| Source accuracy | Share of material claims supported by exact retrieved evidence | Unsupported figures or stale citations |
| Contradiction recall | Known source conflicts correctly surfaced | Conflicts silently harmonised |
| Review time | Analyst minutes to reach a decision-ready record | Faster output at the cost of errors |
| Equity audit | Coverage of specified communities and vulnerability factors | Unjustified omission of marginalised groups |
| Decision traceability | Ability to reconstruct claims, revisions and approvals | No reliable source or reviewer trail |
Table 1. Pre-specified evaluation measures. No results have yet been observed.
Tests should include fabricated-source traps, conflicting policy versions, interrupted tool calls, prompt injection inside retrieved documents and a case where every proposed action is ineligible. These tests evaluate not only the accuracy of agents but also whether the framework can fail safely. Statistical comparisons should report uncertainty intervals and reviewer agreement, rather than treating one favourable example as a general result.
7. Discussion and governance limits
The governance problem is not solved merely by adding a human approval button. Officials require time and information to challenge outputs; otherwise, nominal human oversight may become automatic endorsement. The review record therefore needs accessible citations, visible disagreement, a mechanism to request missing evidence and procedures to stop a transaction when source confidence is inadequate.
Some evidence cannot be made public, particularly data about vulnerable households or critical infrastructure. Access rights, data minimisation and retention controls must be established by the deploying institution. The workflow must also respect jurisdiction-specific law; a general NIST or UNESCO framework does not replace local statutory duties. Agent-generated recommendation text should never be classified as an authoritative policy interpretation without institutional review.
There is no validated demonstration that the framework improves climate-finance allocation. A credible finding may be that it reduces document-search time but does not improve substantive choices, or that gains depend heavily on clean administrative data. Those are legitimate research outcomes.
8. Conclusion
This research establishes a human-governed decision architecture in which climate hazards, policy authority and finance eligibility can be examined together without being merged into an opaque automated recommendation. The next scientific step is an independently reviewed comparison against ordinary analysts and a single assistant. The benchmark must establish whether provenance, contradiction handling and equity improve while keeping the authority to spend public money with accountable people.
References
- United Nations Office for Disaster Risk Reduction (2017). Sendai Framework Terminology: Disaster Risk. https://www.undrr.org/terminology/disaster-risk.
- National Institute of Standards and Technology (2023). AI Risk Management Framework 1.0. https://www.nist.gov/itl/ai-risk-management-framework.
- UNESCO (2021). Recommendation on the Ethics of Artificial Intelligence. https://www.unesco.org/en/artificial-intelligence/recommendation-ethics.