Over the past several months, we have worked with a growing number of leading financial institutions on AI governance. The conversations tend to start in the same place. A security or risk leader explains that the institution is still defining its long-term AI strategy, while employees are already using AI every day. They’re summarizing filings, reviewing customer documents, and adopting tools outside established procurement processes. Lloyds reports that 59% of institutions were already reporting productivity improvements from AI. The efficiency is real, but so is the exposure.
Financial institutions handle client data, nonpublic personal information, investment models, and regulated decisions, and a written policy alone cannot protect them. Security teams often lack visibility into which tools are in use, what is being shared, and what actions AI systems are taking. What institutions need is a clear path that begins with real usage and becomes more rigorous as AI moves closer to customers, regulated decisions, and financial actions.
Stage 1. Gain Visibility Into How AI Is Actually Used
The first step is understanding the AI activity already taking place. That includes public chatbots, coding assistants, meeting tools, and autonomous agents. It also includes employees using personal accounts for tools that the organization licenses centrally. And increasingly, it includes AI features that vendors add to applications the institution already owns: capabilities that arrive through routine product updates, inherit the application's existing access to data, and never pass through a procurement or security review of their own.
Approved vendor lists reveal only a fraction of actual AI use. Shadow AI often appears because employees are under pressure to move faster, not because they intend to bypass security. An analyst who needs to summarize a dense report or an operations employee trying to automate a repetitive task may choose the fastest available tool. In other cases, no one chooses anything: a document platform or CRM simply begins offering AI summarization or drafting, and sensitive data starts flowing to a model without a single deliberate decision inside the institution.
Visibility also needs to extend below the application layer. A single AI tool may depend on a cloud provider, a foundation-model provider, external datasets, connectors, and several subprocessors, and a vendor can change any of these without changing the interface employees see. Discovering an application is not the same as understanding where its data goes.
Visibility gives the institution evidence to shape its policy. The goal is a deep understanding of how people actually work with AI: what information employees share, through which accounts, and what actions AI systems take on their behalf. Security leaders should identify where sensitive data enters prompts and attachments, where personal accounts stand in for enterprise ones, and where agents are already accessing internal systems. This is the evidence that classification, data policies, and runtime controls in later stages will depend on.
Stage 2. Classify the Use Case by Potential Impact
Not every AI interaction needs the same governance process. Summarizing a public filing does not carry the same risk as recommending whether to approve a loan or preparing a trade.
A practical classification model should consider data sensitivity, customer impact, financial exposure, autonomy, reversibility, and the importance of the affected process.
Where a given use case lands will differ by institution. The same activity may be routine at one firm and material at another, depending on its business model, customer base, regulatory obligations, and risk appetite, so the tiers below are illustrative rather than universal.
Low-impact use cases rely on public or nonsensitive information for internal productivity. Examples include summarizing public research, drafting internal notes, or generating boilerplate code.
Moderate-impact use cases involve internal data or support operational decisions. This may include internal investment research, customer-service drafting, fraud investigation support, or document review.
High-impact use cases influence regulated decisions, customer outcomes, or financial activity. Credit decisions, investment recommendations, trading, payments, account changes, and regulatory filings belong in this category.
This classification should determine the depth of review, testing, approval, logging, and human oversight required. It also prevents governance from becoming a single slow process applied equally to every experiment.
Stage 3. Define the AI Policy for Data and Actions
Once the use case is understood, the next questions are straightforward. What information does the AI need, what should it never receive, and what is it permitted to do with the systems it can reach?
On the data side, controls cannot stop at structured fields or documents with known labels. Sensitive context often appears in prompts, copied text, screenshots, attachments, and conversational threads. Despite not matching any DLP pattern, they can still create significant regulatory risk if shared with a public model.
On the action side, the policy should state which systems an AI may connect to, which operations it may perform there, and which remain reserved for humans. For example, an assistant permitted to read a CRM is not permitted to update it. An agent that can draft communications is not permitted to send them. Defining these permissions in policy gives later stages something concrete to assign and enforce.
Policies should account for the content, context, and purpose of the interaction, and distinguish between enterprise and personal accounts, since the same application may handle data differently depending on how it is accessed. Firms can then enforce the policy with AI-specific controls that allow, restrict, redact, or block information before it reaches an AI system, and that constrain the actions an AI attempts in connected applications.
Stage 4. Define the Role AI May Play in Decisions
Financial institutions should define the authority of an AI system before focusing on the sophistication of the model.
A useful framework separates four roles.
Inform. The system retrieves, organizes, or summarizes information.
Recommend. The system proposes a decision, while a human remains responsible for evaluating and approving it.
Prepare. The system drafts an action, communication, filing, payment, or transaction for human review.
Execute. The system takes action in another system.
Each step increases the governance burden. An assistant may be permitted to summarize a credit file without recommending an outcome. A validated system may recommend an outcome, while a qualified employee remains accountable for the decision. An agent may prepare a payment, but submission may require a separate approver.
This distinction is especially important as AI shifts from generating content to operating software. Access to an account system, trading platform, payment workflow, or customer communication channel creates a different risk than access to a chat interface.
Stage 5. Constrain Actions at Runtime and Log Everything
Policies and predeployment reviews cannot anticipate every prompt, output, or agent decision. Controls are needed at the moment an AI system attempts to act.
The right control depends on the activity. Institutions may apply transaction thresholds, role-based permissions, approved-tool restrictions, dual approval, separation of duties, customer-level restrictions, time-based limits, or requirements that high-risk actions remain reversible.
An agent preparing a payment may be allowed to populate the workflow but blocked from submitting it. An investment assistant may retrieve approved research but be prevented from accessing restricted deal information. A customer-service agent may draft a response but require review before it reaches the customer.
Every access and action should leave a record. The audit trail should capture what data an AI system touched, what it did, under which account and authority, and whether a human reviewed or approved the result. In a regulated institution this is not optional telemetry. It is what allows the firm to answer an examiner, reconstruct an incident, or demonstrate that the accountable human, not the model, made the decision.
Stage 6. Continuously Monitor the System After Launch
Approval is not the end of governance. Monitoring should watch two things: the systems themselves, and the behavior recorded in the audit trail.
AI systems change constantly. Model updates, new data, altered prompts, and new integrations can shift behavior without any visible change to the application employees see, and a vendor can swap an underlying model or subprocessor without notice. Institutions should track changes in outputs, accuracy, override rates, decision disparities, and customer complaints, and connect findings back to the original risk classification so that a low-impact tool cannot quietly evolve into a high-impact workflow without another review. The same discipline applies to the discovery work from Stage 1, which has to run continuously rather than as an annual assessment, since new tools and features will keep arriving.
The audit trail established at runtime becomes a primary monitoring instrument. Reviewing it for anomalies, such as an agent accessing data outside its normal scope, actions taken without a corresponding approval, unusual volumes or timing, or activity under unexpected accounts, is how to identify control failures before they become incidents.
What This Looks Like in Practice
Consider three common scenarios.
An asset-management analyst uses an approved AI assistant to summarize public earnings reports. The use case is low impact. The firm requires a corporate account, redacts sensitive information before transmission to the LLM, records the activity, and keeps the analyst responsible for reviewing the summary.
A bank uses AI to support lending decisions. The institution validates the data and model, tests for disparate outcomes, requires explainable recommendations, documents human accountability, and monitors overrides and customer outcomes. The system may support the decision, but it does not silently replace the accountable decision-maker.
An AI agent prepares payments or account updates. The institution limits the agent’s permissions, applies amount thresholds, separates preparation from approval, records every action, and maintains an emergency shutdown mechanism. The agent can accelerate the workflow without receiving unlimited authority.
These scenarios belong to the same AI governance program, but they should not be governed identically. The controls become stronger as the system gains access to more sensitive data, greater influence over decisions, and more ability to act.
How Lumia Helps Financial Institutions Govern AI at Every Stage
Lumia gives security teams visibility into AI activity across applications, accounts, prompts, files, responses, and agent workflows.
Institutions can enforce policy in real time based on the content, context, and intent of each interaction. Enforcement actions can include redacting sensitive information before it is transmitted, restricting activity to approved accounts, or blocking an AI system from taking an unauthorized action.
This gives financial institutions a practical foundation for governing AI as it moves from employee experimentation into regulated processes and financial execution.
If you are working through these same challenges, see how Lumia can help: https://www.lumia.security/book-a-demo.

