A regional cash service intelligent decision system and method

By combining edge computing and multi-agent reinforcement learning, the problem of optimizing the disconnect between the vault and external cash logistics was solved, achieving global optimization and real-time risk control, and improving the security and adaptability of the cash service system.

CN122434409APending Publication Date: 2026-07-21南通融鑫信息技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
南通融鑫信息技术有限公司
Filing Date
2026-04-03
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as localized bottlenecks, poor adaptability of decision-making models, lagging operational risk control, and a lack of continuous self-optimization capabilities due to the disconnect between internal vault operations and external cash logistics.

Method used

Real-time video analysis is performed using edge computing nodes. By combining causal inference and multi-agent reinforcement learning, a causal graph model is constructed to estimate the intervention effect, generate decision actions, and achieve global optimization and real-time risk control through multi-agent collaborative decision-making and autonomous evolution layers.

Benefits of technology

It achieves end-to-end global optimal allocation of cash resources, improves security and compliance, enhances the traceability and robustness of decision-making, and has continuous adaptive capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434409A_ABST
    Figure CN122434409A_ABST
Patent Text Reader

Abstract

The present application relates to the field of cash logistics and vault operation intelligence, and particularly relates to a regional cash service intelligent decision system and method, wherein the system comprises: an edge risk perception and prevention layer, which is used for identifying abnormal behaviors and generating risk event data; a data perception and fusion layer, which is used for collecting and fusing multi-source data, and constructing system state time sequence features; a causal inference layer, which is used for constructing a causal graph model and performing intervention effect estimation; a multi-agent reinforcement learning decision layer, which is used for generating decision actions according to the risk event data, the system state time sequence features and the intervention effect estimation results; a business rule constraint and explanation module, which is used for performing compliance verification on the decision actions and generating decision explanation information; and an autonomous evolution layer, which is used for generating cross-supervision signals according to mutual evaluation between the agents, and driving evolution update of the agent strategies. The present application realizes global collaboration, real-time risk control, interpretable decision and autonomous evolution of regional cash service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cash logistics and vault operation, specifically to an intelligent decision-making system and method for regional cash services. Background Technology

[0002] Current cash service management, especially the coordinated processes involving vault sorting, ATM cash replenishment, and external security transport, primarily relies on rule-based systems or independent optimization models. These existing technologies have significant drawbacks: First, they typically treat internal vault operations and external cash logistics as separate systems for independent optimization, lacking global coordination and leading to overall bottlenecks caused by localized optimization. Second, even with the introduction of AI technologies such as reinforcement learning, they generally face convergence difficulties and strategy fragility due to the instability of the financial environment and sparse decision feedback; furthermore, the model's decision-making process is like a "black box," lacking interpretability. Third, in terms of operational risk control, they heavily rely on post-event audits and manual video monitoring, lacking real-time, proactive, and intelligent risk control interception capabilities, and are unable to identify and intervene in violations (such as single-person operation, process deviations, and area intrusion) in key areas like vaults and branches within milliseconds. Finally, existing systems are rigid; once deployed, their strategies or rules are difficult to automatically adjust and evolve according to changes in business models, relying on expensive and slow human expert intervention.

[0003] Although existing research has explored the combination of causal inference and reinforcement learning, or attempted to apply multi-agent systems in financial decision-making, no specific technical solution has been recorded in the prior art for deeply integrating causal inference, multi-agent reinforcement learning (MARL), real-time operational risk control based on edge computing, and autonomous evolution mechanism, and applying it to dynamic decision-making throughout the entire process of regional cash services.

[0004] Therefore, there is an urgent need for a technical solution that can achieve global collaborative optimization, real-time risk control and blocking, explainable decision-making and continuous self-evolution, so as to realize end-to-end, explainable, secure adaptive and continuously evolving intelligent decision-making from vault to branch. Summary of the Invention

[0005] To address the technical problems in existing technologies, such as insufficient coordination between internal vault operations and external allocation, poor adaptability of decision-making models, lagging operational risk control, and lack of continuous self-optimization capabilities, this invention provides a regional cash service intelligent decision-making system and method.

[0006] In a first aspect, the present invention provides a regional cash service intelligent decision-making system, comprising:

[0007] The edge risk perception and prevention layer includes several edge computing nodes deployed in key areas of the vault and branches, which are used to identify abnormal operational behaviors and generate risk event data by analyzing local video. The data perception and fusion layer, connected to the edge risk perception and prevention layer, is used to collect and fuse multi-source data to construct system state temporal characteristics. The causal inference layer, connected to the data perception and fusion layer, is used to construct a causal graph model and estimate the intervention effect. The multi-agent reinforcement learning decision layer includes multiple agents for generating decision actions based on the risk event data, the temporal characteristics of the system state, and the results of the intervention effect estimation. The business rule constraint and interpretation module is connected to the multi-agent reinforcement learning decision layer and is used to perform compliance verification on the decision actions and generate decision interpretation information. The autonomous evolution layer, connected to the multi-agent reinforcement learning decision layer, is used to drive the evolution and update of each agent's strategy based on the cross-supervision signals generated by mutual evaluation among the agents.

[0008] Furthermore, the edge computing node includes a video acquisition unit, a lightweight AI processing unit, a video analysis unit, a central business rules unit, a local rules engine, and a data interface. The video analysis unit uses a combination of traditional computer vision algorithms and lightweight convolutional neural networks to process the video stream in real time and identify targets, including personnel identity and compliant actions, work tool status, compliance of cash box handling paths, and unauthorized area intrusion. The local rules engine is used to execute localized alarms and device control operations when abnormal behavior is detected. The data interface is used to report structured risk event data through a collaborative data bus.

[0009] Furthermore, the data perception and fusion layer is used to collect and fuse multi-source data, including risk event data, vault internal operation data, external branch inventory data, and cash logistics data.

[0010] Furthermore, the causal inference layer includes: The causal graph construction unit, based on domain knowledge and historical data learning algorithms, constructs a causal structure graph covering macroeconomic indicators, regional consumption index, vault operation efficiency, branch cash gap and operational risk variables; The causal inference engine uses the causal structure graph to estimate the intervention effect and perform counterfactual queries, and outputs the estimated intervention effect value to the multi-agent reinforcement learning decision layer.

[0011] Furthermore, the multi-agent reinforcement learning decision layer includes multiple agents, including an internal vault agent and an external allocation agent. The internal vault agent includes a sorting machine agent, an ATM cash replenishment team agent, and a warehouse inventory management agent. The external allocation agent includes a branch agent, an armored vehicle agent, and a regional allocation center agent.

[0012] Furthermore, the multi-agent reinforcement learning decision layer uses a reward function for decision-making, the reward function... Represented as:

[0013] in, Incentives for basic business operations As a cause-and-effect reward, To constrain rewards through rules, For collaborative rewards, As a reward for risk prevention and control, , , , These are the weighting coefficients.

[0014] Furthermore, the business rule constraint and interpretation module includes: A business rules knowledge base is used to store security rules, compliance rules, and operational rules expressed in logical form. The rule distribution unit is used to compile and distribute core risk control rules to the local rule engine of each edge computing node; The decision verification unit is used to perform real-time logical verification on each decision action output by the multi-agent reinforcement learning decision layer. If the hard constraints are violated, a negative reward is generated and a policy adjustment is triggered. The explanation generation unit is used to automatically generate natural language explanations for decision-making actions using the causal graph model and risk event data.

[0015] Furthermore, the self-evolutionary layer includes: The supervisory network unit is used to observe the historical behavioral trajectory of the agent and generate cross-supervision signals based on public performance indicators and a shared rule base; The fitness evaluation unit is used to calculate the fitness of the current agent's decision parameters based on the cross-supervision signals and risk prevention rewards received by each agent. An evolutionary operation unit is used to perform crossover and mutation operations on the decision parameters of the agent according to the fitness and generate candidate decision parameters. Based on short-term simulation verification, candidate decision parameters that are better than the current decision parameters are updated to the multi-agent reinforcement learning decision layer.

[0016] Secondly, the present invention provides a regional cash service intelligent decision-making method, comprising the following steps: Real-time video analysis is performed using edge computing nodes deployed in key areas of vaults and branches to identify abnormal operational behaviors and generate risk event data. Collect and integrate the risk event data, vault internal operation data, external branch inventory data, and cash logistics data to construct the system status time series characteristics; Construct a causal graph model and estimate the intervention effect based on the causal graph model; Multi-agent reinforcement learning is used for collaborative decision-making, with each agent generating a decision action based on the risk event data, system state temporal characteristics, and the results of intervention effect estimation. The decision-making actions are verified for compliance, and decision explanation information is generated. Cross-supervision signals are generated through mutual evaluation among agents, and the evolution and updating of decision parameters of each agent are driven by risk event feedback.

[0017] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention fundamentally solves the problem of local optimization and overall bottleneck caused by the separation of internal vault operations and external cash logistics in the prior art, and realizes the global optimal allocation of cash resources from end to end.

[0018] (2) This invention changes the operational risk control from post-event tracing to in-event blocking, which significantly improves the security and compliance of the cash service chain.

[0019] (3) The present invention significantly enhances the traceability, understandability and credibility of the decision-making process, and the introduction of causal rewards and risk rewards makes strategy learning more stable and has stronger robustness to the non-stationarity of the financial environment.

[0020] (4) The present invention enables the system to continuously and autonomously optimize by utilizing interactive experience and risk feedback during operation, and has the continuous adaptive capability that traditional systems lack. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a general framework diagram of a regional cash service intelligent decision-making system according to the present invention.

[0023] Figure 2 This is a diagram of the edge computing node framework of the present invention.

[0024] Figure 3 This is a schematic diagram of the causal structure of the regional cash service of the present invention.

[0025] Figure 4 This is a flowchart of the multi-agent reinforcement learning collaborative decision-making process of the present invention.

[0026] Figure 5 This is a closed-loop flowchart of the cross-supervised evolution mechanism of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] This invention provides the following technical solutions: like Figure 1 The regional cash service intelligent decision-making system shown includes: The edge risk perception and prevention layer includes several edge computing nodes deployed in key areas of the vault and branches, which are used to identify abnormal operational behaviors and generate risk event data by analyzing local video. The data perception and fusion layer, connected to the edge risk perception and prevention layer, is used to collect and fuse multi-source data to construct system state temporal characteristics. The causal inference layer, connected to the data perception and fusion layer, is used to construct a causal graph model and estimate the intervention effect; The multi-agent reinforcement learning decision layer includes multiple agents for generating decision actions based on the risk event data, the temporal characteristics of the system state, and the results of the intervention effect estimation. The business rule constraint and interpretation module is connected to the multi-agent reinforcement learning decision layer and is used to perform compliance verification on the decision actions and generate decision interpretation information. The autonomous evolution layer, connected to the multi-agent reinforcement learning decision layer, is used to drive the evolution and update of each agent's strategy based on the cross-supervision signals generated by mutual evaluation among the agents.

[0030] As a preferred embodiment of this application, such as Figure 2 The diagram shows an edge computing node, which includes a video acquisition unit, a lightweight AI processing unit, a video analysis unit, a central business rules unit, a local rules engine, and a data interface. The video analysis unit uses a combination of traditional computer vision algorithms and lightweight convolutional neural networks to process video streams in real time and identify targets, including personnel identity and compliant actions, tool status, compliance of cash box handling paths, and unauthorized area intrusions. The local rules engine is used to execute localized alarms and device control operations when abnormal behavior is detected. The data interface is used to report structured risk event data through a collaborative data bus.

[0031] In a preferred embodiment of this application, the data perception and fusion layer is used to collect and fuse multi-source data, including risk event data, vault internal operation data, external branch inventory data, and cash logistics data.

[0032] In a preferred embodiment of this application, the causal inference layer includes: The causal graph construction unit, based on domain knowledge and historical data learning algorithms, constructs a causal structure graph covering macroeconomic indicators, regional consumption indices, vault operation efficiency, branch cash shortages, and operational risk variables, such as... Figure 3 The diagram shown is a schematic representation of the cause-and-effect relationship of regional cash services. The causal inference engine uses the causal structure graph to estimate the intervention effect and perform counterfactual queries, and outputs the estimated intervention effect value to the multi-agent reinforcement learning decision layer.

[0033] In a preferred embodiment of this application, the multi-agent reinforcement learning decision layer includes multiple agents, including an internal vault agent and an external allocation agent. The internal vault agent includes a sorting machine agent, an ATM cash replenishment team agent, and a warehouse inventory management agent. The external allocation agent includes a branch agent, an armored vehicle agent, and a regional allocation center agent.

[0034] As a preferred embodiment of this application, such as Figure 4 The diagram shown illustrates the multi-agent reinforcement learning collaborative decision-making process of this invention. The multi-agent reinforcement learning decision-making layer employs a reward function for decision-making. Represented as:

[0035] in, Incentives for basic business operations As a cause-and-effect reward, To constrain rewards through rules, For collaborative rewards, As a reward for risk prevention and control, , , , These are the weighting coefficients.

[0036] In a preferred embodiment of this application, the business rule constraint and interpretation module includes: A business rules knowledge base is used to store security rules, compliance rules, and operational rules expressed in logical form. The rule distribution unit is used to compile and distribute core risk control rules to the local rule engine of each edge computing node; The decision verification unit is used to perform real-time logical verification on each decision action output by the multi-agent reinforcement learning decision layer. If the hard constraints are violated, a negative reward is generated and a policy adjustment is triggered. The explanation generation unit is used to automatically generate natural language explanations for decision-making actions using the causal graph model and risk event data.

[0037] As a preferred embodiment of this application, such as Figure 5 The diagram shown is a closed-loop flowchart of the cross-supervised evolution mechanism of this invention. The autonomous evolution layer includes: The supervisory network unit is used to observe the historical behavioral trajectory of the agent and generate cross-supervision signals based on public performance indicators and a shared rule base; The fitness evaluation unit is used to calculate the fitness of the current agent's decision parameters based on the cross-supervision signals and risk prevention rewards received by each agent. An evolutionary operation unit is used to perform crossover and mutation operations on the decision parameters of the agent according to the fitness and generate candidate decision parameters. Based on short-term simulation verification, candidate decision parameters that are better than the current decision parameters are updated to the multi-agent reinforcement learning decision layer.

[0038] A smart decision-making method for regional cash services includes the following steps: S1. Real-time video analysis is performed based on edge computing nodes deployed in key areas of vaults and branches to identify abnormal operational behaviors and generate risk event data.

[0039] S2. Collect and integrate the risk event data, vault internal operation data, external branch inventory data, and cash logistics data to construct system status time-series characteristics.

[0040] S3. Construct a causal graph model and estimate the intervention effect based on the causal graph model.

[0041] S4. Multi-agent reinforcement learning is used for collaborative decision-making. Each agent generates a decision action based on the risk event data, the system state time sequence characteristics, and the results of intervention effect estimation.

[0042] S5. Perform compliance verification on the decision-making action and generate decision explanation information.

[0043] S6. Cross-supervision signals are generated through mutual evaluation among agents, and the evolution and updating of decision parameters of each agent are driven by risk event feedback.

[0044] Example Taking a regional cash service network as an example, the intelligent decision-making system and method for regional cash services provided by this invention will be described in detail. The regional cash service network includes: 1 regional vault (equipped with 3 cash sorting machines, 2 ATM cash replenishment teams, and 1 warehouse), 50 business outlets, and 10 armored vehicles.

[0045] In the edge computing node of the vault clearing section, the video acquisition unit collects video streams of the work locations in real time. The video analysis unit uses a background subtraction algorithm to detect moving targets and identifies target categories through a lightweight convolutional neural network (MobileNet-SSD). At a certain moment, the system detects that only one operator at the current workstation has been working continuously for more than the prescribed time, and no second operator is detected. The local rule engine matches the clearing operation rule synchronized from the business rule constraint and interpretation module, which requires two people to be present, and immediately executes a localized decision: issuing an audible and visual alarm and locking the next task start button of the clearing machine. At the same time, the data interface reports the violation of single-person operation and clearing section 3 as high-priority risk events to the collaborative data bus.

[0046] The data perception and fusion layer acquires real-time operational data from within the vault via an IoT interface: the real-time status (idle, busy, faulty) of the three sorting machines, the length of each work queue, and the inventory of each cash box in the vault. It also obtains point-in-time inventory and scheduled cash withdrawal data from 50 branches through the bank's core system. Furthermore, it acquires the GPS locations and task lists of 10 armored vehicles through a scheduling platform. Simultaneously, it receives structured risk event data reported by the edge risk perception and control layer. After standardization and time-series alignment, this multi-source data is organized into system state time-series features and input into the causal inference layer and the multi-agent reinforcement learning decision layer.

[0047] The multi-agent reinforcement learning decision layer includes agents for the cash sorting machine, ATM cash replenishment team, warehouse inventory management, branch operations, security vehicle operations, and regional allocation center. Each agent employs a centralized training and distributed execution architecture (MADDPG algorithm).

[0048] At a certain moment, the warehouse inventory management agent receives the following state inputs: local observations (current warehouse inventory level, demand forecasts for each branch), the intervention effect value output by the causal inference layer, and risk events reported by edge nodes. The agent's policy network outputs the action: suspend routine outbound tasks and prioritize personnel to check access control anomalies. This output action receives a positive reward for timely risk response, and the composite reward guides the agent to make a balanced decision between efficiency and safety.

[0049] The business rule constraint and interpretation module performs real-time logical verification of the above decision actions. The decision verification unit checks whether the action violates hard constraints, and if it is verified to be compliant, it is deemed compliant. The interpretation generation unit uses a causal graph model and risk event data to automatically generate a natural language decision interpretation: Because the edge node detected an anomaly in the access control of warehouse X, and causal inference indicates a negative correlation between the violation and operational efficiency, the decision is made to suspend routine outbound tasks and prioritize personnel verification to prevent potential security risks.

[0050] Taking the escort vehicle agent A and the vault dispatch agent B as examples: Escort vehicle agent A designed an efficient route that required rapid vault exit, leading to frequent reports of congestion and unauthorized personnel gathering at the vault edge nodes. Agent B observes agent A's behavioral trajectory through the supervisory network unit and sends negative supervisory signals to agent A based on public performance indicators and a shared rule base. The fitness evaluation unit integrates the supervisory signals and environmental rewards to calculate the fitness of agent A's decision parameters. The evolutionary operation unit performs crossover and mutation operations based on the fitness, generating candidate decision parameters, which are then updated selectively after short-term simulation verification.

[0051] This embodiment fully demonstrates the closed-loop operation process of the present invention, from edge risk perception, multi-source data fusion, causal inference guidance, multi-agent collaborative decision-making, rule verification and interpretation to autonomous evolution, and verifies the technical effectiveness of the system in achieving global collaborative optimization, real-time risk prevention and control, interpretable decision-making, and continuous adaptive capabilities.

[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A regional cash service intelligent decision-making system, characterized in that, include: The edge risk perception and prevention layer includes several edge computing nodes deployed in key areas of the vault and branches, which are used to identify abnormal operational behaviors and generate risk event data by analyzing local video. The data perception and fusion layer, connected to the edge risk perception and prevention layer, is used to collect and fuse multi-source data to construct system state temporal characteristics. The causal inference layer, connected to the data perception and fusion layer, is used to construct a causal graph model and estimate the intervention effect. The multi-agent reinforcement learning decision layer includes multiple agents for generating decision actions based on the risk event data, the temporal characteristics of the system state, and the results of the intervention effect estimation. The business rule constraint and interpretation module is connected to the multi-agent reinforcement learning decision layer and is used to perform compliance verification on the decision actions and generate decision interpretation information. The autonomous evolution layer, connected to the multi-agent reinforcement learning decision layer, is used to drive the evolution and update of each agent's strategy based on the cross-supervision signals generated by mutual evaluation among the agents.

2. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The edge computing node includes a video acquisition unit, a lightweight AI processing unit, a video analysis unit, a central business rules unit, a local rules engine, and a data interface. The video analysis unit uses a combination of traditional computer vision algorithms and lightweight convolutional neural networks to process the video stream in real time and identify targets, including personnel identity and compliant actions, status of work tools, compliance of cash box handling paths, and unauthorized area intrusion. The local rule engine is used to execute localized alarms and device control operations when abnormal behavior is detected; the data interface is used to report structured risk event data through the collaborative data bus.

3. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The data perception and fusion layer is used to collect and fuse multi-source data, including risk event data, vault internal operation data, external branch inventory data, and cash logistics data.

4. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The causal inference layer includes: The causal graph construction unit, based on domain knowledge and historical data learning algorithms, constructs a causal structure graph covering macroeconomic indicators, regional consumption index, vault operation efficiency, branch cash gap and operational risk variables; The causal inference engine uses the causal structure graph to estimate the intervention effect and perform counterfactual queries, and outputs the estimated intervention effect value to the multi-agent reinforcement learning decision layer.

5. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The multi-agent reinforcement learning decision layer includes multiple agents, including an internal vault agent and an external allocation agent. The internal vault agent includes a sorting machine agent, an ATM cash replenishment team agent, and a warehouse inventory management agent. The external allocation agent includes a branch agent, an escort vehicle agent, and a regional allocation center agent.

6. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The multi-agent reinforcement learning decision layer uses a reward function for decision-making. Represented as: in, Incentives for basic business operations As a cause-and-effect reward, To constrain rewards through rules, For collaborative rewards, As a reward for risk prevention and control, , , , These are the weighting coefficients.

7. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The business rule constraint and interpretation module includes: A business rules knowledge base is used to store security rules, compliance rules, and operational rules expressed in logical form. The rule distribution unit is used to compile and distribute core risk control rules to the local rule engine of each edge computing node; The decision verification unit is used to perform real-time logical verification on each decision action output by the multi-agent reinforcement learning decision layer. If the hard constraints are violated, a negative reward is generated and a policy adjustment is triggered. The explanation generation unit is used to automatically generate natural language explanations for decision-making actions using the causal graph model and risk event data.

8. The regional cash service intelligent decision-making system according to claim 1, characterized in that, The self-evolutionary layer includes: The supervisory network unit is used to observe the historical behavioral trajectory of the agent and generate cross-supervision signals based on public performance indicators and a shared rule base; The fitness evaluation unit is used to calculate the fitness of the current agent's decision parameters based on the cross-supervision signals and risk prevention rewards received by each agent. An evolutionary operation unit is used to perform crossover and mutation operations on the decision parameters of the agent according to the fitness and generate candidate decision parameters. Based on short-term simulation verification, candidate decision parameters that are better than the current decision parameters are updated to the multi-agent reinforcement learning decision layer.

9. A regional cash service intelligent decision-making method, implemented based on the system described in any one of claims 1-8, characterized in that, Includes the following steps: Real-time video analysis is performed using edge computing nodes deployed in key areas of vaults and branches to identify abnormal operational behaviors and generate risk event data. Collect and integrate the risk event data, vault internal operation data, external branch inventory data, and cash logistics data to construct the system status time series characteristics; Construct a causal graph model and estimate the intervention effect based on the causal graph model; Multi-agent reinforcement learning is used for collaborative decision-making, with each agent generating a decision action based on the risk event data, system state temporal characteristics, and the results of intervention effect estimation. The decision-making actions are verified for compliance, and decision explanation information is generated. Cross-supervision signals are generated through mutual evaluation among agents, and the evolution and updating of decision parameters of each agent are driven by risk event feedback.