Financial network security defense method and system based on multiple Agents and dynamic large model

By deploying detection Agent, decision Agent, intelligence Agent and dynamic game engine in the financial network system, the inefficiency problem of traditional defense methods in the face of fast attack methods is solved, efficient and intelligent threat identification and defense are achieved, and the system's defense capabilities are improved.

CN120498759APending Publication Date: 2025-08-15HUAYING (SHANGHAI) INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510625019.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When modern financial network systems face rapidly evolving attack methods, traditional defense methods are difficult to achieve efficient and intelligent threat identification and defense, resulting in low detection rates and high false alarm rates, and they are unable to adapt to new security threats in a timely manner.

Method used

Deploy detection Agent at the edge layer, collect traffic data in real time and generate abnormal scores; build decision Agents in the cloud to perform multimodal feature fusion and defense strategy generation; build cross-agency federated learning networks and dynamic game engines through intelligence Agents to optimize defense strategies; dynamically assign detection tasks to deal with different threat levels.

Benefits of technology

It realizes efficient collection and analysis of real-time traffic data of network nodes, improves the ability to identify abnormal behaviors, generates the optimal defense instruction set, improves the intelligence, automation and efficiency of network security defense, and enhances system resilience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498759A_ABST
    Figure CN120498759A_ABST
Patent Text Reader

Abstract

The invention discloses a financial network security defense method and system based on multiple Agents and a dynamic large model. A detection Agent is deployed in an edge layer, financial network node flow data and system logs are collected in real time, time sequence features are extracted through a lightweight convolutional network, and a preliminary anomaly score is generated. And the cloud layer constructs a decision Agent, receives the feature abstract transmitted by the edge node in an encrypted manner, inputs the feature abstract into a dynamic large model for multi-modal feature fusion, and outputs defense action probability distribution. And the intelligence Agent constructs a cross-institution federated learning network. And constructing a dynamic game engine, constructing a revenue matrix based on the attack cost and the defense revenue, solving a Nash equilibrium strategy, and generating an optimal defense instruction set. And dynamically allocating detection tasks according to the threat level and the edge computing power state. According to the method, efficient acquisition and analysis are realized, the abnormal behavior recognition capability is improved, support is provided for making a defense strategy, the defense strategy is optimized, and the intelligent, automatic and efficient levels of defense are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field to which the present invention relates is financial network security defense technology, specifically, a financial network security defense method and system based on multi-agent collaboration and dynamic large models. Background Art

[0002] In the modern financial network, multiple financial ecosystems are interconnected through advanced digital technologies, forming an efficient, convenient, and inclusive financial services network. Leveraging cutting-edge technologies such as blockchain and cloud computing, this network has significantly improved the efficiency and security of financial transactions. Blockchain technology, with its decentralized, transparent, and tamper-proof nature, provides a reliable record-keeping and verification mechanism for financial transactions, reducing transaction costs and mitigating the risk of fraud. Cloud computing, by offering elastic and scalable computing and storage resources, supports the processing and analysis of large-scale financial data, empowering financial institutions with powerful data processing capabilities and driving innovation and development in financial services.

[0003] Modern financial network systems prioritize user experience and personalized services. Leveraging big data analytics and artificial intelligence technologies, financial institutions can gain a deep understanding of user needs and preferences, providing customized financial products and services. This user-centric service philosophy not only enhances customer satisfaction and loyalty but also promotes segmentation and differentiated competition in the financial market. Furthermore, through mobile and instant service, financial network systems enable users to enjoy convenient financial services anytime, anywhere, further broadening the reach of financial services.

[0004] Modern financial network systems are also undergoing continuous improvement in terms of regulation and compliance. With the rapid development of financial technology, regulators have strengthened their oversight of financial network systems, promoting the formulation and improvement of relevant laws and regulations. At the same time, financial institutions have also actively responded to regulatory requirements, strengthening internal controls and risk management to ensure the compliance and security of financial services. This dual guarantee of regulation and compliance provides strong support for the healthy development of financial network systems.

[0005] In modern financial network systems, network security defense faces increasingly complex and diverse threats. Traditional defense methods rely on static rules and signature libraries, which are difficult to cope with rapidly evolving attack methods, resulting in low detection rates, high false alarm rates, and an inability to adapt to new security threats in a timely manner. Patent publication number CN119071051B discloses a network security risk control system based on traffic identification, which includes a traffic identification module, a data analysis module, and a behavior identification module. Among them, the traffic identification module is responsible for real-time analysis and monitoring of network data flows, checking the size, frequency, and protocol of the data; the data analysis module conducts in-depth analysis of the source and destination of the content identified by the traffic identification module to identify anomalies or potential attack patterns. The rule construction module collects historical anomaly log data and combines it with predefined rule information to gradually optimize and generate rule matching and network security identification mechanisms that adapt to the current operating environment, thereby achieving more accurate identification of abnormal user traffic behavior.

[0006] With the advancement of computing power and big data technologies, technologies such as deep learning and reinforcement learning have been introduced into the cybersecurity field, further enhancing the adaptability and predictive capabilities of defense systems. Researchers have begun exploring dynamic defense mechanisms based on artificial intelligence to improve the intelligence level of financial network security, thereby maintaining high defense effectiveness even in the face of unknown attacks. Summary of the Invention

[0007] This invention aims to overcome the shortcomings of existing technologies by providing a financial network security defense method and system based on multiple agents and a dynamic large-scale model. This method deploys detection agents at the edge layer to efficiently collect and analyze real-time traffic data and system logs from network nodes. It also builds cloud-layer decision-making and intelligence agents to provide strong support for developing effective defense strategies.

[0008] The object of the present invention is achieved through the following technical solutions: A financial network security defense method based on multi-agent collaboration and dynamic large models includes the following contents: S1: deploying detection agents at the edge layer to collect traffic data and system logs of financial network nodes in real time, extracting time series features through lightweight convolutional networks and generating preliminary anomaly scores; S2: building a decision agent at the cloud layer to receive the encrypted feature summaries transmitted by each edge node, inputting them into a dynamic large model for multimodal feature fusion, the dynamic large model is trained based on a reinforcement learning framework, and outputs a probability distribution of defense actions; S3: building a cross-institutional federated learning network through intelligence agents, aggregating threat features using secure multi-party computing protocols, and generating a global attack pattern knowledge graph; S4: building a dynamic game engine, constructing a benefit matrix based on attack costs and defense benefits, solving the Nash equilibrium strategy through linear programming, and generating the optimal defense instruction set; S5: dynamically allocating detection tasks to the edge or cloud for execution based on the attack threat level and edge computing power status, and triggering full cloud analysis when the edge computing power is lower than 20% or the threat level is high.

[0009] Through a multi-agent collaborative architecture, a closed loop of rapid edge response, in-depth cloud analysis, global intelligence sharing, and dynamic strategy optimization is achieved. This deeply integrates real-time data processing, cross-institutional collaboration, game theory optimization, and adaptive resource scheduling, improving the financial network's defense intelligence and system resilience in the face of new attacks.

[0010] This invention operates from edge to cloud: S1's anomaly scores and feature summaries are encrypted and transmitted to S2, serving as input to a dynamic large-scale model. From cloud to edge: S2 outputs a probability distribution of defensive actions, which is fed back to the edge for execution. Simultaneously, S5 dynamically adjusts task allocation based on cloud-based analysis results. Intelligence sharing: S3's knowledge graph provides S4 with attack pattern data (such as the payoff parameters of historically similar attacks), optimizing the payoff matrix. Game feedback: The Nash equilibrium strategy generated by S4 guides S2's reinforcement learning reward function (such as the weight of the blocking success rate), forming a closed-loop strategy optimization loop.

[0011] As a preferred method, it also includes: S6: quantizing the dynamic large model after training, converting the model parameters from FP32 to INT8 format, adjusting the dynamic range using the calibration set, and then deploying it to the edge node; S7: after the defense instruction is executed, collecting the blocking success rate and resource consumption indicators and feeding them back to the reinforcement learning reward function to update the parameters of the dynamic large model.

[0012] As a preferred method, the structure of the lightweight convolutional network described in S1 is: using the MobileNetV3 architecture, the input layer converts network traffic data into a grayscale image, and the width factor is set to 0.5; the output layer is connected to the Sigmoid activation function to generate an anomaly score of 0-1; when deployed on edge devices, the graph is optimized through the TensorRT engine to control the inference delay within 25ms.

[0013] As a preferred method, the training method of the dynamic large model in S2 includes: constructing a deep reinforcement learning framework, the policy network adopts a 12-layer Transformer structure, and the hidden layer dimension is 1024; the reward function R = 0.7·blocking success rate + 0.2·(1-false alarm rate)-0.1·resource consumption; using the PPO algorithm for distributed training, setting the KL divergence threshold δ = 0.01, and saving checkpoints every 100,000 steps; in the model fine-tuning stage, injecting 20% adversarial samples to improve robustness, and the adversarial sample generation adopts the FGSM method.

[0014] As a preferred method, the knowledge graph construction method in S3 includes: extracting entity nodes from multi-source logs, including IP addresses, vulnerability numbers, and malicious file hashes; defining the relationship type as {exploitation, propagation, control, penetration}, and constructing an attribute graph model; using a graph attention network to calculate node importance scores and identify key attack paths; when a new attack event occurs, similar historical patterns are retrieved through a subgraph matching algorithm, and the top-3 disposal solutions are returned.

[0015] As a preferred method, the implementation of the dynamic game engine in S4 includes: constructing an attack action set A={a1,a2,...,an} and a defense action set D={d1,d2,...,dm}; quantitatively calculating the defense benefit value Rij=α·blocking success rate+β·asset protection degree-γ·resource consumption of each attack and defense pair (ai,dj), where α=0.7, β=0.2, γ=0.1; inputting the benefit matrix into the Gurobi optimizer to solve the mixed strategy Nash equilibrium and outputting the defense action probability distribution vector P=[p1,p2,...,pm]; when a new attack mode is not defined in the matrix, the action set is expanded through the clustering algorithm and the equilibrium strategy is recalculated.

[0016] As a preferred method, the task allocation logic in S5 is as follows: define the threat level assessment function L = ω1*anomaly score + ω2*attack propagation speed + ω3*asset value, where ω1=0.6, ω2=0.3, and ω3=0.1; when L>0.8 or the edge CPU usage is>80%, the original traffic data is mirrored and transmitted to the cloud analysis cluster; otherwise, lightweight model inference is performed at the edge, and only the feature vector summary is uploaded to the cloud; the task allocation decision is updated every 200ms and broadcast to all Agent nodes through the publish-subscribe model.

[0017] As a preferred method, the specific process of model quantization in S6 includes: statistical model weight distribution histogram, determining the INT8 dynamic range parameter scale = 255 / (max_weight - min_weight); when forward propagating the calibration set, recording the 0.999 quantile of the activation value of each layer as the activation scaling factor; after quantization, the model is optimized through layer fusion, merging Conv-BN-ReLU into a single computational graph; enabling TensorCore acceleration during deployment, converting FP16 calculations to INT8 precision, and increasing the inference speed by 3 times.

[0018] A financial network security defense system based on multi-agent collaboration and dynamic large models includes: an edge detection module, deployed on financial network terminals, containing a lightweight CNN model and a rule engine for generating anomaly scores in real time; a cloud-based decision-making module, integrating a dynamic large model with a reinforcement learning framework, receiving multi-node feature data and outputting defense strategies; a federated learning module, which aggregates cross-institutional threat intelligence through a secure multi-party computing protocol, and contains an SM4 encryption unit and a homomorphic operation unit; a dynamic game engine, with a built-in attack and defense payoff matrix and a Nash equilibrium solver, supporting dynamic strategy optimization; a resource scheduling module, which allocates detection tasks based on threat level and computing power status, and contains an edge-cloud collaborative controller; and a blockchain audit module, which records attack events and disposal logs and provides Merkle tree verification and PBFT consensus functions.

[0019] As a preferred method, the edge detection module includes: a network traffic mirroring unit that converts the original data packet into a grayscale image; a lightweight inference unit that uses a MobileNetV3 model accelerated by TensorRT; and a local rule base that stores more than 200 predefined threshold rules.

[0020] This invention has at least the following beneficial effects: By deploying a detection agent at the edge layer, it enables efficient collection and analysis of real-time traffic data and system logs from network nodes. The construction of a cloud-layer decision agent enables the system to receive and process feature summaries from each edge node. By integrating multimodal features into a dynamic large-scale model, it further improves the ability to identify abnormal behavior. The intelligence agent provides strong support for developing effective defense strategies.

[0021] Furthermore, the construction of a dynamic game engine builds a payoff matrix based on attack costs and defense benefits. By solving the Nash equilibrium strategy through linear programming, it generates an optimal defense instruction set, optimizing the defense strategy and improving the efficiency and effectiveness of network security defense. By combining multi-agent collaboration with a dynamic large-scale model, financial network security defense significantly enhances the intelligence, automation, and efficiency of network security defense, providing a strong guarantee for the secure and stable operation of financial networks. BRIEF DESCRIPTION OF THE DRAWINGS To reveal the technical details of the embodiments of the present invention, the following is a brief introduction to the drawings involved in the embodiments. It should be emphasized that these drawings only illustrate several embodiments of the present invention and should not be considered as defining the scope of the invention. Those skilled in the art can deduce other relevant drawings based on these drawings without engaging in creative work.

[0022] Figure 1 This is a diagram of the core defense system of the present invention; Figure 2 This is a schematic diagram of the Agent collaborative closed loop mentioned in the embodiment. DETAILED DESCRIPTION

[0023] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.

[0024] In the following, embodiments of the present disclosure are described in detail with the aid of accompanying drawings. However, please be aware that the present disclosure is not limited to the specific forms shown herein. Rather, it should be understood to encompass various variations, equivalents, and / or alternatives to the embodiments of the present disclosure. In describing the drawings, the same reference numerals will be used to indicate similar components.

[0025] In this disclosure, terms are used to illustrate specific embodiments and do not constitute limitations of this disclosure. In this context, the use of the singular also encompasses the plural, unless the text clearly indicates otherwise. In the process of explanation, it should be understood that terms such as "including" or "having" are intended to indicate the presence of a feature, quantity, step, operation, structural component, part, or combination thereof, and do not preclude the possibility or addition of one or more other features, quantities, steps, operations, structural components, parts, or combinations thereof.

[0026] It should be understood that while the following description provides extensive specific details intended to facilitate a comprehensive understanding of the example embodiments, those skilled in the art will appreciate that the example embodiments can be implemented without these specific details. For example, systems may be presented in block diagram form to avoid excessive detail that would obscure the clarity of the examples. In other cases, unnecessary details regarding well-known processes, structures, and techniques may be omitted to maintain clarity of the examples.

[0027] A financial network security defense method based on multi-agent collaboration and dynamic large model, such as Figure 1 As shown, it includes the following contents: S1: Deploy detection agents at the edge layer to collect traffic data and system logs of financial network nodes in real time, extract time series features through lightweight convolutional networks and generate preliminary anomaly scores; S2: Construct decision agents at the cloud layer to receive encrypted feature summaries transmitted by each edge node, input dynamic large models for multimodal feature fusion, and the dynamic large models are trained based on the reinforcement learning framework to output the probability distribution of defense actions; S3: Construct a cross-institutional federated learning network through intelligence agents, use secure multi-party computing protocols to encrypt and aggregate threat features, and generate a global attack pattern knowledge graph; S4: Construct a dynamic game engine, construct a profit matrix based on attack costs and defense benefits, solve the Nash equilibrium strategy through linear programming, and generate the optimal defense instruction set; S5: Dynamically allocate detection tasks to the edge or cloud for execution based on the attack threat level and edge computing power status. When the edge computing power is less than 20% or the threat level is high, full cloud analysis is triggered. Reference Figure 2 Through a multi-agent collaborative architecture, it realizes a closed loop of rapid edge response - deep cloud analysis - global intelligence sharing - dynamic strategy optimization, deeply integrating real-time data processing, cross-institutional collaboration, game theory optimization and adaptive resource scheduling, and improving the defense intelligence level and system resilience of the financial network in the face of new attacks.

[0028] This embodiment integrates real-time response, cross-organizational collaboration, and dynamic policy tuning through a four-layer linkage: lightweight edge detection, intelligent cloud-based decision-making, global knowledge sharing, and flexible resource scheduling. Multiple agents, each performing their own functions while sharing data, ensure millisecond-level responses to high-frequency attacks on the financial network. Through continuous learning and game-playing, they build an adaptive shield against emerging threats.

[0029] From real-time detection at the edge to deep decision-making in the cloud, a lightweight detection agent is deployed at the edge layer to convert network traffic into grayscale images in real time, extract time-series features, and generate anomaly scores (0-1). 95% of normal traffic is filtered, and a summary of suspicious data is uploaded to reduce cloud load. The anomaly score is transmitted to the cloud via SM4 encryption, triggering subsequent analysis. Cloud-based decision-making utilizes a dynamic large-scale model to integrate multi-node features and output a probability distribution for defensive actions (e.g., a 70% probability of blocking an IP address). Model weights are dynamically adjusted based on historical defense effectiveness (blocking rate / false alarm rate), forming a closed loop of decision-making, feedback, and optimization. Federated learning units aggregate threat features (e.g., malicious IP addresses and vulnerability numbers) through encrypted encryption to construct a global knowledge graph. Historical attack paths in the graph (e.g., "vulnerability A to lateral penetration to data theft") are directly mapped to the payoff matrix of the game engine. For example, if the graph shows that an IP address has been associated with three DDoS attacks, its defense payoff weight in the payoff matrix is automatically increased by 20%. The game engine quantifies attack behaviors (such as DDoS and SQL injection) and defensive actions (such as throttling and WAF interception) into a payoff matrix and calculates the optimal hybrid strategy through Nash equilibrium. New attacks are automatically classified into similar historical patterns (clustering algorithm), and the payoff matrix is expanded and the strategy is re-solved in real time. When the cloud is overloaded, idle edge nodes are automatically utilized, and bandwidth is dynamically allocated through SDN.

[0030] In a preferred embodiment, the following steps are also included: S6: quantizing the dynamic large model after training, converting model parameters from FP32 to INT8 format, adjusting the dynamic range using a calibration set, and deploying it to edge nodes; performing distribution statistics on the FP32 model weights to generate a histogram; setting a scaling factor for mapping FP32 to INT8, and calculating the dynamic range parameter; using 5,000 attack traffic samples as a calibration set (covering scenarios such as DDoS and SQL injection), recording the 0.999th percentile of each layer's activation value (excluding extreme outliers), and calculating the activation scaling factor; completing the quantization conversion through weight conversion and activation value conversion; merging the Conv-BN-ReLU computation graph into a single computational graph to reduce memory accesses on edge devices; testing the quantized model on a Jetson Nano, comparing FP32 and INT8 accuracy: the accuracy drop was kept within 3%; and the inference speed was increased by more than 3 times. S7: After the defense instruction is executed, the blocking success rate and resource consumption metrics are collected and fed back into the reinforcement learning reward function to update the parameters of the dynamic large model. Feedback metrics collected include blocking success rate, resource consumption, and false alarm rate. Rewards are updated in real time. After the defense system executes an interception command, it collects three key metrics in real time: blocking success rate (the percentage of attacks successfully intercepted), resource consumption (CPU / memory usage), and false positive rate (the percentage of legitimate traffic incorrectly intercepted). This data is converted into a comprehensive "reward score." A high score indicates an effective strategy (e.g., a high blocking rate and low resource consumption), while a low score indicates problems (e.g., excessive false positives). Based on this reward score, the dynamic large-scale model automatically adjusts its internal parameters through reinforcement learning. These parameters primarily control the attention weights (which determine which attack features the model focuses on) and the connection strengths of the fully connected layers (which influences the probability of which defensive action is ultimately selected). If the "rate limiting" action is found to have a high success rate in blocking DDoS attacks but consumes too many resources, the model will gradually reduce the weight of this action while increasing the probability of selecting more effective actions such as "cloud traffic cleaning." This closed-loop learning mechanism continuously optimizes the strategy based on real-world results, ultimately achieving a dynamic defense that becomes smarter with use. In a preferred embodiment, S8: Attack event characteristics and handling logs are recorded on the blockchain, establishing a Merkle tree-based software package verification mechanism for real-time supply chain attack detection. Each time a cyberattack occurs, detailed attack characteristics (such as attack method, source IP address, exploited vulnerability) and handling results (such as how the attack was intercepted and by whom) are encrypted and packaged into a "block" and linked to the previous record. Once the record is on the chain, it cannot be secretly modified or deleted, effectively providing a permanent anti-counterfeiting label for the attack. To prevent software supply chain attacks (such as malicious code injection), the system generates a "digital ID" for each software package—a Merkle tree. During the development phase, the software package is divided into multiple small file blocks, and a unique hash value (similar to a fingerprint) is generated for each block. This is then aggregated layer by layer to generate a final root hash and stored on the blockchain.Download Verification: When a user downloads software, the system automatically recalculates the hashes of all blocks to generate a new root hash. This is then compared to the original root hash stored on the blockchain. If the hashes match, the software has not been tampered with and is running safely. If they do not, it indicates that a file block has been maliciously modified (e.g., by a virus), triggering an immediate alert. Once an attack is uploaded to the chain, all institutions can trace the attack pattern and jointly prevent and control it. Even if a punctuation mark is tampered with in the software, the Merkle tree verification can detect it in seconds, preventing "infected software" from infiltrating the financial system.

[0031] S9: When a new attack pattern is detected, the federated learning model update process is triggered. Each participant calculates encrypted gradients based on local data and achieves secure aggregation through a threshold signature protocol. When a new cyberattack is discovered, the system activates the federated learning collaborative defense mechanism. Local encrypted learning: Each financial institution trains the model locally using its own attack data (such as virus samples and anomaly logs), but only generates "encrypted learning results" (gradients). This is similar to how each doctor writes an encrypted treatment plan after studying a case, without revealing specific patient information. Fragmented signature: Each participant breaks the encrypted result into multiple fragments and distributes them to other institutions for safekeeping. This is similar to splitting a safe combination into three parts; at least two institutions must provide the fragments to complete the combination. Secure aggregate update: All institutions use the combined combination to decrypt the aggregated results, merging them into a more powerful global defense model. Even if someone hacks one institution, they cannot steal the complete model independently, and the new model can recognize the latest attack methods encountered by all institutions. This approach protects data privacy while enabling all institutions to "share wisdom," quickly forming a collective immunity to new viruses, just like the immune system. S10: Software-defined networking dynamically adjusts edge-to-cloud communication bandwidth based on real-time network load, prioritizing 1Gbps dedicated channels during high-risk attacks. Financial networks monitor network load (such as data congestion and attack frequency) at each node in real time through an intelligent traffic scheduling system. Under normal circumstances, the system automatically balances communication bandwidth between edge devices and the cloud to ensure smooth operation of regular services, similar to how intelligent transportation systems dynamically adjust lane traffic flow. Once a high-risk attack (such as a DDoS flood) is detected, the system immediately initiates an emergency response: Bandwidth reallocation: The SDN controller (network command center) switches communication traffic in the attacked area to a backup channel to ensure core services (such as payment transactions) are not impacted. Dedicated channel prioritization: A dedicated 1Gbps channel (similar to an ambulance lane) is established for security analysis data, mirroring attack traffic directly to the cloud detection cluster to avoid bandwidth competition with normal traffic. After the attack is mitigated, bandwidth balance is automatically restored to prevent resource waste. Even under large-scale attacks, critical service latency remains within 10 milliseconds, improving defense response speed.

[0032] In a preferred embodiment, the lightweight convolutional network described in S1 is structured as follows: Using an improved MobileNetV3 architecture, the input layer converts network traffic data into grayscale images with a width factor set to 0.5; the output layer uses a Sigmoid activation function to generate anomaly scores ranging from 0 to 1; when deployed on edge devices, graph optimization is performed using the TensorRT engine to keep inference latency below 25ms. This embodiment, based on the lightweight MobileNetV3 architecture, is specifically designed for edge devices. Input processing: Network traffic (such as packet size and protocol type) is converted into grayscale images in time series, similar to visualizing traffic fluctuations as black and white heat maps, making it easier for the model to detect anomaly patterns. Model slimming: By using a "width factor of 0.5" to compress the number of network layer channels, the model size is reduced by half (for example, from 4MB to 2MB), saving more resources when running on edge devices. Anomaly scoring: The output layer uses a Sigmoid function to convert the model's judgment results into a score between 0 and 1 (for example, 0.9 indicates a high-risk attack). Exceeding a threshold (for example, 0.7) immediately triggers an alert. Ultra-fast inference: During deployment, the TensorRT engine combines computational steps (such as convolution and normalization) to accelerate inference, ensuring each detection takes less than 25 milliseconds. On edge devices like the Raspberry Pi, this allows for rapid attack identification without slowing down normal operations.

[0033] In a preferred embodiment, the training method of the dynamic large model in S2 includes: constructing a deep reinforcement learning framework, the policy network adopts a 12-layer Transformer structure, and the hidden layer dimension is 1024; the reward function R=0.7*blocking success rate+0.2*(1-false alarm rate)-0.1*resource consumption; using the PPO algorithm for distributed training, setting the KL divergence threshold δ=0.01, and saving checkpoints every 100,000 steps; in the model fine-tuning stage, injecting 20% adversarial samples to improve robustness, and the adversarial sample generation adopts the FGSM method.

[0034] The intelligent decision-making network utilizes a 12-layer Transformer architecture (similar to a multi-layer information processing pipeline), with 1024 neurons per layer. It excels at capturing complex patterns in cyberattacks. Scoring criteria include: 70% blocking success rate, which rewards accurate interception of real attacks; 20% false positive rate, which penalizes accidental impact on legitimate traffic; and 10% resource consumption, which prevents excessive computing power consumption by defensive actions. Training strategy: The PPO algorithm (similar to a coach's error correction mechanism) is used to control the model update amplitude and prevent sudden policy failures. Training is automatically archived every 100,000 iterations to prevent unexpected training interruptions. 20% adversarial examples (artificially constructed "attack feints") are added to enable the model to learn to identify attack variants, similar to how a vaccine injects a weakened virus to boost immunity. This model can both quickly block real attacks and flexibly balance security with system resource consumption.

[0035] In a preferred embodiment, the knowledge graph construction method in S3 includes: extracting entity nodes from multi-source logs, including IP addresses, vulnerability numbers, and malicious file hashes; defining the relationship type as {exploit, spread, control, penetration} and constructing an attribute graph model; using a graph attention network to calculate node importance scores and identify key attack paths; and when a new attack event occurs, using a subgraph matching algorithm to retrieve similar historical patterns and return the top-three remediation solutions. The knowledge graph in this embodiment resembles a map of the attack relationship network. Its construction and use consists of four steps: Node mapping: Extracting key information from sources such as firewall logs and vulnerability databases as "strongholds," including hacker IP addresses, vulnerability numbers, and virus file hashes. Connection relationships: Hackers infiltrate servers through a vulnerability; Viruses spread from a single device to the intranet; Attackers remotely control compromised hosts. Using a graph attention network, the most critical nodes (such as frequently exploited vulnerabilities or recurring attack IP addresses) are automatically analyzed, highlighted in red, and prioritized for defense. When a new attack (such as a ransomware intrusion) is discovered, the system automatically compares it to historical attack maps, quickly identifying the three most similar cases (e.g., a similar attack from last year), and recommending successful solutions from that time (e.g., isolating the device and plugging the vulnerability). Security personnel, like detectives with access to a "criminal database," can instantly access their experience solving new cases, significantly reducing emergency response time.

[0036] In a preferred embodiment, the implementation of the dynamic game engine in S4 includes: constructing an attack action set A = {a1, a2, ..., an} and a defense action set D = {d1, d2, ..., dm}; for example: a1 = DDoS attack, a2 = SQL injection, a3 = ransomware propagation. d1 = IP throttling, d2 = WAF interception, d3 = virtual machine isolation. For each attack-defense pair (ai, dj), the defense benefit Rij is quantified as: α·blocking success rate + β·asset protection - γ·resource consumption, where α = 0.7, β = 0.2, and γ = 0.1 (weights can be adjusted dynamically). The blocking success rate is calculated as a historical statistical value (e.g., a throttling blocking rate of 80% for DDoS attacks); the asset protection is calculated as the weighted CVSS score of the protected asset (e.g., a weight of 0.9 for core financial systems); and the resource consumption is calculated as CPU / bandwidth utilization normalized to [0, 1]. The payoff matrix is input into the Gurobi optimizer to solve the mixed strategy Nash equilibrium and output the defense action probability distribution vector P = [p1, p2, ..., pm]. When a new attack mode is not defined in the matrix, the action set is expanded through the clustering algorithm and the equilibrium strategy is recalculated.

[0037] Assume that the current profit matrix is: Construct the linear programming equation: in, It is the minimum guaranteed profit for the defender. in It's a defensive move The probability of selection, j Can be 1 or 2. Solved by Gurobi optimizer .

[0038] When a new attack is detected When: Feature clustering: Using DBSCAN algorithm, calculate Similarity with existing attack features (such as traffic patterns, vulnerability exploitation methods). If the similarity is greater than 80%, it is classified as the same type of attack. Profit matrix expansion: Add a new row , whose benefit value is generated by similarity weighting: ; Recalculate equilibrium: The expanded matrix is solved through the same linear programming process to generate a defense strategy that includes the new attack.

[0039] In a preferred embodiment, the task allocation logic in S5 is as follows: define a threat level assessment function L = ω1*anomaly score + ω2*attack propagation speed + ω3*asset value, where ω1=0.6, ω2=0.3, and ω3=0.1; when L>0.8 or the edge CPU usage is>80%, the original traffic data is mirrored and transmitted to the cloud analysis cluster; otherwise, lightweight model inference is performed at the edge, and only the feature vector summary is uploaded to the cloud; the task allocation decision is updated every 200ms and broadcast to all Agent nodes through a publish-subscribe model.

[0040] Every 200 milliseconds, the current attack threat is assigned a comprehensive score (L). The calculation method is as follows: 60% is based on the anomaly score (the model's assessment of the likelihood of an attack); 30% is based on the spread rate (e.g., the number of devices infected per second); and 10% is based on the value of the attack target (i.e., the asset value, such as whether it targets core trading systems). Traffic diversion decisions: In high-risk scenarios (L > 0.8 or edge CPU utilization exceeding 80%), raw traffic is mirrored and uploaded to the cloud cluster for in-depth analysis to ensure that complex attacks are not missed. In normal scenarios, edge devices quickly process traffic (using lightweight models to generate results) and only upload key feature summaries to reduce bandwidth pressure. Real-time synchronization: After each decision is made, it is instantly synchronized to all nodes through a publish-subscribe model to ensure consistent policies across the entire network. When ransomware spreads rapidly (high spread rate score) and attacks database servers (high asset value score), the system immediately mirrors traffic to the cloud and activates expert-level models for a comprehensive attack. In the case of a simple port scan (low threat), the edge device handles it independently without alerting the cloud.

[0041] In a preferred embodiment, the specific process of model quantization in S6 includes: statistically analyzing the model weight distribution histogram to determine the INT8 dynamic range parameter scale = 255 / (max_weight-min_weight); when forward propagating the calibration set, recording the 0.999 quantile of the activation value of each layer as the activation scaling factor; after quantization, the model is optimized through layer fusion, merging Conv-BN-ReLU into a single computational graph; enabling TensorCore acceleration during deployment, converting FP16 calculations to INT8 precision, and increasing the inference speed by 3 times.

[0042] Parameter range measurement: Statistically analyze the distribution of model weights (for example, 90% of the weights are between -1 and 1), determine the compression ratio, and evenly map FP32 floating-point numbers to integers in the 0-255 range, similar to adjusting the brightness of an image to prevent overexposure. Activation value calibration: Test the model with a batch of typical attack data (calibration set) and record the maximum output range of each layer. For example, if 99.9% of the output values of a layer are less than 50, they are scaled to 50 to avoid computational overflow. Calculation step merging: Combine the previously separate convolution, normalization, and activation function calculations into a single step, reducing data transfer time and improving inference speed.

[0043] Hardware-accelerated deployment: Quantized models are run on dedicated chips (TensorCores) on edge devices that support INT8 computing, such as the NVIDIA Jetson. The quantized models run incredibly fast on devices like the Raspberry Pi, identifying an attack in the blink of an eye with less than a 3% drop in accuracy, achieving a perfect balance between efficiency and precision.

[0044] In a preferred embodiment, the blockchain traceability mechanism in S8 includes: writing the attack event feature hash value H=SM3 (attack IP+timestamp+vulnerability number) into the blockchain; constructing a Merkle tree to verify the integrity of the supply chain software, with the leaf node being the software block hash value Hi=SHA256(Block_i); the downloading node recalculates the root hash H_root' and compares it with the signature value stored in the blockchain. If there is any inconsistency, isolation is triggered; using the PBFT consensus algorithm, four verification nodes are set, and three nodes must reach a consensus before a new block can be written.

[0045] This mechanism ensures the immutability of attack records and the security of the software supply chain through dual anti-counterfeiting verification: Attack event evidence: Each attack feature (such as the attacking IP address and vulnerability ID) is packaged and generated into a unique digital fingerprint (hash) and written to the blockchain. This acts as a "security ID" for each attack event; any modification will cause the fingerprint to change, immediately exposing any tampering. Software integrity verification: On the developer side, the software is divided into multiple blocks, each of which generates a unique fingerprint. These are then aggregated to create a total fingerprint (Merkle root) and stored on the blockchain. On the user side, after downloading the software, the fingerprints of each block and the total fingerprint are recalculated and compared with the blockchain record. If even one block is tampered with (for example, by inserting a virus), the total fingerprint will not match, and the system will automatically quarantine and dispose of it. Consensus anti-cheating: The blockchain network has four verification nodes and utilizes the PBFT consensus mechanism. Any new data (attack records or software fingerprints) must be verified by at least three nodes before it can be uploaded to the blockchain. Even if a single node is controlled by another party, data forgery is impossible, ensuring global trustworthiness.

[0046] In a preferred embodiment, the specific process of updating the federated learning model in S9 includes: each participating institution uses the SM4 algorithm to encrypt the local model gradient Gi (the local model gradient of the i-th participating institution, that is, the model parameter adjustment direction) to generate the ciphertext Ci=Enc(Ki,Gi) (the ciphertext generated by encrypting Gi using the SM4 algorithm and the key Ki); the key Ki (the symmetric key used to encrypt Gi) is sharded and stored on 5 nodes through the Shamir secret sharing protocol, and can only be decrypted when 3 nodes cooperate; the aggregation server performs a homomorphic addition operation C_sum=ΣCi mod p (directly summing the encrypted gradient ciphertext and then performing a modulo operation, p is a large prime number to control the numerical range to prevent overflow), and obtains the global gradient ΔG after decryption (the global gradient obtained after aggregating the gradients of all participants, which guides the model update direction); the FedAvg algorithm is used to update the global model parameter θ=θ-η·ΔG, where θ is the global model parameter (such as the neural network weight), and the learning rate η is dynamically adjusted according to the amount of data from the participating parties; a digital signature is attached when the model is issued, and the edge node verifies the signature integrity through the national secret SM9 algorithm before loading the model.

[0047] This mechanism enables financial institutions to jointly train risk control models, sharing security intelligence without exfiltrating data locally. Encrypted local training and data privacy protection: Each bank encrypts its training results (such as fraud detection parameters) using the national secret SM4 algorithm, generating encrypted data packets. Sharded key escrow: The decryption key is split into five shares and distributed across five independent nodes (e.g., the central bank, UnionPay, and a third-party certification agency). Decryption requires the collaboration of at least three parties to prevent single-point leakage. Secure aggregated updates and homomorphic encrypted computation: A cloud-based aggregation server directly performs mathematical operations on the encrypted data packets, merging them to generate global model update parameters. Similar to joint accounting: Each bank submits encrypted transaction risk parameters, which are aggregated and calculated in the cloud without exposing individual bank data details. Dynamic weight allocation: Contribution weights are automatically adjusted based on the data volume of participating parties, with banks with larger transaction data volumes having a greater influence on the model. Trusted model delivery and digital signature verification: The updated model is digitally signed with the national secret SM9 algorithm. Edge nodes (e.g., ATMs and payment systems) verify the signature before loading, ensuring the model has not been tampered with. Secure synchronization in seconds: Delivered via a dedicated financial network channel, the upgrade is completed within a minute across thousands of nodes across the network. Defensive effectiveness and data leakage prevention: Even if an attacker steals a single bank's data, they cannot decrypt the complete model. Co-evolution: When a bank discovers a new fraud method, the model is updated and defense capabilities are synchronized across the entire network. Compliance auditing: Contribution records of all participants are recorded on-chain to meet financial regulatory requirements. When a bank discovers a new phishing attack pattern in a cross-border transaction, this mechanism can update the interception rules of the risk control model across the entire network within three hours without sharing specific customer data.

[0048] In a preferred embodiment, the dynamic bandwidth adjustment method in S10 includes: deploying a traffic classifier on the SDN controller to identify three types of traffic: transaction data, detection data, and management signaling; setting a baseline bandwidth allocation policy: 50% for transaction data, 30% for detection data, and 20% for management signaling; when a DDoS attack is detected, increasing the detection data bandwidth to 60% and using compressed transmission for management signaling; and adjusting the switch queue weights in real time through the OpenFlow protocol to ensure that the latency of high-priority tasks is less than 10ms. This mechanism ensures the stable operation of core financial services through intelligent bandwidth management and intelligent traffic classification. The SDN controller monitors network traffic in real time and divides it into three categories based on service priority: Core transaction data (payment, clearing, etc.): 50% of the bandwidth is allocated by default to ensure millisecond-level response for high-priority services; security detection data (attack signature analysis): 30% of the bandwidth is allocated for real-time scanning of abnormal traffic; management and control signaling (device heartbeat, configuration instructions): 20% of the bandwidth is reserved to maintain basic system operations and maintenance.

[0049] Attack emergency response: When a DDoS attack is detected (such as a massive number of false requests impacting a payment gateway): security detection bandwidth is increased to 60%, enhancing the ability to deeply analyze attack traffic and quickly locate the attack source IP address; management signaling intelligent compression: Operation and maintenance instructions are compressed into a streamlined format (such as JSON to binary protocol), reducing bandwidth usage from 20% to 8%, freeing up resources to prioritize transaction protection and defense. Real-time priority scheduling dynamically adjusts network device queues through the OpenFlow protocol: Core transaction channel: Payment packets are marked as "urgent" and forwarded by switches with priority, with latency strictly controlled within 10ms; Attack traffic rate limiting: Packets identified as malicious requests are automatically downgraded to the lowest priority queue to prevent them from squeezing out normal bandwidth; Defense linkage: When the security analysis cluster discovers new attack characteristics, it automatically applies for temporary bandwidth expansion (for example, from 30% to 45%).

[0050] A financial network security defense system based on multi-agent collaboration and dynamic large models includes: an edge detection module, deployed on financial network terminals, containing a lightweight CNN model and a rule engine for generating anomaly scores in real time; a cloud decision module, integrating a dynamic large model and a reinforcement learning framework, receiving multi-node feature data and outputting defense strategies; a federated learning module, which realizes cross-institutional threat intelligence aggregation through a secure multi-party computing protocol, and contains an SM4 encryption unit and a homomorphic operation unit; a dynamic game engine, with a built-in attack and defense payoff matrix and a Nash equilibrium solver, supporting dynamic strategy optimization; a resource scheduling module, which allocates detection tasks according to threat level and computing power status, and contains an edge-cloud collaborative controller; a blockchain audit module, which records attack events and disposal logs, and provides Merkle tree verification and PBFT consensus functions; a model management module, which is responsible for post-training quantization and version control, and supports the distribution and verification of encryption models; a network control module, which dynamically adjusts bandwidth allocation based on software-defined networks and integrates the OpenFlow protocol stack.

[0051] This system builds a layered, collaborative, and dynamically evolving intelligent defense system. Real-time edge sentinels (edge detection modules) are deployed on bank terminals (such as ATMs and payment gateways). Using lightweight AI models (such as MobileNetV3), they scan network traffic in seconds and generate anomaly scores. Suspicious behavior (such as high-frequency, abnormal access) is immediately reported to the cloud. The cloud-based intelligent brain (cloud-based decision module + dynamic game engine) aggregates feature data from multiple nodes. The dynamic large-scale model combines historical attack patterns (knowledge graph) to generate defense strategies. For example, when a new DDoS attack is detected, the game engine calculates the cost-benefit of attack and defense and selects the optimal solution (such as "limiting traffic by 70% and enabling cloud-based scrubbing"). A cross-institutional joint defense network (federated learning module + blockchain audit) allows financial institutions to share threat signatures (such as ransomware hashes) through encrypted aggregation, but the original data remains private. The blockchain records the fingerprints of all attack events, and any tampering attempts are verified and blocked by Merkle tree verification. Elastic resource scheduling (resource scheduling module + network control module) dynamically allocates computing power based on the threat level: standard attacks are handled autonomously by the edge, while high-risk attacks trigger full traffic analysis in the cloud. SDN also dynamically allocates bandwidth (for example, increasing security analysis bandwidth from 30% to 60% during a DDoS attack). A closed-loop evolutionary mechanism (model management module) provides real-time feedback on defense effectiveness to the reinforcement learning model, automatically optimizing policy network parameters. The updated model is encrypted and quantized (e.g., FP32 to INT8) before being delivered to edge devices, ensuring continuous improvement of defense capabilities across the entire network.

[0052] In this embodiment, lightweight edge detection (S1) and cloud-based deep decision-making (S2) work together; federated learning (S3) and dynamic game (S4) achieve global strategy optimization; and threat classification scheduling (S5) ensures efficient resource utilization, forming a complete closed loop of perception-decision-action-evolution.

[0053] In a preferred embodiment, the edge detection module includes: a network traffic mirroring unit that converts raw data packets into grayscale images; a lightweight inference unit that uses a MobileNetV3 model accelerated by TensorRT; a local rule base that stores more than 200 predefined threshold rules; and an encrypted communication unit that implements SM4 national encryption algorithm and TLS 1.3 dual-channel encryption.

[0054] This module is an intelligent security checkpoint deployed at financial terminals (such as ATMs and POS terminals), including: Data mirroring conversion: The network traffic mirroring unit captures data packets in real time and converts traffic features (such as protocol type, packet size, and time interval) into grayscale images to facilitate the model's identification of abnormal timing patterns.

[0055] Lightweight real-time detection: MobileNetV3 model (accelerated by TensorRT): Rapidly analyzes grayscale images on edge devices and outputs an anomaly score of 0-1 (e.g., 0.9 indicates a high-risk attack). Local rule library: Built-in 200+ predefined rules (e.g., "50 login requests from the same IP within 1 second"), enabling rapid rule matching and blocking of known attacks. Double-security transmission: SM4 national secret encryption encrypts anomaly scores and feature summaries to prevent data leakage; TLS 1.3 tunnel establishes a bank-grade transmission tunnel to prevent middleman eavesdropping. Data is delivered to the cloud with dual internal and external protection, like an "armed escort vehicle." This implementation enables the local rule base to intercept common attacks (such as brute force attacks) in seconds. Complex attacks (such as APT penetration) are analyzed and reported encrypted after model analysis. Edge resource utilization is kept below 20%, ensuring that terminal services are not impacted. If a branch ATM detects an abnormal transfer request, the local rule base quickly blocks the operation. Simultaneously, the model identifies the IP address associated with historical attack patterns and reports it encrypted to the cloud, triggering a network-wide ban.

[0056] In a preferred embodiment, the federated learning module includes a key management unit that implements sharded key storage based on the Shamir threshold protocol; a security aggregation unit that supports Paillier homomorphic encryption and the MPC protocol; a compliance check unit that integrates knowledge graphs to verify the legitimacy of attack signatures; and a disconnection recovery unit that enables historical gradient caching when participants are offline. This module enables secure collaborative modeling for financial institutions, ensuring both data privacy and model effectiveness.

[0057] Key Sharding: The encryption key is split into multiple shares (e.g., five) using the Shamir threshold protocol and stored in distributed locations across different nodes. Decryption requires at least three shares of the key. Even if one or two nodes are compromised, the attacker cannot obtain the complete key.

[0058] Privacy-focused aggregate computing and homomorphic encryption: Each institution uploads encrypted model parameters, and the cloud directly performs mathematical operations (such as addition) on the ciphertext, aggregating global results without decryption. Multi-party secure computation (MPC): Collaborative computing with fully encrypted sensitive data. Compliance and attack resistance: Knowledge graph verification automatically verifies whether the data characteristics submitted by institutions (such as IP addresses and vulnerability numbers) match historical attack patterns, blocking the injection of forged data. Disconnection recovery: If an institution loses network access, its historical data cache is automatically used to continue training, and participation can be resumed even after offline for 72 hours. This invention prevents raw data from being exported locally, and key sharding prevents single-point breaches. Abnormal data submissions are intercepted by the knowledge graph to ensure model reliability. Temporary disconnection of institutions is supported to prevent training interruptions from impacting risk control model iteration. A bank discovered a new phishing attack signature, and its federated learning module aggregated data from 88 institutions within three hours, generating global interception rules without leaking any customer transaction details.

[0059] In a preferred embodiment, the dynamic game engine includes: a payoff matrix construction unit that collects attack cost and defense benefit indicators in real time; a linear programming solver that calls the Gurobi algorithm library to calculate hybrid strategies; a policy interpreter that converts probability distributions into executable defense instructions; and a feedback learning unit that updates payoff matrix parameters based on the response results. This engine formulates the optimal defense strategy in real time. The payoff matrix construction collects the costs of attack methods (such as DDoS attack bandwidth consumption) and the benefits of defense actions (such as blocking success rate and resource consumption) in real time, creating a dynamic payoff comparison table. For example, intercepting ransomware is highly profitable but time-consuming; throttling suspicious IPs is less costly but may result in missed attacks. The optimal strategy deduction (linear programming solver) calls the Gurobi algorithm library to calculate the balance point (Nash equilibrium) between the attacker and defender, generating a hybrid strategy. Probability distribution: For example, a 70% probability of deep ransomware interception and a 30% probability of rapid throttling; resource constraints: ensuring that computing power consumption does not exceed the node's tolerance limit. The policy interpreter converts the probability distribution into executable instructions. Block IP addresses and limit bandwidth by 60% (corresponding to a 70% deep interception strategy); temporarily isolate virtual machines and manually review alerts (corresponding to a 30% rapid response strategy). In a preferred embodiment, the resource scheduling module includes: a computing power monitoring unit that collects real-time CPU / GPU utilization and memory usage; a task classifier that assesses threat levels based on a random forest model; a load balancer that optimizes resource allocation using a Q-Learning algorithm; and an emergency response unit that triggers a full cloud takeover mechanism in the event of a high-risk attack. This module dynamically allocates computing power to balance efficiency and security. The computing power monitoring unit continuously monitors metrics such as CPU / GPU load and memory usage on each node. The task classifier uses a random forest model to quickly determine the threat level of the attack. Low-risk (e.g., port scan) attacks are marked as "green" and allocated 10% of computing power; medium-risk (e.g., abnormal login) attacks are marked as "yellow" and allocated 30% of computing power; high-risk (e.g., 0-day vulnerability exploit) attacks are marked as "red," triggering a subsequent emergency response. The load balancer uses the Q-Learning algorithm to learn the optimal allocation strategy. For example, the normal allocation is: 60% of computing power is allocated to the trading system, 30% to security detection, and 10% to operations and maintenance. During an attack, if ransomware is detected, security detection computing power is increased to 50% and operations and maintenance resources are reduced to 5%. If the emergency response unit identifies a "red flag" attack, it automatically mirrors traffic to the cloud analysis cluster and activates the full detection model. Edge devices only perform basic filtering, conserving resources to deal with ongoing attacks.

[0060] This paper proposes a financial network security defense system and method based on multi-agent collaboration and a dynamic large-scale model. This system achieves efficient protection through a four-layer linkage mechanism: lightweight edge detection, intelligent cloud-based decision-making, federated knowledge sharing, and dynamic resource scheduling. Edge nodes use the MobileNetV3 lightweight model to generate anomaly scores in real time, combining it with a local rule base to intercept conventional attacks in seconds. The cloud-based dynamic large-scale model integrates multi-node data, combines it with an attack-defense game engine to calculate Nash equilibrium strategies, and continuously optimizes model weights through reinforcement learning. The cross-institutional federated learning module utilizes SM4 encryption and Shamir key sharding to securely aggregate threat signatures and construct a global attack knowledge graph. The resource scheduling module dynamically allocates computing power and bandwidth based on threat scores. The system integrates blockchain auditing and model quantification technologies to ensure tamper-proof attack traceability and accelerate edge inference, forming a full-cycle defense system consisting of real-time perception, game decision-making, elastic response, and closed-loop evolution. This system is suitable for highly sensitive financial scenarios such as payment clearing and securities trading.

[0061] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as covering the preferred embodiments and all changes and modifications that fall within the scope of the invention. The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A financial network security defense method based on multi-agent collaboration and dynamic large model, characterized by: Includes the following: S1: Deploy a detection agent at the edge layer to collect traffic data and system logs from financial network nodes in real time, extract time series features through a lightweight convolutional network, and generate preliminary anomaly scores; S2: Build a decision-making agent at the cloud layer, receive encrypted feature summaries from each edge node, and input them into a dynamic large model for multimodal feature fusion. The dynamic large model is trained based on a reinforcement learning framework and outputs a probability distribution of defensive actions. S3: Build a cross-institutional federated learning network through intelligence agents, aggregate threat features using a secure multi-party computing protocol, and generate a global attack pattern knowledge graph. S4: Build a dynamic game engine, construct a profit matrix based on attack costs and defense benefits, solve the Nash equilibrium strategy through linear programming, and generate the optimal defense instruction set; S5: Dynamically allocate detection tasks to the edge or cloud for execution based on the attack threat level and edge computing power status. When the edge computing power is lower than 20% or the threat level is high, full cloud analysis is triggered.

2. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 1 is characterized by: Also includes: S6: Quantize the dynamic large model after training, convert the model parameters from FP32 to INT8 format, adjust the dynamic range using the calibration set, and then deploy it to the edge node; S7: After the defense command is executed, the blocking success rate and resource consumption indicators are collected and fed back to the reinforcement learning reward function to update the dynamic large model parameters.

3. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 1 is characterized in that: The structure of the lightweight convolutional network described in S1 is as follows: the architecture of MobileNetV3 is adopted, the input layer converts the network traffic data into a grayscale image with a width factor set to 0.5; the output layer is connected to the Sigmoid activation function to generate anomaly scores of 0-1; When deployed on edge devices, graph optimization is performed through the TensorRT engine to control inference latency to less than 25ms.

4. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 1 is characterized in that: The training method for the dynamic large model in S2 includes: building a deep reinforcement learning framework, using a 12-layer Transformer structure for the policy network with a hidden layer dimension of 1024; the reward function R = 0.7·blocking success rate + 0.2·(1-false alarm rate)-0.1·resource consumption; using the PPO algorithm for distributed training, setting the KL divergence threshold δ = 0.01, and saving checkpoints every 100,000 steps; during the model fine-tuning stage, injecting 20% adversarial samples to improve robustness, and using the FGSM method to generate adversarial samples.

5. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 1 is characterized in that: The knowledge graph construction method in S3 includes: extracting entity nodes from multi-source logs, including IP addresses, vulnerability numbers, and malicious file hashes; defining the relationship type as {exploitation, propagation, control, penetration} and building an attribute graph model; using a graph attention network to calculate node importance scores and identify key attack paths; when a new attack event occurs, similar historical patterns are retrieved through a subgraph matching algorithm and the top-3 disposal solutions are returned.

6. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 1 is characterized in that: The implementation of the dynamic game engine in S4 includes: constructing an attack action set A={a1,a2,...,an} and a defense action set D={d1,d2,...,dm}; quantitatively calculating the defense payoff value Rij=α·blocking success rate+β·asset protection degree-γ·resource consumption for each attack and defense pair (ai,dj); inputting the payoff matrix into the Gurobi optimizer to solve the mixed strategy Nash equilibrium and outputting the defense action probability distribution vector P=[p1,p2,...,pm]; when a new attack mode is not defined in the matrix, the action set is expanded through the clustering algorithm and the equilibrium strategy is recalculated.

7. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 1 is characterized in that: The task allocation logic in S5 is as follows: define the threat level assessment function L = ω1*anomaly score + ω2*attack propagation speed + ω3*asset value, where ω1=0.6, ω2=0.3, and ω3=0.1; when L>0.8 or the edge CPU usage is >80%, the original traffic data is mirrored and transmitted to the cloud analysis cluster; otherwise, lightweight model inference is performed at the edge, and only the feature vector summary is uploaded to the cloud; the task allocation decision is updated every 200ms and broadcast to all agent nodes through a publish-subscribe model.

8. The financial network security defense method based on multi-agent collaboration and dynamic large model according to claim 2 is characterized in that: The specific process of model quantization in S6 includes: statistically analyzing the model weight distribution histogram and determining the INT8 dynamic range parameter scale = 255 / (max_weight - min_weight); when forward propagating the calibration set, recording the 0.999 quantile of the activation value of each layer as the activation scaling factor; after quantization, the model is optimized through layer fusion, merging Conv-BN-ReLU into a single computational graph; enabling TensorCore acceleration during deployment and converting FP16 calculations to INT8 precision.

9. A financial network security defense system based on multi-agent collaboration and dynamic large model, characterized by: include: The edge detection module, deployed on financial network terminals, includes a lightweight CNN model and a rule engine for generating anomaly scores in real time; The cloud-based decision-making module integrates a dynamic large model with a reinforcement learning framework, receives multi-node feature data, and outputs defense strategies; The federated learning module aggregates cross-institutional threat intelligence through a secure multi-party computing protocol, including an SM4 encryption unit and a homomorphic computing unit. The dynamic game engine includes a built-in attack and defense payoff matrix and a Nash equilibrium solver, supporting dynamic strategy optimization. The resource scheduling module allocates detection tasks based on threat levels and computing power status, and includes an edge-cloud collaborative controller; the blockchain audit module records attack events and disposal logs, and provides Merkle tree verification and PBFT consensus functions.

10. The financial network security defense system based on multi-agent collaboration and dynamic large model according to claim 9 is characterized in that: The edge detection module includes: a network traffic mirroring unit that converts raw data packets into grayscale images; a lightweight inference unit that uses a MobileNetV3 model accelerated by TensorRT; and a local rule library that stores more than 200 predefined threshold rules.

Citation Information

Patent Citations

  • A network security risk control system based on traffic identification

    CN119071051B

Cited By

  • Distributed network security early warning method based on cloud computing

    CN120979834A

  • Controllable full-stack security detection method and device, electronic equipment and storage medium

    CN121056251A

  • High-efficiency network monitoring and scheduling system based on cloud edge collaborative architecture

    CN121334239A

  • A high-efficiency network monitoring and scheduling system based on cloud-edge collaborative architecture

    CN121334239B

  • Multi-tenant adaptive cooperative defense method and system in hybrid cloud scene

    CN121530767A