Financial black and gray production attack and defense self-evolution method and system based on adversarial reinforcement learning
Patent Information
- Application Number
- CN202610799286.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]为了改善现有方法容易滞后和失效的问题,本申请公开如下技术方案:
[0013] The beneficial effects of this embodiment are as follows: the attack strategy agent automatically generates an attack strategy that simulates black and gray market attacks, and then the defense decision agent generates a corresponding defense strategy based on the attack strategy. Thus, this solution can dynamically update the defense strategy and the attack strategy used to train the defense decision agent through the confrontation between the attack strategy agent and the defense decision agent, so that the attack strategy and the defense strategy can be updated in real time in practical applications, solving the problem that the strategy is prone to lag and failure in practical applications.
Smart Images

Figure CN122597055A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of adversarial learning technology, and in particular to a self-evolving method and system for attacking and defending against financial black and gray market activities based on adversarial reinforcement learning. Background Technology
[0002] Traditional risk control systems often rely on business rules and machine learning models to monitor and score financial transactions and account behavior. Algorithms such as logistic regression, support vector machines, decision trees, and random forests are used to identify potential attack behaviors, while big data feature engineering is combined to improve model performance. With the development of artificial intelligence, Generative Adversarial Networks (GANs) have been introduced into the field of financial fraud detection to address the problems of imbalanced data and the scarcity of rare fraud samples. As research continues, reinforcement learning, as a machine learning method that can optimize decision-making strategies based on environmental interactions and feedback, is also gradually being introduced into the fields of financial risk control and fraud detection. Reinforcement learning models the risk control problem as a Markov decision process by defining a state space, action space, and reward function, enabling the agent to learn the optimal risk control strategy in a dynamic environment.
[0003] However, existing traditional risk control models and most existing reinforcement learning solutions are usually based on historical data or fixed training sets for strategy optimization. Their decision-making strategies remain relatively static after training, making it difficult to adapt to the dynamic changes and strategic evolution of black and gray market attack behaviors, which leads to the strategy being prone to lag and failure in practical applications. Summary of the Invention
[0004] To address the issues of lag and failure in existing methods, this application discloses the following technical solution:
[0005] The first aspect of this application provides a self-evolving method for attacking and defending against financial black and gray market activities based on adversarial reinforcement learning, including:
[0006] Based on the attack strategy, the intelligent agent processes the attack state of the target object and obtains the attack strategy for simulating financial black and gray industry attacks.
[0007] The defense decision-making agent processes the defense status of the target object to generate a defense strategy against the attack strategy.
[0008] The results of the countermeasures are evaluated based on the attack strategy and the defense strategy. If the results of the countermeasures meet the target defense conditions, the defense strategy is applied in the black and gray industry attack and defense risk control system to identify black and gray industry attacks.
[0009] The second aspect of this application provides a self-evolving offensive and defensive system for financial black market activities based on adversarial reinforcement learning, comprising:
[0010] The attack strategy generation layer is used to obtain attack strategies for simulating financial black and gray market attacks by processing the attack state of the target object based on the attack strategy agent.
[0011] The defense decision execution layer is used to generate a defense strategy against the attack strategy based on the defense status of the target object processed by the defense decision agent.
[0012] The attack and defense evolution control layer is used to evaluate the adversarial results based on the attack strategy and the defense strategy, and apply the defense strategy in the black and gray industry attack and defense risk control system to identify black and gray industry attacks when the adversarial results meet the target defense conditions.
[0013] The beneficial effects of this embodiment are as follows: the attack strategy agent automatically generates an attack strategy that simulates black and gray market attacks, and then the defense decision agent generates a corresponding defense strategy based on the attack strategy. Thus, this solution can dynamically update the defense strategy and the attack strategy used to train the defense decision agent through the confrontation between the attack strategy agent and the defense decision agent, so that the attack strategy and the defense strategy can be updated in real time in practical applications, solving the problem that the strategy is prone to lag and failure in practical applications. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0015] Figure 1 This is a flowchart of a self-evolving attack and defense method for financial black and gray industries based on adversarial reinforcement learning, provided in an embodiment of this application.
[0016] Figure 2 This is a schematic diagram of the architecture of a self-evolving attack and defense system for financial black and gray industries based on adversarial reinforcement learning, provided in an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] Existing traditional risk control models and most existing reinforcement learning solutions suffer from relatively static decision-making strategies after training, making it difficult to adapt to the dynamic changes and strategic evolution of black market attacks. This leads to the strategies becoming outdated and ineffective in practical applications. The following problems also exist.
[0019] The existing methods for generating adversarial examples for reinforcement learning (e.g., generating fraudulent examples based on GANs) mainly rely on expanding training data and addressing class imbalance issues. However, the generated examples lack continuous strategic steps and cannot fully cover unknown or novel fraudulent paths, resulting in limited system capabilities for identifying potential risks.
[0020] Strategy updates rely on human experience. Existing reinforcement learning risk control solutions are mostly single-agent strategy optimizations, lacking a closed-loop mechanism for adversarial game between multiple agents. Strategy updates still rely to some extent on human experience or static rules for calibration, making it difficult to achieve autonomous evolution and continuous optimization of risk control strategies.
[0021] To address the problems of static strategies, insufficient coverage of unknown risks, and reliance on manual experience for strategy updates in existing financial risk control systems when dealing with attacks from cybercriminals, this application proposes a self-evolving method and system for attacking and defending against cybercrime in the financial sector based on adversarial reinforcement learning. This application constructs a three-layer adversarial reinforcement learning architecture, specifically including: an attack strategy generation layer that simulates strategy-aware cybercriminal attack behaviors; a defense decision execution layer that learns the optimal risk management strategy in an adversarial environment; and an attack-defense evolution control layer that schedules and constrains the adversarial process.
[0022] Through the above three-layer architecture, an adversarial reinforcement learning mechanism is introduced. Without relying on a large number of real fraud samples for expansion, the risk control system can proactively discover potential risk blind spots and continuously optimize defense strategies in the process of attack and defense. This improves the financial risk control system's ability to identify and handle unknown black and gray market behaviors and its overall robustness.
[0023] Specifically, this application proposes a self-evolving method and system for attacking and defending against financial black market activities based on adversarial reinforcement learning. This system is based on real-world financial business risk control scenarios. By introducing a multi-layered adversarial agent collaborative mechanism, it integrates the generation of black market attack behaviors, the defensive decision-making of risk strategies, and the evolution and scheduling of the attack and defense process into a single reinforcement learning closed loop, thereby achieving continuous self-evolution of risk control strategies in an adversarial game environment. By constructing a three-layer adversarial structure—a red team attack strategy layer, a blue team defense decision layer, and an attack and defense evolution control layer—the risk control system can proactively expose potential risk blind spots without relying on the expansion of real fraud samples, and continuously improve its ability to identify and handle unknown black market activities.
[0024] The three-tier architecture used to implement the method of this application is described in detail below.
[0025] The first layer is the attack strategy generation layer, also known as the red team attack strategy layer. The attack strategy generation layer is used to simulate the strategic attack behavior of black and gray market operators in real business environments, and generates attack actions that can bypass the current risk control strategy through reinforcement learning, so as to drive the adversarial training and evolution of defense strategies.
[0026] The attack strategy generation layer includes at least one attack strategy agent (also known as a red team attack strategy agent). Each attack strategy agent can generate at least one attack action within the attack state space and attack action space.
[0027] The attack state space is used to characterize the environmental state of the current attack action (also known as the current attack decision). Specifically, the attack state space includes at least one or more of the following information: the risk assessment result of the current defense strategy (also known as the strategy output state); the transaction characteristics, account attributes, or behavioral sequence characteristics of the target; historical attack actions and their corresponding risk control handling results; and the stage identifier or risk accumulation state of the attack path.
[0028] The attack status reflects the state of the target being attacked, and its relevance to black and gray market attacks.
[0029] Current attack action refers to the attack action currently generated by the attack strategy agent, while historical attack action refers to the attack action that the attack strategy agent has generated in the past.
[0030] The current defense strategy (also known as the current defense action) refers to the defense strategy (or action) generated by the agent in the defense decision execution layer (also known as the blue team defense decision execution layer) to prevent historical attack actions in response to the latest attack action generated by the attack strategy agent. The method for determining the risk assessment result is described in the section on the defense decision execution layer. The target of attack can be understood as one or more specified systems used for attack and defense testing (e.g., an online trading platform). Transaction characteristics and behavioral sequence characteristics respectively characterize the system's transaction records and system operation behavior in a recent period. Account attributes include, but are not limited to, one or more of the following: account level, account holder's personal attributes, account purpose, account opening time, etc., of all or some of the accounts registered on this system.
[0031] Historical attack actions refer to attack actions generated by the attack strategy agent within a recent period. The risk control handling results corresponding to historical attack actions are determined by the defense decision execution layer based on historical attack actions. For specific determination methods, please refer to the relevant instructions of the defense decision execution layer.
[0032] An attack path is determined based on specific attack actions; it is a set of specific operational instructions taken to carry out attacks in the black and gray market. Each attack action corresponds to an attack path, and the attack path is determined by its corresponding attack action. For example, an attack path might include: stealing account information via instruction 1, forging transaction records via instruction 2, and calling the account holder via instruction 3 to commit fraud.
[0033] The stage identifier of the attack path is used to identify which business stages this attack path targets. Business stages include, but are not limited to, any one or more of the following: transfer, payment, deposit, withdrawal, and wealth management.
[0034] The risk accumulation state of an attack path can be understood as the anticipated level of harm that may be caused by carrying out black and gray market attacks along this attack path.
[0035] In a specific implementation, the attack strategy agent's attack state SR in the attack state space at any given time is... t It can be represented as: SR t ={x t PaiB t (x) t ), h t}. Where, x t This represents the transaction characteristics and behavioral sequence characteristics of the target object; PaiB t (x) t This includes the current defense strategy generated by the defense decision execution layer and the corresponding risk assessment results; h t This includes any one or more of the following: historical attack actions, stage markers of the attack path, and the risk accumulation status of the attack path.
[0036] The state within the attack state space can be denoted as the attack state.
[0037] The action space of the attack strategy agent is defined as follows.
[0038] The action space of the attack strategy agent describes the attack behaviors it can take. Unlike traditional methods that directly generate complete fraud samples, the actions of the attack strategy agent in this application are defined as strategic adjustments to attack paths or behavioral characteristics to gradually approach the target of bypassing risk control.
[0039] In this application, the attack action generated by an attack strategy agent at any time (e.g., time t) can be denoted as AR. t AR t It can be any of the multiple optional attack actions contained in the attack action space {AR}. The attack action space can include any one or more of the following optional attack actions:
[0040] Disturbances to any one or more of the numerical characteristics such as transaction amount, frequency, and time interval;
[0041] Adjustments to the order of account actions, call paths, or interaction modes;
[0042] Gradual weakening or diversification of risk-sensitive characteristics;
[0043] Operations to advance or retreat during the attack phase.
[0044] During the operation of the attack strategy agent, the attack actions (AR) generated by it at any given moment are... t The attack state SR used at this moment t To form the attack state SR in the next moment. t+1 .
[0045] Regarding the design of the reward function for the attack strategy agent, the reward function is used to guide it to generate more covert and bypassable attack actions. Its core objective is to achieve a breakthrough against the defense strategy with minimal risk exposure. The reward function of the attack strategy agent is defined as: Q(AR) t =C1*I bypass (AR) t -C2*C exposed (AR) t -C3*C cost (AR) t ).
[0046] Among them, I bypass (AR) t () indicates the attack action AR t Whether the current defense strategy was successfully bypassed is equivalent to the first component; C exposed (AR) t () indicates the attack action AR t The resulting risk exposure or degree of abnormality corresponds to the second component; C cost (AR) t ) represents the deviation cost of the attack action relative to normal business behavior, which is equivalent to the third component; C1, C2 and C3 are preset weight parameters used to balance the success rate and stealth of the attack. The values of these parameters can be preset according to the actual situation and experience, without limitation.
[0047] Through the above reward design, the attack strategy agent can not only pursue the result of bypassing the defense when generating attack actions, but also be constrained by the stealth of the attack and the rationality of the behavior, thereby generating attack actions that are closer to the real black and gray industry behavior patterns.
[0048] The second layer is the defense decision execution layer, also known as the blue team defense decision execution layer. It is used to conduct risk assessment and handling decisions on transaction requests, account behaviors, or business events during the operation of financial business. It is the core execution layer responsible for actual risk prevention and control in this application. This layer includes at least one defense decision agent. Under the adversarial reinforcement learning framework, the defense decision agent models the handling behavior of the risk control system as a sequential decision problem, enabling the defense strategy to be continuously optimized in a continuous offensive and defensive game environment, thereby improving the ability to identify and respond to complex, covert, and unknown black and gray market activities.
[0049] The defense decision agent can generate corresponding defense actions to prevent an attack action generated by the attack strategy agent, within the defense state space and defense action space.
[0050] In constructing the defense state space, the defense state space of the defense decision-making agent is used to characterize the comprehensive risk situation of the object to be decided. The state of the defense decision-making agent at any given time (called the defense state) can be denoted as SD. t This defensive state can be represented as SD. t ={sx t r t sh t sc t}
[0051] sx t This refers to business characteristic information representing the current transaction, account, or behavior request, including but not limited to amount characteristics, time characteristics, behavior sequence characteristics, and contextual characteristics; t This indicates the basic risk assessment results output by existing risk control models or rule systems; sh t This indicates historical decision-making and feedback information, including the historical handling results of similar behaviors, instances of misjudgment, and feedback on the risk of delays; sc t It indicates information on business constraints and compliance constraints, including regulatory rules, business strategy boundaries, and risk tolerance settings.
[0052] The aforementioned targets for decision-making refer to accounts that have been attacked and require defensive actions, such as those that are equivalent to the targets for attack mentioned above.
[0053] The defense status reflects the state of the target object and the defense strategy to prevent attacks from black and gray industries.
[0054] By constructing the aforementioned defense state space, the defense decision-making agent can simultaneously perceive business risks, historical experience, and compliance constraints during the decision-making process of generating defense actions, thus avoiding strategy imbalances caused by a single feature.
[0055] In defining the defensive action space of a defensive decision-making agent, the defensive action space describes the possible responses of a risk control system to risk events. The defensive action space can be represented by {AD}, which contains multiple optional defensive actions, including but not limited to:
[0056] Directly block or reject transaction or business requests;
[0057] Allow the transaction or business request;
[0058] Trigger enhanced verification or two-factor authentication;
[0059] Adjust risk thresholds or risk control strategy parameters;
[0060] The request will be routed to a manual review or post-review process.
[0061] By discretizing or parameterizing risk control actions into a defensive action space, the defensive decision-making agent can learn the optimal defensive actions under different risk situations, rather than relying on fixed rules or static thresholds.
[0062] The defensive action generated by the defensive decision-making agent at any given time (e.g., time t) can be denoted as AD. t This defensive action can be any of the optional defensive actions contained in the aforementioned defensive action space.
[0063] In designing the reward function for the defense decision-making agent, this application constructs a reward function based on multi-objective weighting to guide the defense decision-making agent to achieve a balance between security, business availability and compliance.
[0064] Specifically, the reward function Q(AD) for defensive actions t ) can be represented as: Q(AD) t =-B1*L risk (AD) t -B2*L false (AD) t -B3*L cost (AD) t )+B4*G business (AD) t ).
[0065] Among them, L risk (AD) t This represents the risk loss caused by defensive decision-making errors, equivalent to the fourth component; L false (AD) t ) represents the negative business effects caused by mistaken interception or mistaken release, equivalent to the fifth component; L cost (AD) tThis represents the operational costs associated with manual review and verification processes, equivalent to the sixth component; G business (AD) t The first component represents the positive return from successful business completion under compliance conditions (i.e., meeting compliance requirements), equivalent to the seventh component. B1, B2, B3, and B4 are configurable weight parameters used to reflect the emphasis on risk control and business efficiency at different business stages. Their specific values can be preset based on actual conditions and relevant experience, without limitation.
[0066] The third is the attack and defense evolution control layer. The attack and defense evolution control layer is used to uniformly schedule, constrain, and manage the adversarial process between the attack strategy generation layer and the defense decision execution layer. It is the core control layer for achieving long-term stable operation and continuous optimization of the risk control system.
[0067] The attack and defense evolution control layer includes at least one attack and defense evolution control agent. The attack and defense evolution control agent is used to monitor and dynamically adjust the training process of the attack strategy agent and the defense decision agent in a global manner, so as to avoid problems such as attack strategies deviating from real business scenarios, defense strategies overfitting adversarial samples, or loss of control of system risk boundaries during adversarial training.
[0068] Specifically, the attack-defense evolution control agent can dynamically control the training intensity, adversarial frequency, and strategy update rhythm of the attack strategy generation layer and the defense decision execution layer based on the strategy convergence state, risk exposure level, and changes in business indicators generated during the adversarial process. When the complexity or concealment of the attack strategy exceeds the preset reasonable business range, the adversarial weight of the attack action space is automatically reduced; when the performance of the defense strategy against known attack patterns is detected to be degraded, a strategy backtracking or retraining mechanism is triggered, thereby ensuring that both the attacker and defender continue to evolve within a controlled range.
[0069] Among them, the adversarial weights of the attack action space refer to the weight parameters C1, C2 and C3 used in the aforementioned embodiments to balance the attack success rate and stealth.
[0070] Furthermore, the attack-defense evolution control layer is also used to uniformly manage the attack strategies, defense strategies, and corresponding adversarial results formed during each attack-defense confrontation, constructing an attack-defense strategy evolution library. This library records the attack strategies, defense strategies, risk characteristic distribution, and adversarial results of different adversarial processes, providing a basis for subsequent adversarial training, strategy verification, and audit backtracking. It also serves as an important reference constraint when updating defense strategies to prevent performance regression or risk amplification during strategy updates.
[0071] In this process, the attack strategy agent generates an attack strategy, the defense decision agent generates a defense strategy in response to the attack strategy, and then the attack strategy and the defense strategy are applied in the simulation environment to obtain the adversarial results and risk feature distribution. The process until the adversarial results and risk feature distribution are obtained is called an adversarial process (also known as an attack-defense adversarial process or an attack-defense process).
[0072] The outcome of the countermeasures characterizes the extent to which the application of this defense strategy can prevent attacks launched based on this attack strategy. Specifically, the defense success rate can be used as the outcome of the countermeasures. The defense success rate represents the proportion of all attacks launched based on this attack strategy that are intercepted by the black and gray industry attack and defense risk control system after the application of this defense strategy. The risk characteristic distribution characterizes what vulnerabilities may still exist after the application of this defense strategy and how great the risk of these vulnerabilities being exploited is. Specifically, this can be obtained by scanning the vulnerabilities of online trading platforms monitored by the black and gray industry attack and defense risk control system after the application of this defense strategy.
[0073] Through the aforementioned attack and defense evolution control mechanism, this application can achieve continuous game and co-evolution of attack and defense strategies while ensuring the security and business stability of the risk control system, enabling the risk control system to adapt to the long-term evolution of black and gray market behaviors.
[0074] The beneficial effects of the method in this application are as follows:
[0075] First, it significantly enhances the ability to identify unknown black and gray market attacks. Through a strategy-aware attack generation mechanism based on the defense strategy state in the attack strategy generation layer, the system can proactively construct attack paths that bypass existing risk control strategies, effectively exposing potential risk blind spots, thereby improving its coverage of unknown fraud patterns;
[0076] Second, it enables the continuous self-evolution of risk control strategies in adversarial environments. The defense decision execution layer continuously adjusts risk handling strategies through ongoing adversarial training with attack strategies, so that risk control decisions no longer rely on static rules or single model training, but achieve dynamic optimization driven by real adversarial feedback;
[0077] Third, ensure the stability and operational controllability of the adversarial training process. Through the unified scheduling of adversarial intensity, strategy update frequency, and training rhythm by the attack and defense evolution control layer, it avoids attack strategies deviating from the actual business distribution or defense strategies overfitting to adversarial samples, ensuring that the system operates stably within the security boundary;
[0078] Fourth, it promotes the accumulation and reuse of offensive and defensive strategy knowledge. This application constructs an offensive and defensive strategy evolution library to systematically manage the attack paths, defense strategies, and adversarial results generated during the offensive and defensive process, providing reusable strategy assets for subsequent risk control strategy optimization and model training;
[0079] Fifth, reduce the cost of manual intervention and improve the overall response efficiency of the risk control system. By embedding the discovery, verification, and optimization of risk strategies into an adversarial reinforcement learning closed loop, the reliance on manual rule revision and post-event analysis is reduced, thereby improving the risk control system's response speed to attack changes and its overall operational efficiency.
[0080] In some embodiments, an adversarial reinforcement learning-based black market attack and defense risk control system, based on the above-mentioned attack strategy generation layer, defense decision execution layer, and attack and defense evolution control layer, can be applied to an online trading platform. System deployment includes:
[0081] Attack strategy agent: Deployed in the mining environment, isolated from the actual transaction database, to simulate potential black and gray market attack behaviors;
[0082] Defense decision-making intelligent agent: It interfaces with the online transaction monitoring system and is responsible for real-time risk scoring and interception;
[0083] Offensive and defensive evolution control agent: responsible for scheduling the adversarial training cycle between the attack strategy agent and the defense decision agent, controlling the training intensity, and maintaining the strategy evolution library.
[0084] During the one-month gray-scale testing period, the black-market attack and defense risk control system scored and blocked online transactions in real time, with the main effects including:
[0085] The attack simulation samples cover over 95% of novel fraud paths, representing an approximately 20% improvement in coverage compared to traditional static rule systems.
[0086] The defensive decision-making agent reduces the false alarm rate of the defense strategy generated by the defense strategy by 10% and increases the interception success rate by 5%.
[0087] The closed-loop offensive and defensive evolution ensures continuous updates to system strategies without frequent manual intervention.
[0088] The risk control strategy has a self-evolution cycle of approximately one day, and is automatically updated after daily manual review.
[0089] As can be seen, the method provided in this application has advantages such as strategy-aware red team attack sample generation, blue team defense strategy self-evolution, and attack-defense evolution closed-loop mechanism. At the same time, it can quantitatively demonstrate the effect of being superior to the traditional static risk control system.
[0090] This application also provides a self-evolving method for attacking and defending against financial black market activities based on adversarial reinforcement learning. Please refer to [link to relevant documentation]. Figure 1 Here is a flowchart of the method, which may include the following steps.
[0091] S101, Based on the attack strategy, the intelligent agent processes the attack state of the target to be attacked and obtains an attack strategy for simulating financial black and gray industry attacks.
[0092] S102, Based on the defense decision-making agent's processing of the defense state of the target object, a defense strategy is generated to counter the attack strategy.
[0093] S103 evaluates the results of the confrontation based on the attack strategy and the defense strategy, and applies the defense strategy in the black and gray industry attack and defense risk control system to identify black and gray industry attacks when the confrontation results meet the target defense conditions.
[0094] The outcome of the confrontation can be, for example, the success rate of the defense. The outcome of the confrontation can meet the target conditions, such as the success rate of the defense being greater than a preset success rate threshold.
[0095] The methods for generating attack strategies and obtaining attack status in step S101 can be found in the description of the attack strategy generation layer in the foregoing embodiments. The methods for generating defense strategies and obtaining defense status in step S102 can be found in the description of the defense decision execution layer in the foregoing embodiments.
[0096] It should be noted that if the result of this confrontation does not meet the target defense conditions, steps S101 and S102 can be repeated, so that the attack strategy agent and the defense decision agent alternately carry out the confrontation process multiple times until the confrontation result meets the target defense conditions, and then step S103 is executed.
[0097] The beneficial effects of this embodiment are as follows: the attack strategy agent automatically generates an attack strategy that simulates black and gray market attacks, and then the defense decision agent generates a corresponding defense strategy based on the attack strategy. Thus, this solution can dynamically update the defense strategy and the attack strategy used to train the defense decision agent through the confrontation between the attack strategy agent and the defense decision agent, so that the attack strategy and the defense strategy can be updated in real time in practical applications, solving the problem that the strategy is prone to lag and failure in practical applications.
[0098] In some optional embodiments, if the adversarial results meet the target defense conditions, the defense decision agent can also be applied to the black and gray industry attack and defense risk control system, and the defense decision agent can monitor in real time whether the target to be attacked is attacked by black and gray industries.
[0099] Optionally, based on the attack strategy, the intelligent agent processes the attack state of the target to obtain an attack strategy for simulating attacks on financial black and gray industries, including:
[0100] Based on the attack strategy, the intelligent agent processes the attack state of the target to be attacked and obtains an attack strategy for simulating financial black and gray industry attacks, and the corresponding attack reward function satisfies the first condition.
[0101] The attack reward function is formed by weighting and fusing a first component, which indicates whether the attack strategy successfully bypasses the current defense strategy, a second component, which indicates the degree of risk exposure or anomaly caused by the attack strategy, and a third component, which indicates the deviation cost of the attack action relative to normal business behavior, according to a preset adversarial weight.
[0102] The attack reward function corresponding to the above embodiments can be, for example, the aforementioned Q(AR) function. t The attack reward function satisfies the first condition, which means that the attack reward function is greater than or equal to the preset attack reward threshold.
[0103] Optionally, based on the defense decision-making agent's processing of the defense state of the target object, a defense strategy is generated to counter the attack strategy, including:
[0104] The defense decision-making agent processes the defense state of the target object to generate a defense strategy against the attack, and the corresponding defense reward function satisfies the second condition.
[0105] The defense reward strategy is formed by integrating the fourth component, which represents the risk loss caused by defense decision errors; the sixth component, which represents the negative business effects caused by false interception or false release; the seventh component, which represents the operating costs caused by manual review and verification processes; and the seventh component, which represents the positive benefits brought by the successful completion of business under compliance conditions.
[0106] The defense reward function corresponding to the above embodiments can be, for example, the aforementioned Q(AD) function. t The defense reward function satisfies the second condition, which means that the defense reward function is greater than or equal to the preset defense reward threshold.
[0107] Optional, also includes:
[0108] Detect the complexity and stealth of attack strategies;
[0109] If either the complexity or the stealth of the attack strategy exceeds the preset reasonable range for business operations, the first, second, and third components will be reduced.
[0110] The preset reasonable range for business operations may include, but is not limited to, the preset range for complexity and the preset range for concealment.
[0111] The complexity and stealth of an attack strategy can be determined by the attack-defense evolution control agent calling a pre-trained neural network model to evaluate the attack strategy. For specific evaluation methods, please refer to relevant technologies in the field of artificial intelligence, which will not be elaborated here.
[0112] Optional, also includes:
[0113] Attack strategies, defense strategies, adversarial outcomes, and risk characteristic distributions determined based on defense strategies are stored in an attack and defense strategy evolution library to generate new attack strategies and new defense strategies based on the library.
[0114] The method for obtaining the outcome of the confrontation and the distribution of risk characteristics is described in the aforementioned embodiment regarding the attack and defense evolution control layer, and will not be repeated here.
[0115] This application provides a self-evolving attack and defense system for financial black market activities based on adversarial reinforcement learning. Please refer to [link to relevant documentation]. Figure 2 This is a schematic diagram of the system architecture, which can include the following three-layer architecture.
[0116] Attack strategy generation layer 201 is used to obtain an attack strategy for simulating financial black and gray industry attacks by processing the attack state of the target object according to the attack strategy agent.
[0117] The defense decision execution layer 202 is used to generate a defense strategy against the attack strategy based on the defense state of the target object processed by the defense decision agent.
[0118] The attack and defense evolution control layer 203 is used to evaluate the adversarial results based on attack and defense strategies, and apply defense strategies in the black and gray industry attack and defense risk control system to identify black and gray industry attacks when the adversarial results meet the target defense conditions.
[0119] Optionally, when the attack strategy generation layer 201 obtains an attack strategy for simulating financial black market attacks based on the attack strategy agent's processing of the attack state of the target, it is specifically used for:
[0120] Based on the attack strategy, the intelligent agent processes the attack state of the target to be attacked and obtains an attack strategy for simulating financial black and gray industry attacks, and the corresponding attack reward function satisfies the first condition.
[0121] The attack reward function is formed by weighting and fusing a first component, which indicates whether the attack strategy successfully bypasses the current defense strategy, a second component, which indicates the degree of risk exposure or anomaly caused by the attack strategy, and a third component, which indicates the deviation cost of the attack action relative to normal business behavior, according to a preset adversarial weight.
[0122] Optionally, when the defense decision execution layer 202 generates a defense strategy against the attack strategy based on the defense state of the target object processed by the defense decision agent, it is specifically used for:
[0123] The defense decision-making agent processes the defense state of the target object to generate a defense strategy against the attack, and the corresponding defense reward function satisfies the second condition.
[0124] The defense reward strategy is formed by integrating the fourth component, which represents the risk loss caused by defense decision errors; the sixth component, which represents the negative business effects caused by false interception or false release; the seventh component, which represents the operating costs caused by manual review and verification processes; and the seventh component, which represents the positive benefits brought by the successful completion of business under compliance conditions.
[0125] Optionally, the attack and defense evolution control layer 203 is also used for:
[0126] Detect the complexity and stealth of attack strategies;
[0127] If either the complexity or the stealth of the attack strategy exceeds the preset reasonable range for business operations, the first, second, and third components will be reduced.
[0128] Optionally, the attack and defense evolution control layer 203 is also used for:
[0129] Attack strategies, defense strategies, adversarial outcomes, and risk characteristic distributions determined based on defense strategies are stored in an attack and defense strategy evolution library to generate new attack strategies and new defense strategies based on the library.
[0130] The specific working principle of the financial black and gray market attack and defense self-evolution system based on adversarial reinforcement learning in this embodiment can be found in the relevant steps of the financial black and gray market attack and defense self-evolution method based on adversarial reinforcement learning in the foregoing embodiments, and will not be repeated here.
[0131] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0132] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.
[0133] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0134] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0135] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A self-evolving method for attacking and defending against financial black and gray market activities based on adversarial reinforcement learning, characterized in that, include: Based on the attack strategy, the intelligent agent processes the attack state of the target object and obtains the attack strategy for simulating financial black and gray industry attacks. The defense decision-making agent processes the defense status of the target object to generate a defense strategy against the attack strategy. The results of the countermeasures are evaluated based on the attack strategy and the defense strategy. If the results of the countermeasures meet the target defense conditions, the defense strategy is applied in the black and gray industry attack and defense risk control system to identify black and gray industry attacks.
2. The method according to claim 1, characterized in that, The process of the attack agent processing the attack state of the target object according to the attack strategy to obtain an attack strategy for simulating financial black and gray market attacks includes: Based on the attack strategy, the intelligent agent processes the attack state of the target to be attacked and obtains an attack strategy for simulating financial black and gray industry attacks, and the corresponding attack reward function satisfies the first condition. The attack reward function is formed by weighting and fusing a first component, which indicates whether the attack strategy successfully bypasses the current defense strategy, a second component, which indicates the degree of risk exposure or anomaly caused by the attack strategy, and a third component, which indicates the deviation cost of the attack action relative to normal business behavior, according to a preset adversarial weight.
3. The method according to claim 1, characterized in that, The step of processing the defense state of the target object according to the defense decision-making agent to generate a defense strategy against the attack strategy includes: The defense decision agent processes the defense state of the target object to generate a defense strategy against the attack strategy, and the corresponding defense reward function satisfies the second condition. The defense reward strategy is formed by integrating the fourth component, which represents the risk loss caused by defense decision errors; the sixth component, which represents the negative business effects caused by false interception or false release; the seventh component, which represents the operating costs caused by manual review and verification processes; and the seventh component, which represents the positive benefits brought by the successful completion of business under compliance conditions.
4. The method according to claim 2, characterized in that, Also includes: The complexity and stealth of the attack strategy were detected. If either the complexity or the stealth of the attack strategy exceeds a preset reasonable range for business operations, the first component, the second component, and the third component shall be reduced.
5. The method according to claim 1, characterized in that, Also includes: The attack strategy, the defense strategy, the confrontation result, and the risk characteristic distribution determined based on the defense strategy are stored in the attack and defense strategy evolution library to generate new attack strategies and new defense strategies according to the attack and defense strategy evolution library.
6. A self-evolving offensive and defensive system for financial black market activities based on adversarial reinforcement learning, characterized in that: include: The attack strategy generation layer is used to obtain attack strategies for simulating financial black and gray market attacks by processing the attack state of the target object based on the attack strategy agent. The defense decision execution layer is used to generate a defense strategy against the attack strategy based on the defense status of the target object processed by the defense decision agent. The attack and defense evolution control layer is used to evaluate the adversarial results based on the attack strategy and the defense strategy, and apply the defense strategy in the black and gray industry attack and defense risk control system to identify black and gray industry attacks when the adversarial results meet the target defense conditions.
7. The system according to claim 6, characterized in that, When the attack strategy generation layer obtains an attack strategy for simulating financial black market attacks based on the attack strategy agent's processing of the attack state of the target object, it is specifically used for: Based on the attack strategy, the intelligent agent processes the attack state of the target to be attacked and obtains an attack strategy for simulating financial black and gray industry attacks, and the corresponding attack reward function satisfies the first condition. The attack reward function is formed by weighting and fusing a first component, which indicates whether the attack strategy successfully bypasses the current defense strategy, a second component, which indicates the degree of risk exposure or anomaly caused by the attack strategy, and a third component, which indicates the deviation cost of the attack action relative to normal business behavior, according to a preset adversarial weight.
8. The system according to claim 6, characterized in that, When the defense decision execution layer processes the defense state of the target object based on the defense decision agent to generate a defense strategy against the attack strategy, it is specifically used for: The defense decision agent processes the defense state of the target object to generate a defense strategy against the attack strategy, and the corresponding defense reward function satisfies the second condition. The defense reward strategy is formed by integrating the fourth component, which represents the risk loss caused by defense decision errors; the sixth component, which represents the negative business effects caused by false interception or false release; the seventh component, which represents the operating costs caused by manual review and verification processes; and the seventh component, which represents the positive benefits brought by the successful completion of business under compliance conditions.
9. The system according to claim 7, characterized in that, The attack and defense evolution control layer is also used for: The complexity and stealth of the attack strategy were detected. If either the complexity or the stealth of the attack strategy exceeds a preset reasonable range for business operations, the first component, the second component, and the third component shall be reduced.
10. The system according to claim 6, characterized in that, The attack and defense evolution control layer is also used for: The attack strategy, the defense strategy, the confrontation result, and the risk characteristic distribution determined based on the defense strategy are stored in the attack and defense strategy evolution library to generate new attack strategies and new defense strategies according to the attack and defense strategy evolution library.