Confrontation strategy evolution method based on evolution
By training multiple RL learners in parallel to generate multiple strategy execution bodies, combining task behavior trees and causal-driven data sets, and dynamically adjusting the strategies, solving the problems of insufficient diversity, poor adaptability and incomplete evaluation in the existing technology, and achieving flexibility, diversity and high adaptability of the adversarial strategy.
Patent Information
- Application Number
- CN202510187571.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing evolution methods of adversarial strategies lack diversity, insufficient adaptability, imperfect evaluation mechanisms, and it is difficult to effectively deal with complex and dynamic adversarial environments.
By creating a parallel adversarial game training environment, multiple reinforcement learning (RL) learners are used to generate multiple versions of RL strategy execution bodies, record and analyze performance data, generate task behavior trees and causal-driven data sets, and dynamically adjust the strategies to achieve evolutionary optimization of adversarial capabilities.
The flexibility and diversity of strategies are achieved, the adaptability and accuracy of adversarial strategies are improved, and the strategies can be adapted in real time and maintained stable cross-situation adaptability in complex and changeable adversarial scenarios.
Smart Images

Figure CN120124708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and machine learning, and specifically to an evolution method for adversarial strategies based on evolution. Background Art
[0002] In the fields of artificial intelligence and machine learning, especially in reinforcement learning (RL), evolution methods for adversarial strategies are widely applied to various complex tasks, such as games, robot control, and network security. The core of these methods lies in simulating an adversarial environment to train an agent to optimize its decision-making strategy, thereby improving its performance in a dynamic and uncertain environment.
[0003] Existing evolution methods for adversarial strategies usually rely on a single strategy executor for training. These methods continuously interact with the environment, collect feedback, and gradually improve the strategy. The agent will face different opponents and environmental changes during the training process, and the goal is to maximize its cumulative reward by learning the optimal strategy. During this process, the agent can gradually adapt to the environment and improve its adversarial ability.
[0004] However, there are some obvious defects in the existing technologies:
[0005] Lack of diversity: Most methods rely on a fixed strategy executor, resulting in a lack of diversity in the strategy and being unable to effectively cope with complex and dynamic adversarial environments.
[0006] Insufficient adaptability: When facing rapidly changing adversarial situations, the adaptability of the agent is limited, and it is difficult to update and optimize the strategy in real time.
[0007] Imperfect evaluation mechanism: Traditional evaluation methods often cannot accurately quantify the effectiveness of the strategy, resulting in a lack of effective basis for strategy selection and possibly leading to the selection of suboptimal strategies. Summary of the Invention
[0008] In view of the deficiencies of the existing technologies, the present invention provides an evolution method for adversarial strategies based on evolution, which solves the problems mentioned in the above background art.
[0009] To achieve the above objectives, the present invention is realized through the following technical solutions: An evolution method for adversarial strategies based on evolution, comprising the following steps:
[0010] Create a parallel adversarial game training environment, and use multiple reinforcement learning (RL) learners to generate multiple versions of RL strategy executors to cope with different strategy adversarial situations;
[0011] Record and analyze the performance data of each RL strategy executor, and generate a task behavior tree for policy decision-making based on the performance data;
[0012] Construct a causality-driven dataset based on the causality features in the task behavior tree to guide the policy optimization and evolution of the RL policy executor;
[0013] Dynamically adjust the adversarial policy of the RL policy executor using the causality-driven dataset, and achieve the evolutionary optimization of the policy adversarial ability through continuous iterative training.
[0014] Preferably, the generation process of the RL policy executor includes: through a parallel adversarial game training mode, enabling each RL learner to generate policy versions with specific behavioral characteristics, and the specific behavioral characteristics cover different policy advantages and policy weaknesses.
[0015] Preferably, the generation of the task behavior tree includes the following steps:
[0016] Establish behavior tree nodes based on policy adversarial experiment data, and each node contains the key behaviors of the RL policy executor during the adversarial process;
[0017] Analyze the causal links of the nodes in the behavior tree to generate multiple policy decision paths, and the policy decision paths are used to determine the optimal policy execution sequence in different adversarial situations.
[0018] Preferably, the generation of the causality-driven dataset includes:
[0019] Extract the causal relationship feature data of the RL policy executor in multiple situations from the policy adversarial experiment data, and the causal relationship feature data contains the association information of the task antecedents, behavior nodes, and policy results;
[0020] Construct a causality-driven dataset applicable to dynamic situations based on the extracted causal relationship features to enhance the policy adjustment ability of the RL policy executor in different adversarial situations.
[0021] Preferably, the causality-driven dataset is used to establish a policy feedback mechanism during the optimization process of the RL policy executor, specifically including:
[0022] Use the causality-driven dataset combined with the policy feedback results to adjust the decision rules of the RL policy executor in real time;
[0023] Through the feedback mechanism, gradually optimize the performance of the RL policy executor in different game adversarial situations.
[0024] Preferably, the optimization of the feedback mechanism includes the following steps:
[0025] Dynamically adjust the priority of the causal relationship feature data according to the adversarial performance of the RL policy executor in different situations;
[0026] For a specific version of the RL policy executor, by dynamically adjusting the priorities of causal data features, more targeted policy optimization is achieved.
[0027] Preferably, the calculation formula for policy evaluation and optimization of the RL policy executor is:
[0028] Let the performance index of the RL policy executor in a specific game scenario be P i , the weight of the causality-driven data feature be w i , and the expression of the comprehensive evaluation score S of the adversarial strategy be:
[0029]
[0030] where P i is the specific performance index of the policy executor in the adversarial game, and w i is the causality feature weight in this scenario.
[0031] Preferably, through the adversarial experiments of multiple versions of the RL policy executor, a task scenario-driven policy rule framework is generated, and the framework includes:
[0032] Based on the causality relationship feature data and policy adversarial performance, a policy rule library in typical task scenarios is established;
[0033] Using the policy rule library, adaptative policy rules are automatically generated in different task scenarios.
[0034] Preferably, the generation method of the policy rule library includes:
[0035] Based on the performance of the RL policy executor in typical task scenarios, high-frequency causal chain nodes are extracted from the task behavior tree;
[0036] The high-frequency causal chain nodes are constructed as the basic units of the policy rule library, and the execution effect of the adversarial strategy is optimized according to the following formula:
[0037]
[0038] where R represents the average effectiveness score of the policy rule, S j is the comprehensive evaluation score of the policy in the j-th adversarial experiment, and T is the number of experiments.
[0039] The present invention provides an evolutionary method for adversarial strategies based on evolution. It has the following beneficial effects:
[0040] 1. Through parallel training, the present invention uses multiple RL learners to generate different versions of policy executors, dynamically updates the policy based on the performance data in the adversarial game, and then forms a task behavior tree and a causality-driven dataset. Since each RL policy executor has specific behavior characteristics in different adversarial scenarios, it can adapt to different policy combinations in real time in complex and changeable adversarial scenarios, ensuring the flexibility and diversity of policy evolution. Under the action of the causality-driven data feedback mechanism, it can dynamically optimize the policy rules according to the changes in the adversarial scenario, realizing the intelligent evolution of adversarial strategies, and making the policy executor more adaptable in diverse scenarios.
[0041] 2. The present invention uses the calculation methods of comprehensive evaluation score and average effectiveness score to evaluate the performance of different policy executors in the adversarial game in real time, effectively improving the accuracy of the policy optimization process. The comprehensive evaluation score is calculated by weighting the causal feature weights and performance data, and can accurately quantify the performance of different policy versions; while the average effectiveness score further evaluates the stability of the policy rule base, helping to optimize policy selection and adjustment, so that the present invention can identify and retain efficient policy paths during the policy evolution process, ensuring that the adversarial strategy is not only effective in a single scenario, but also has stable cross-scenario adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment:
[0045] Please refer to the attached Figure 1 , the embodiment of the present invention provides an adversarial strategy evolution method based on evolution, including the following steps:
[0046] Create a parallel adversarial game training environment, and use multiple reinforcement learning (RL) learners to generate multiple versions of RL policy executors to cope with different policy adversarial scenarios;
[0047] Record and analyze the performance data of each RL policy executor, and generate a task behavior tree for policy decision-making based on the performance data;
[0048] Construct a causality-driven dataset based on the causal features in the task behavior tree to guide the policy optimization and evolution of the RL policy executor;
[0049] Dynamically adjust the adversarial strategy of the RL policy executor using the causality-driven dataset, and through continuous iterative training, achieve the evolutionary optimization of the policy adversarial ability.
[0050] Specifically, through the combination of reinforcement learning (RL) and causal analysis, continuously optimize and dynamically evolve the policy in a multi-agent adversarial environment. The specific method includes first deploying multiple RL learners in a parallel adversarial game training environment. Each learner generates different versions of the RL policy executor during the game process to effectively handle changing adversarial situations and complex policy requirements. On this basis, further record and analyze the performance data of each policy executor, and generate a task behavior tree based on this. The task behavior tree not only helps identify key policy decision points but also enables in-depth analysis of policy performance through causal relationships. Subsequently, construct a causality-driven dataset based on these causal features to provide accurate data support for policy optimization. Using the causality-driven dataset, the adversarial strategy of the RL policy executor can be dynamically adjusted, enabling it to gradually optimize policy selection in multi-round adversarial situations, thereby achieving the continuous evolution and improvement of the policy adversarial ability. The advantage of this method is that it does not rely excessively on high-computing-power devices. Through an effective data extraction and iterative optimization mechanism, it can be efficiently executed in resource-constrained environments and adapt to changing adversarial scenarios, providing a more robust and efficient adversarial policy update scheme for the intelligent agent.
[0051] The generation process of the RL policy executor includes: through a parallel adversarial game training mode, enable each RL learner to generate policy versions with specific behavioral characteristics, and the specific behavioral characteristics cover different policy advantages and policy weaknesses to improve the diversity and adaptability of the overall adversarial policy.
[0052] Specifically, through a parallel adversarial game training mode, prompt multiple RL learners to generate policy versions with specific behavioral characteristics in different adversarial situations. Each policy version exhibits unique policy characteristics when dealing with different situations, thus forming a rich policy library that contains various policy advantages and policy weaknesses. Through this multi-version policy generation method, dynamic interaction between policies can be achieved during the game process, enabling the evolutionary method to not only quickly adapt to a changing adversarial environment but also more precisely match corresponding adversarial situations based on the different characteristics of each policy version, enhancing the diversity and adaptability of the overall adversarial policy.
[0053] The generation of the task behavior tree includes the following steps:
[0054] Build behavior tree nodes based on policy adversarial experiment data, where each node contains the key behaviors of the RL policy executor during the adversarial process;
[0055] Analyze the causal links between nodes in the behavior tree to generate multiple policy decision paths, which are used to determine the optimal policy execution sequences in different adversarial scenarios.
[0056] Specifically, first construct behavior tree nodes based on policy adversarial experiment data. Each node records the key behavior performances of the RL policy executor in different adversarial scenarios, so as to ensure that the behavior tree can reflect the core decision-making points and behavior patterns of the policy executor during the confrontation. On this basis, conduct causal link analysis on each node in the behavior tree. By mining the causal relationships between nodes, generate multiple policy decision paths. These paths help to identify and optimize the policy combinations that perform well in different adversarial scenarios, enabling the evolution method to quickly locate the optimal policy execution sequence according to the current adversarial scenario, thereby improving the accuracy of the policy in practical applications.
[0057] Causality-driven dataset generation includes:
[0058] Extract the causal relationship feature data of the RL policy executor in multiple scenarios from the policy adversarial experiment data. The causal relationship feature data contains the association information of task antecedents, behavior nodes, and policy results;
[0059] Based on the extracted causal relationship features, construct a causality-driven dataset applicable to dynamic scenarios to improve the policy adjustment ability of the RL policy executor in different adversarial scenarios.
[0060] The causality-driven dataset is used to establish a policy feedback mechanism during the optimization process of the RL policy executor, specifically including:
[0061] Use the causality-driven dataset combined with the policy feedback results to adjust the decision-making rules of the RL policy executor in real time;
[0062] Through the feedback mechanism, gradually optimize the performance of the RL policy executor in different game adversarial scenarios.
[0063] Specifically, through the extraction of this causal relationship, the key decisions and results of the policy executor can be effectively associated, so as to more accurately analyze the applicability and effects of each policy in different scenarios. Based on the extracted causal relationship features, further construct a causality-driven dataset applicable to dynamic scenarios, thereby improving the ability of the RL policy executor to flexibly adjust policies in complex and changing adversarial environments. This causality-driven dataset is used to establish a real-time policy feedback mechanism during the optimization process of the policy executor. By combining the feedback results of policy performance with the causal feature data, the decision-making rules of the policy executor can be dynamically adjusted.
[0064] The optimization of the feedback mechanism includes the following steps:
[0065] Dynamically adjust the priority of causal feature data based on the confrontational performance of the RL strategy executor in different scenarios;
[0066] For a specific policy execution version, more targeted policy optimization can be achieved by dynamically adjusting the priority of causal data features.
[0067] Specifically, through this dynamic adjustment method, key features in adversarial situations can be flexibly identified, thereby more accurately guiding strategy optimization. When targeting a specific version of the strategy executor, the priority of these causal features can be further adjusted to make strategy optimization more targeted, thereby enhancing the effectiveness of the strategy executor in a variety of situations.
[0068] The calculation formula for strategy evaluation and optimization of the RL strategy executor is:
[0069] Suppose the performance index of the RL strategy execution body in a specific game situation is P i , the weight of the causal driven data feature is w i , the comprehensive evaluation score S of the adversarial strategy is expressed as:
[0070]
[0071] Among them, P i is the specific performance indicator of the strategy execution body in the confrontation game, w i is the causal feature weight in this situation, and by maximizing S, the adversarial optimization of the strategy executor is achieved.
[0072] Through adversarial experiments on multiple versions of RL policy executors, a task-context-driven policy rule framework is generated. The framework includes:
[0073] Based on the causal feature data and strategy confrontation performance, a strategy rule library for typical task scenarios is established;
[0074] Utilize the policy rule library to automatically generate adaptive policy rules in different task scenarios.
[0075] Specifically, the strategy rule library includes verified strategy paths that can cover a variety of adversarial situation characteristics, thereby forming a multi-level strategy alternative plan. When the framework is applied, it uses the strategy rule library to achieve automatic strategy matching and generation, and automatically generates strategy rules that are adapted to the characteristic requirements of the current task situation.
[0076] The generation method of the policy rule base includes:
[0077] Extract high-frequency causal chain nodes from the task behavior tree based on the performance of the RL policy executor in typical task scenarios;
[0078] Construct the high-frequency causal chain nodes as the basic units of the policy rule library, and optimize the execution effect of the adversarial policy according to the following formula:
[0079]
[0080] where R represents the average effectiveness score of the policy rule, and S j is the comprehensive evaluation score of the policy in the j-th adversarial experiment, T is the number of experiments, and by maximizing R, the stability of the policy rule in different scenarios is achieved.
[0081] Specifically, it can be obtained from the formula that the average effectiveness score R of the policy rule is calculated based on the multiple evaluation results in the adversarial experiment. Specifically, this formula calculates the average value of multiple adversarial experiment scores S j where S j is the comprehensive evaluation score of the policy in the j-th adversarial experiment, T is the total number of experiments. By calculating the mean of all experimental results, an index R that measures the stability of the policy rule is generated, which reflects the overall effectiveness of the policy rule library in different scenarios. By maximizing R, the policy rules that are stable and highly adaptable in various task scenarios can be selected, thus ensuring the robustness and continuous excellent performance of the policy executor in a dynamic environment.
[0082] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An evolution-based confrontation strategy evolution method, characterized in that: The following steps are involved: Create a parallel adversarial game training environment and use multiple reinforcement learning (RL) learners to generate multiple versions of RL policy actors to deal with different policy adversarial scenarios; Record and analyze the performance data of each RL strategy execution body, and generate a task behavior tree for strategy decision-making based on the performance data; Based on the causal features in the task behavior tree, a causal-driven dataset is constructed to guide the strategy optimization and evolution of the RL strategy executor; The causal-driven data set is used to dynamically adjust the adversarial strategy of the RL strategy executor, and the evolutionary optimization of the strategy adversarial capability is achieved through continuous iterative training.
2. The evolution-based confrontation strategy evolution method according to claim 1, characterized in that: The generation process of the RL strategy executor includes: through a parallel adversarial game training mode, each RL learner generates a strategy version with specific behavior characteristics, and the specific behavior characteristics cover different strategy advantages and strategy weaknesses.
3. The evolution-based confrontation strategy evolution method according to claim 1, characterized in that: The generation of the task behavior tree includes the following steps: Based on the data of the strategy confrontation experiment, behavior tree nodes are established. Each node contains the key behaviors of the RL strategy executor during the confrontation process. The causal links of the nodes in the behavior tree are analyzed to generate multiple strategy decision paths, which are used to determine the optimal strategy execution sequence under different confrontation scenarios.
4. The evolution-based confrontation strategy evolution method according to claim 1, characterized in that: The causal driven dataset generation includes: Extracting causal feature data of the RL strategy execution body in various scenarios from the strategy confrontation experiment data, wherein the causal feature data includes the correlation information of the task antecedents, behavior nodes and strategy results; Based on the extracted causal features, a causal-driven dataset suitable for dynamic scenarios is constructed to enhance the strategy adjustment capability of the RL strategy executor in different adversarial scenarios.
5. The method for evolving a countermeasure strategy based on evolution according to claim 1, characterized in that: The causal driving dataset is used to establish a policy feedback mechanism during the optimization process of the RL policy executor, specifically including: Use causal driven data sets combined with policy feedback results to adjust the decision rules of the RL policy executor in real time; Through the feedback mechanism, the performance of the RL strategy executor in different game confrontation scenarios is gradually optimized.
6. The evolution-based confrontation strategy evolution method according to claim 5, characterized in that: The optimization of the feedback mechanism comprises the following steps: Dynamically adjust the priority of causal feature data based on the confrontational performance of the RL strategy executor in different scenarios; For a specific policy execution version, more targeted policy optimization can be achieved by dynamically adjusting the priority of causal data features.
7. The method for evolving a countermeasure strategy based on evolution according to claim 6, characterized in that: The calculation formula for strategy evaluation and optimization of the RL strategy execution body is: Suppose the performance index of the RL strategy execution body in a specific game situation is P i , the weight of the causal driven data feature is w i , the comprehensive evaluation score S of the adversarial strategy is expressed as: Among them, P i is the specific performance indicator of the strategy execution body in the confrontation game, w i is the causal feature weight in this context.
8. The method for evolving a countermeasure strategy based on evolution according to claim 1, characterized in that: Through adversarial experiments on multiple versions of RL policy execution bodies, a task context-driven policy rule framework is generated. The framework includes: Based on the causal feature data and strategy confrontation performance, a strategy rule library for typical task scenarios is established; Utilize the policy rule library to automatically generate adaptive policy rules in different task scenarios.
9. The evolution-based confrontation strategy evolution method according to claim 8, characterized in that: The method for generating the policy rule base includes: Based on the performance of the RL policy executor in typical task scenarios, high-frequency causal chain nodes are extracted from the task behavior tree; The high-frequency causal chain nodes are constructed as the basic units of the strategy rule library, and the execution effect of the adversarial strategy is optimized according to the following formula: Among them, R represents the average effectiveness score of the policy rules, S j is the comprehensive evaluation score of the strategy in the jth adversarial experiment, and T is the number of experiments.