Intelligent opponent selection training framework based on rule-intelligent double-strategy library and fuzzy logic
By adopting an intelligent opponent selection training framework based on the ‘rule-intelligence’ dual strategy library and fuzzy logic in complex game scenarios, the problems of insufficient generalization capabilities of agents and instability in training in the existing technology are solved, and more efficient strategy optimization and adaptability are achieved.
Patent Information
- Application Number
- CN202510280465.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In complex game games, existing reinforcement learning training methods are difficult to effectively improve the generalization ability and adversarial performance of agents in different adversarial environments, especially when the opponent's strategy is inappropriate.
The intelligent opponent selection training framework based on the ‘rule-intelligent’ dual policy library and fuzzy logic is adopted, and the rule opponent strategy library and intelligent opponent strategy library are built, and the fuzzy comprehensive evaluation model is used to dynamically switch the opponent strategy library.
The generalization ability and game level of agents in complex game scenarios has been improved, and the training stability and strategy adaptability and optimization capabilities have been enhanced.
Smart Images

Figure CN120218279A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of reinforcement learning training methods, and specifically relates to an intelligent opponent selection training framework based on a "rule-intelligence" dual policy library and fuzzy logic. Background Art
[0002] In highly dynamic and uncertain adversarial environments such as aerial combat games, one of the core challenges of reinforcement learning is how to design a reasonable training mechanism so that the agent can not only converge stably but also have the ability to confront unknown opponents in complex environments. During the training process, the selection of opponents is one of the key factors affecting the learning effect of the agent. A suitable opponent can not only provide effective training signals to promote the optimization of the agent's strategy but also affect its exploration efficiency and generalization ability.
[0003] Existing research mainly adopts three types of reinforcement learning training methods. The first type of method uses rule-based opponent strategies, such as fixed rules or opponent strategies designed by experts, to provide a stable training environment. However, the generalization ability of this type of method is poor, and it is difficult to deal with dynamically changing opponents. The second type of method is based on self-play and its improved forms, which improve the agent's confrontation ability by competing with historical intelligent strategies. However, this method usually has problems such as low exploration efficiency and unstable training, which affect the final decision-making ability. The third type of method adopts curriculum learning, which guides the agent to learn by gradually increasing the difficulty of the opponent. Although it can improve the training efficiency, it lacks dynamic adaptability and is difficult to adjust according to the actual training state of the agent. These methods have their own advantages and disadvantages, but they all face certain limitations.
[0004] In summary, the research on the generalization ability of agents in complex game scenarios has received increasing attention. When reinforcement learning agents are trained in game scenarios, they usually face problems such as insufficient generalization ability and unstable training processes due to inappropriate opponent strategies. Therefore, it is very necessary to propose a reinforcement learning training framework that can perform intelligent opponent selection, so as to provide opponents of different levels and styles at different training stages to improve the generalization ability and confrontation performance of the agent. Summary of the Invention
[0005] The present invention provides an intelligent opponent selection training framework based on a "rule-intelligence" dual policy library and fuzzy logic. This training framework is a reinforcement learning training method, especially for training methods in complex game scenarios. The present invention introduces an intelligent opponent selection training framework, and introduces a "rule-intelligence" dual policy library and a policy library switching method based on fuzzy logic, which can improve training stability and policy generalization ability.
[0006] To solve the above problems, the technical solution adopted by the present invention is: an intelligent opponent selection training framework based on a "rule-intelligence" dual strategy library and fuzzy logic, and the training framework includes:
[0007] S1. The rule-based strategy combines expert experience with a behavior tree to establish a rule-based opponent strategy library, that is, a rule opponent strategy library, to provide a stable and reliable opponent for the training of reinforcement learning in the game scenario;
[0008] S2. An intelligent opponent strategy library generated by interacting with the rule strategy is established. The intelligent opponent strategy library is obtained by screening and sorting the agent models in the historical training process, and finally provides more novel and unpredictable opponents, effectively improving the diversity and flexibility of the opponent strategies;
[0009] S3. In a fixed number of training iterations, select an opponent from the opponent strategy library according to the decision result;
[0010] S4. Conduct an evaluation, that is, let the latest agent model play a game with a fixed high-level rule strategy;
[0011] S5. For the performance of the currently evaluated model and the real-time training results, select the indicators that can represent both as the factor set of the fuzzy comprehensive evaluation model;
[0012] S6. First, construct a fuzzy comprehensive evaluation model, use the fuzzy comprehensive evaluation model to decide whether to switch the opponent strategy library, and select the opponent strategy library to be used in the next fixed number of training iterations according to the result of this time.
[0013] In step S1 of the present invention, a pure rule attack strategy and a defense strategy are constructed based on expert experience; by combining the behavior tree with the pure rule, a more flexible and complex behavior tree strategy is constructed.
[0014] In step S2 of the present invention, an opponent is randomly selected from the rule opponent strategy library for training, and the agent model is saved regularly; the confrontation performance of the agent model is tested, and the unstable or deviated from human cognition models are eliminated; according to the win-loss results, the intelligent strategy models are divided into three levels: low, medium, and high, and models are evenly selected from each level and added.
[0015] In step S3 of the present invention, in the next fixed number of training, an opponent is selected from the opponent strategy library; if the rule opponent strategy library is selected, a uniform distribution random selection mechanism is adopted; if the intelligent opponent strategy library is selected, a priority virtual self-play mechanism is adopted for opponent matching.
[0016] In step S4 of the present invention, let the latest agent model play a game with a fixed high-level rule strategy and conduct an evaluation.
[0017] In step S5 of the present invention, the evaluation reward of the latest intelligent policy model is selected to measure its win-loss situation against the high-level rule policy; the difference between the current evaluation reward and the previous round of evaluation reward is calculated to measure the training status.
[0018] In step S6 of the present invention, the specific method for constructing the fuzzy comprehensive evaluation model is as follows: the membership function of each evaluation factor is introduced; the fuzzy evaluation matrix is constructed based on the membership function; the fuzzy vector is calculated to obtain the probabilities of selecting the rule opponent strategy library and the intelligent opponent strategy library, so as to make the final decision on switching the opponent strategy library.
[0019] The beneficial effects of adopting the above technical solutions are as follows: the present invention introduces an intelligent opponent selection training framework to systematically improve the generalization ability and game level of the agent's strategy. First, by constructing a "rule-intelligent" dual strategy library, the agent can rely on the rule opponent to learn basic strategies and skills in the initial stage of training, and gradually adapt to diverse intelligent opponents in the later stage of training, thereby enhancing its generalization ability and adaptability. In addition, the present invention adopts an opponent strategy library switching method based on fuzzy logic, enabling the agent to dynamically match opponents with different levels and styles at different training stages to promote the continuous optimization and evolution of the strategy. The experimental results verify the significant advantages of this method in improving the agent's generalization ability and game performance, indicating its wide applicability and practical value in complex game environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of an intelligent opponent selection training framework based on a "rule-intelligent" dual strategy library and fuzzy logic according to the present invention;
[0021] Figure 2 is a schematic diagram of the confrontation of the trained agent in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0023] A typical embodiment of the present invention is an intelligent opponent selection training framework based on a "rule-intelligent" dual strategy library and fuzzy logic in a high-fidelity air game scenario, as Figure 1 shown, and includes the following steps:
[0024] S1. Establish the rule-based opponent strategy library S rule , which contains two types of strategies: pure rule-based strategies S PR and strategies generated based on behavior trees S BT . Among them, S PR is derived from human expert experience and can be divided into attack strategies and defense strategies. S BT can dynamically select and adjust strategies according to different situations, and its complexity and flexibility are higher than S PR . This type of strategy combines S PR with the behavior tree to enable the agent to dynamically select appropriate S PR according to the current situation. The behavior tree determines the S PR that should be adopted in the current scenario by constructing conditional nodes and behavior nodes based on factors such as distance, the number of available missiles, and whether being locked by the opponent's missiles.
[0025] Therefore, the rule-based opponent strategy library S rule is defined as:
[0026] S rule = {S PR , S BT}
[0027] S2. The strategies in the intelligent opponent strategy library S AI are derived from the historical training process of the agent and are obtained through confrontation with the opponents in the rule-based opponent strategy library S rule . First, randomly select a strategy from S rule as the training opponent. During the training process, the intelligent strategy model is saved regularly. Subsequently, a large number of intelligent strategy models with different styles and levels are obtained, and their performance in confrontation is tested. The models that are unstable or seriously deviate from human cognition are discarded, and the remaining models are graded. The grading is based on the win-loss results in the evaluation stage, where the intelligent strategy model confronts a designated opponent and is divided into three levels: low, medium, and high according to the confrontation results. Finally, a certain number of intelligent strategy models are evenly selected from each level and added to S AI .
[0028] S3. In the next fixed number of training iterations, select opponents from the opponent strategy library according to the decision result in S6. If selected from the rule-based opponent strategy library, a uniform random selection mechanism is adopted; if selected from the intelligent opponent strategy library, the priority virtual self-play mechanism is used. The priority virtual self-play introduces a priority sampling mechanism to determine the probability of a candidate opponent being selected according to its win rate, so that the training focuses more on confrontations that can provide more effective learning signals and avoids repeated confrontations with significantly weaker opponents.
[0029]
[0030] Among them, \(f:[0,1]\to[0,\infty)\) is a weight function. In this paper, it is set that \(f(x)=1 - x\), that is, the agent tends to choose opponents that are more difficult to defeat and will not choose opponents with a 100% winning rate.
[0031] S4. Conduct an evaluation, that is, let the latest agent model play against a fixed high-level rule strategy.
[0032] S5. After each evaluation, calculate the return \(u_1\) of this evaluation and the difference \(u_2\) between it and the return of the previous evaluation.
[0033] S6. Establish a fuzzy evaluation matrix and assign membership degrees for each factor at each evaluation level. The membership function \(\mu_1\) describing different evaluation returns \(u_1\) under the concept of "selecting the rule-based opponent strategy library" is defined as follows:
[0034]
[0035] Among them, \(n\) represents the number of UAVs on both sides of the confrontation, \(x\in[-n,n]\) represents the value of the evaluation return, that is, the average number of opponent UAVs defeated during the evaluation process. \(c\in[1,2]\) is the return evaluation coefficient, and a larger \(c\) value indicates that a higher level of gaming ability is required when selecting the intelligent opponent strategy library.
[0036] The membership function \(\mu_2\) describing the difference \(u_2\) between two evaluation returns under the concept of "selecting the rule-based opponent strategy library" is defined as follows:
[0037]
[0038] Among them, \(x\in[-2n,2n]\) represents the numerical difference between the two evaluation returns.
[0039] Therefore, the fuzzy evaluation matrix \(R\) is defined as:
[0040]
[0041] Among them, each row of the matrix represents the membership degree of a certain factor at different evaluation levels.
[0042] Determine the weights of each factor as \(A = [0.5, 0.5]\). Select the weighted average method as the fuzzy logic rule, that is, through the fuzzy transformation, the fuzzy vector \(A\) on the set \(U\) is converted into the fuzzy vector \(B\) on the set \(V\). This process enables the derivation of the final comprehensive evaluation result based on the evaluation results and weights of each evaluation index. The calculation method of the fuzzy vector \(B\) is defined as follows:
[0043] \(B = A\cdot E=[b_1,b_2]\)
[0044] b1 and b2 respectively represent the probabilities of the selection rule for the opponent strategy library and the intelligent opponent strategy library, so as to obtain the opponent strategy library switching decision in the next fixed number of trainings in S3.
[0045] In summary, the present invention provides an intelligent opponent selection training framework that meets the requirements of reinforcement learning training in complex game scenarios, and improves the training effect by selecting appropriate opponent strategies for the agent at different training stages. The framework constructs a "rule-intelligent" dual strategy library and designs a method for switching the opponent strategy library based on fuzzy logic, enabling the opponent strategy to be dynamically adjusted according to the model performance and real-time training results, thereby improving the generalization ability and game performance of the final strategy.
[0046] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An intelligent opponent selection training framework based on "rule-intelligence" dual strategy library and fuzzy logic, characterized in that: The training framework includes: S1. Rule-based strategy combines expert experience with behavior trees to establish a rule-based opponent strategy library, namely the rule-based opponent strategy library; S2. Establish an intelligent adversary strategy library generated by interacting with rule strategies. The intelligent adversary strategies are obtained by screening and sorting the intelligent agent models in the historical training process. S3, in a fixed number of training iterations, selecting an opponent from the opponent strategy library according to the decision results; S4, conduct an evaluation, that is, let the latest intelligent agent model play against a fixed high-level rule strategy; S5. For the currently evaluated model performance and the real-time training results, select indicators that can represent both as the factor set of the fuzzy comprehensive evaluation model; S6. First, a fuzzy comprehensive evaluation model is constructed, and the fuzzy comprehensive evaluation model is used to decide whether to switch the opponent strategy library, and the opponent strategy library used in the next fixed number of training iterations is selected according to the result.
2. According to claim 1, a smart opponent selection training framework based on a "rule-intelligence" dual strategy library and fuzzy logic is characterized in that: In step S1, a pure rule-based attack strategy and a pure rule-based defense strategy are constructed based on expert experience; a more flexible and complex behavior tree strategy is constructed by combining the behavior tree with the pure rule.
3. The intelligent opponent selection training framework based on the "rule-intelligence" dual strategy library and fuzzy logic as described in claim 1 is characterized in that: The step S2 randomly selects opponents from the rule opponent strategy library for training and regularly saves the intelligent agent model; tests the confrontation performance of the intelligent agent model and eliminates models that are unstable or deviate from human cognition; divides the intelligent strategy model into three levels: low, medium and high according to the winning and losing results, and evenly selects models from each level to join.
4. According to claim 1, a smart opponent selection training framework based on a "rule-intelligence" dual strategy library and fuzzy logic is characterized in that: In step S3, in the next fixed number of trainings, an opponent is selected from the opponent strategy library; if a regular opponent strategy library is selected, a uniformly distributed random selection mechanism is adopted; if an intelligent opponent strategy library is selected, a priority virtual self-game mechanism is adopted for opponent matching.
5. An intelligent opponent selection training framework based on a "rule-intelligence" dual strategy library and fuzzy logic according to any one of claims 1 to 4, characterized in that: The step S4 is to make the latest intelligent agent model play a game with a fixed high-level rule strategy and perform an evaluation.
6. An intelligent opponent selection training framework based on a "rule-intelligence" dual strategy library and fuzzy logic according to any one of claims 1 to 4, characterized in that: The step S5 selects the evaluation reward of the latest intelligent strategy model to measure its success or failure against the high-level rule strategy; calculates the difference between the current evaluation reward and the previous round of evaluation reward to measure the training status.
7. An intelligent opponent selection training framework based on a "rule-intelligence" dual strategy library and fuzzy logic according to any one of claims 1 to 4, characterized in that: The specific method of step S6, constructing the fuzzy comprehensive evaluation model is: introducing the membership function of each evaluation factor; constructing the fuzzy evaluation matrix according to the membership function; The fuzzy vector is calculated to obtain the probability of selecting the rule opponent strategy library and the intelligent opponent strategy library, so as to make the final opponent strategy library switching decision.