Scheme generation method based on deep reinforcement learning

By applying the N-MORL algorithm and improved PPO algorithm in the RTS gaming environment, an agent for drone number matching and orchestration configuration is constructed, which solves the problems of low convergence efficiency and insufficient adaptability of multi-objective optimization tasks in the existing technology, and realizes efficient and flexible tactical decision-making in dynamic environments.

CN120168970APending Publication Date: 2025-06-20BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337881.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In complex RTS gaming environments, existing deep reinforcement learning algorithms have problems of low convergence efficiency and insufficient adaptability in multi-objective optimization tasks, making it difficult to generate real-time optimal solutions in dynamically changing environments.

Method used

Adopting agent training and optimization methods based on N-MORL algorithm, by constructing a unified sampling environment and improved PPO algorithm, an agent for rated drone number matching and actual orchestration configuration is built, and the drone number allocation and orchestration resource configuration scheme is optimized.

Benefits of technology

It improves strategy adaptability and decision-making efficiency in a dynamic environment, can quickly iterate and generate dynamic tactical choices, meet the orchestration requirements under the new situation, and reduce resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120168970A_ABST
    Figure CN120168970A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating an unmanned aerial vehicle arrangement scheme aiming at situation real-time change, in particular to a scheme generation method based on a mask vector motion shielding and deep reinforcement learning algorithm (Deep Reference Learning). Firstly, a reinforcement learning agent matched with the rated unmanned aerial vehicle number is constructed by combining the change of the new situation, and the theoretical unmanned aerial vehicle number required by the arrangement target meeting the new situation is obtained through the agent. And then, a reinforcement learning agent for actual arrangement and configuration is constructed in combination with combat information under the current situation, that is, the position from which target rated arrangement is performed is selected in combination with the unmanned aerial vehicle interception probability of the blue party so as to complete the overall arrangement target, and finally, the blue party target arrangement of the new situation is completed under the condition that the resource consumption is as small as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating a solution based on an improved deep reinforcement learning algorithm, and particularly to a method for training and optimizing an agent based on the N-MORL (Variance-stabilized Multi-objective Proximal Policy Optimization and NSGA-II) algorithm. Background Art

[0002] In the battlefield environment of modern real-time strategy (RTS) games, the dynamic changes and complexity of the situation information pose great challenges to command and decision-making. Traditional fixed decision-making solutions often struggle to meet the immediate battlefield requirements, and existing solution generation methods usually rely on pre-set rules and human experience, making it difficult to generate real-time optimal solutions in a rapidly changing environment. This manual-dependent approach is not only inefficient in complex game environments but also requires a high reaction speed from players, often failing to provide effective support in intense battles.

[0003] Deep Reinforcement Learning (DRL) algorithms offer potential solutions to this problem. By simulating the interaction of an agent in a game environment, DRL can gradually optimize the decision-making process through continuous learning and trial-and-error, enabling it to adaptively adjust tactics in a dynamically changing game situation. For example, in an environment with limited resources and changing enemy threats, deep reinforcement learning can quickly iterate and generate dynamic tactical options, providing the best offensive or defensive solutions for players.

[0004] Although deep reinforcement learning has been widely applied, it still faces the following main challenges when dealing with multi-objective optimization tasks (such as offensive strategies, resource allocation, unit preservation, etc.):

[0005] (1) Low convergence efficiency: In a complex RTS game environment, the large state space and action space often make traditional deep reinforcement learning methods require a large number of iterations to obtain an effective strategy, resulting in a slow convergence speed.

[0006] (2) Insufficient adaptability: When the game environment or the opponent's strategy mutates, existing algorithms are difficult to adapt in a timely manner, easily leading to lagging strategies and affecting the actual decision-making effect.

[0007] Reinforcement learning methods based on the PPO (Proximal Policy Optimization) algorithm have shown excellent performance in complex tasks and have the following advantages:

[0008] (1) Strong stability: By restricting the scope of policy updates, the PPO algorithm effectively avoids the instability problem caused by excessive policy updates, thus ensuring the stability of the learning process.

[0009] (2) High efficiency: PPO adopts a truncated advantage function and a clipping mechanism, reducing the computational complexity and improving the sample utilization rate, and can quickly converge to an effective policy.

[0010] (3) Good adaptability: The PPO algorithm shows good generality in various environments. Even when facing different task scenarios and dynamically changing environments, it can flexibly adapt and adjust the policy. Summary of the Invention

[0011] Aiming at the deficiencies of the above existing technologies, the present invention provides a method for generating a scheme based on a deep reinforcement learning algorithm. By constructing a unified sampling environment and an agent based on the N-MORL algorithm, this method can optimize the matching of the rated number of drones and the actual choreography configuration scheme in a dynamic environment, so as to meet the choreography requirements in the new situation on the basis of minimizing resource consumption. The purpose of the present invention is achieved through the following technical solutions.

[0012] A method for generating a scheme based on a deep reinforcement learning algorithm, the method comprising the following steps:

[0013] 1. Construct a unified sampling environment that can simulate the interaction between the agent and the environment, covering the scenarios of matching the rated number of drones and the actual choreography configuration, and the scenario can give the state and corresponding information mask required by the agent according to the situation information;

[0014] 2. Based on the N-MORL algorithm improved by PPO, construct agents for matching the rated number of drones and actual choreography configuration respectively. Each agent is optimized according to its specific decision-making goal. The agent for matching the rated number of drones is used to determine the drone number allocation scheme in the new situation, and the agent for actual choreography configuration is used to determine the optimal configuration scheme of the choreography resources;

[0015] 3. Through the continuous interaction between the agent and the unified interactive sampling environment, combined with the new situation elements, train the optimized agents for matching the rated number of drones and actual choreography configuration to adapt to the dynamic environment change requirements;

[0016] 4. In practical applications, first input the new situation information into the trained agent for matching the rated number of drones to generate the required rated number of drones; then, input the drone number requirement and the situation information into the agent for actual choreography configuration to determine the deployment theater and its number of specific drones, so as to meet the choreography requirements in the new situation on the basis of minimizing resource consumption. Brief Description of the Drawings

[0017] Figure 1 A method model generation flowchart for generating a solution based on an improved deep reinforcement learning algorithm;

[0018] Figure 2 N-MORL algorithm execution flowchart. Detailed implementation manners

[0019] The following will describe in detail the specific implementation manners and exemplary embodiments of the present invention. To provide a comprehensive understanding of the present invention, the following description covers many specific details. However, it is obvious to those skilled in the art that the present invention can be implemented without some of these specific details. The following description of the embodiments is only to provide an example for a clear understanding of the present invention. The present invention is not limited to the specific configurations and algorithms proposed below, but covers any modifications, substitutions, and improvements of related elements, components, and algorithms without departing from the spirit of the present invention.

[0020] The following will refer to the attached Figure 1 To describe the specific steps of a method for generating a solution based on an improved deep reinforcement learning algorithm according to an embodiment of the present invention as follows:

[0021] In the first step, a rated UAV quantity matching interaction environment model is established for the real-time RTS game battlefield.

[0022] The present invention first constructs a unified sampling environment that can simulate the interaction between an intelligent agent and the game environment, covering scenarios of the red side's unit UAV quantity allocation and choreography resource configuration. The scenario provides state information and corresponding mask masks for the intelligent agent according to the game situation information to ensure that the intelligent agent receives relevant battlefield data.

[0023] Based on the battlefield situation elements, an interaction sampling environment for actually choreographing and configuring the intelligent agent is constructed, including a rated UAV quantity matching interaction environment and an actual choreography configuration interaction environment:

[0024] (1) Based on the known battlefield information, set rules and constraints to limit the construction of the random environment;

[0025] (2) Model the red side's resources, where power i is the choreography ability of UAVs of type i, and num i is the total number of UAVs of type i. If k is the upper limit of UAV types, the overall modeling of the red side's resources is:

[0026] I i =(power i , num i ), i = 1,..., k

[0027] (3) Present resource availability information to the agent in a masked manner:

[0028]

[0029] Among them, m is the total number of enemy targets to be arranged, M j It is a mask space composed of different mask vectors. The mask is a 0, 1 vector related to the size of the action space. The valid position can be used as 1 and the invalid position can be 0. The mask plays a role in screening the effectiveness of the entire action generation process, so that the agent only outputs valid actions and shields invalid actions in the process of generating actions. is a mask for the drone types. For different enemy forces j, only valid drone types can be used as candidate drone types through the mask, and non-valid drone types cannot be used as candidate drone types, i.e., the drone types can be used; It is a mask for arranging the upper limit of the number of drones that can be used by the enemy force j. The vector position of the available number part is 1, and the vector position of the part exceeding the available number is 0.

[0030] (4) Generate the drone quantity matching index through random initialization to improve the data collection efficiency of the orchestration configuration agent:

[0031] S r =f(I r ,M r ,E r )

[0032] Among them, in the rated number of drones matching interactive environment, I r is the game situation information vector, M r is the mask vector, E r The entity information of the currently selected enemy force to be arranged includes the arrangement damage degree and health volume. The three parts of situation information are combined to obtain the real rated drone number matching state information S r .

[0033] (5) Design a reasonable reward value matching the rated number of drones:

[0034]

[0035] Among them, cost w is the cost of the selected drone type, n w is the number of drones selected, profit i The reward represents the reward for the scheduling target with the scheduling number i. This reward is a multi-target reward. and is the weight coefficient. The larger the reward value, the better the effect.

[0036] In the second step, design the action space of the intelligent agent for matching the rated number of drones and the use of masks.

[0037] Based on the scenario of matching the rated number of drones, design a reasonable action space:

[0038] A j =(W j ,N j ), j = 1, 2..., n

[0039] where n is the serial number of the enemy targets to be arranged, W j is the type of drones required for the target to be arranged with serial number j, and N j is the number of drones of type W j required for the target to be arranged with serial number j.

[0040] In the third step, construct an intelligent agent for matching the rated number of drones and conduct data sampling and training.

[0041] Based on the N-MORL algorithm, construct an intelligent agent for allocating the rated number of drones. The N-MORL algorithm combines the advantages of the multi-objective PPO and NSGA-II algorithms, improving the diversity of solutions while ensuring high solution quality, providing more diverse choices for optimizing decisions.

[0042] The intelligent agent for allocating the number of drones optimizes the drone number requirement plan in the new situation according to specific decision-making goals, while the intelligent agent for choreography configuration generates the best allocation of choreography resources according to the decision. The N-MORL algorithm is based on the multi-objective PPO algorithm:

[0043]

[0044] where A* is the best action, is the policy ratio, is the estimated value of the advantage function, ε is the clipping parameter, and t is the t-th sampled data in a round of data sampling.

[0045] Through continuous training of the intelligent agent in the interactive environment, based on the original policy information and the elements of the new situation of the game, optimize the intelligent agent for allocating the number of drones and the intelligent agent for choreography configuration to adapt to the requirements of the dynamically changing game scenario:

[0046]

[0047] where θ is the model parameter, α is the learning rate, is the expected value of the total reward.

[0048] During the process of generating actions by the Actor Net, by replacing the activation function of the variance network with the Sigmoid function, the convergence of the variance network is accelerated, and the volatility of action sampling is reduced:

[0049]

[0050] where X is the output vector of the variance network.

[0051] After the improved multi-objective PPO action is generated, taking this action vector as the initial solution, using the reward function as the fitness function, and making constraints on the numerical values of each dimension of the vector, a better solution set is further obtained through the NSGA-II algorithm, and solutions with higher diversity are given.

[0052] Based on the established interaction environment for matching the rated number of drones and the reinforcement learning agent designed based on the N-MORL algorithm, first, the initial data of the environment is generated by randomly initializing the data. Training with randomly generated data can better simulate the dynamic variability of the real scenario, so as to improve the adaptability of the trained agent. After receiving the situation information, the agent reconstructs the action output according to the mask and only outputs the valid actions masked by the mask, thereby improving the data sampling efficiency.

[0053] In the fourth step, model the interaction environment for the real-time RTS game battlefield choreography and configuration.

[0054] Since the data of the rated number of drones has been given by the rated number of drones agent, the actual choreography and configuration agent only needs to select the force with the minimum cost for using the rated drones according to the two objectives of the choreography probability prob and the choreography cost. Therefore, for this scenario, the optimization problem is a multi-objective optimization that minimizes the choreography cost under the condition of matching the rated number of drones.

[0055] (1) The state space is set as:

[0056]

[0057] M g =(M w ,M n )

[0058] where M g is the state mask in the actual choreography and configuration, which consists of two parts. M w is the force number containing the rated drone type selected, and M n is the quantity mask containing the type of this rated drone, that is, the final valid action is to use the drones that meet the quantity requirements from the force containing the rated matching drone type. The state space S f is composed of the drone resource types of the red force The number of drones of the Red Army forces and the mask M g constitute the rated drone usage plan P for the current scheduling target i i .

[0059] (2) The action space is set as:

[0060]

[0061] where G i is the unit number selected for scheduling the enemy's i theater, W i is the type of drone selected for scheduling the enemy's i theater, and N i is the number of drones used in scheduling the enemy's i theater.

[0062] (3) The reward function is set as:

[0063]

[0064]

[0065] where W1 and W2 are multi-objective weights, cost w is the cost of using drone type w, prob group is the true scheduling probability of the selected unit, is the true amount of drones used calculated by combining the scheduling probability, and cost group is some additional costs for the selected unit (such as the maintenance cost and resource supply cost of drones due to geographical location).

[0066] Step 5: The actual scheduling configuration model interacts with the environment and is trained.

[0067] Several rated drone quantity matching results and Red Army resource configurations are randomly initialized through the environment. Traversing these several rated drone quantity matching results as the interaction process, the agent sampling and the iteration and update of the actual scheduling configuration agent are completed. Since the agents are all based on the N-MORL algorithm, the training process of the model is the same as the rated drone quantity matching process.

[0068] Step 6: Integrate and apply the rated drone quantity matching and the actual scheduling configuration model.

[0069] The training processes of the rated drone quantity matching agent and the actual scheduling configuration agent are independent of each other, and they sample and train in their respective environments. After the agents are trained, they are integrated and used in upstream and downstream:[[]]

[0070] (1) Upstream: First, model the choreography target information and the red-side resource information into an interaction environment with a matching number of rated drones, and use the intelligent agent for matching the number of rated drones to give a decision-making plan for the number of rated drones;

[0071] (2) Downstream: Model the decision-making plan for the number of rated drones, the red-side resources, and the blue-side resources into an interaction environment for actual plan scheduling and configuration, and use the intelligent agent for plan scheduling and configuration to give an actual scheduling plan.

Claims

1. A game strategy generation algorithm based on deep reinforcement learning, characterized in that: The method comprises the following steps: (1) Construct a unified sampling environment that can simulate the interaction between the agent and the game environment, covering the scenarios of unit drone quantity allocation and resource configuration arrangement. The scenario provides the agent with status information and corresponding mask according to the game situation information to ensure that the agent receives relevant battlefield data; (2) Based on the N-MORL (variance-stabilized multi-objective proximal strategy optimization and NSGA-II) algorithm, the UAV number allocation agent and the orchestration configuration agent are constructed respectively; the UAV number allocation agent optimizes the UAV number demand plan under the new situation according to specific decision goals, while the orchestration configuration agent generates the optimal allocation of orchestration resources according to the decision; (3) Through continuous training of the agents in the interactive environment, based on the original strategy information and new game situation elements, the number of drones allocated to the agents and the arrangement of the agents are optimized to adapt to the dynamically changing game scene requirements; (4) In actual application, the original strategy and new situation information are first input into the trained drone quantity allocation agent to generate the current drone quantity demand; then the drone quantity demand and the unit configuration of the original strategy are input into the orchestration configuration agent to determine the drone deployment area and quantity, so as to meet the current game situation requirements with the least resource consumption.

2. The method according to claim 1, characterized in that Based on battlefield situation elements, an interactive sampling environment is constructed to actually arrange and configure intelligent agents, including: An interactive sampling environment for the actual orchestration and configuration agent is constructed based on battlefield situation elements, including: setting rules and constraints based on historical and known battlefield information to limit the construction of random environments; modeling the red team's resources, including troop positions, drone types and quantity allocation, and the probability of successful choreography targets; presenting resource availability information to the agent in a masked manner; and generating drone quantity matching indicators through random initialization to improve the data collection efficiency of the orchestration and configuration agent.

3. The method according to claim 1, characterized in that The rated drone quantity matching agent constructed based on the N-MORL algorithm includes: The state space of the intelligent agent is constructed according to the situation information and mask information given by the environment, and its state space dimension is specified by means of discretization. According to the decision-making goal of the completion degree of the intelligent agent's scheduling goals, the action space based on the number of drones used and the type of drones used is constructed.

4. The method according to claim 1, characterized in that: The actual orchestration configuration agent constructed based on the N-MORL algorithm includes: According to the situation information, mask information and rated UAV number matching information given by the environment, the state space of the intelligent agent is constructed and its state space dimension is specified by means of discretization; according to the rated UAV number matching target of the intelligent agent, the action space based on troop type, war zone type, UAV usage type and UAV usage number is constructed.

5. The method according to claim 3, characterized in that: Matching the intelligent agent and the environment interaction training based on the rated number of drones, including: During the interaction process, different levels of reward values ​​and round end indicators are given according to whether the degree of damage to the choreographed target can be achieved; the rated drone number matching agent interacts with the modeling environment adapted to the rated drone number matching, and the iteration and update of the rated drone number matching agent is completed through environment initialization, agent sampling, and agent training process.

6. The method according to claim 4, characterized in that Based on the actual arrangement, the agent and the environment interaction training are configured, including: During the interaction, different levels of reward values ​​are given according to whether the predetermined number of drone tasks can be completed and the number of drones required to complete these tasks, and whether the round is over is marked; different levels of reward values ​​and round end marks are given according to the achievement of the above two completion indicators; The actual orchestration configuration agent interacts with the adapted modeling environment, and completes the iteration and update of the actual orchestration configuration agent through environment initialization, agent sampling, and agent training.

7. The method according to claim 4, characterized in that The process of training and using intelligent agents based on the rated number of drones and the actual orchestration configuration intelligent agents includes: The training processes of the rated UAV quantity matching agent and the actual orchestration configuration agent are independent of each other, and they are sampled and trained in their own environments respectively; they are integrated after the agent training, and used upstream and downstream when integrated; upstream: first model the orchestration target information and the red team resource information as the rated UAV quantity matching interaction environment, and give the rated UAV quantity decision plan through the rated UAV quantity matching agent; downstream: model the rated UAV quantity decision plan and the red team resources as the actual orchestration configuration interaction environment, and give the UAV use location, the actual combat zone and the actual number of UAVs used through the actual orchestration configuration agent; finally, the final decision plan is obtained through the downstream model output.