Game Strategy Planning Method, System, Electronic Device and Storage Medium

By acquiring and modeling data of multi-agent game areas in real time, using Monte Carlo algorithm and deterministic strategy gradient algorithm to generate game counter strategy sets, solving the problem of insufficient strategy generation accuracy in the existing technology, and significantly improving the generation accuracy and efficiency of multi-agent game strategies.

CN120010265BActive Publication Date: 2025-06-17HARBIN ENGINEERING UNIVERSITY SANYA NANHAI INNOVATION & DEVELOPMENT BASE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457530.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-17
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing multi-agent game strategy planning method relies on static rules and cannot make full use of real-time dynamic data, resulting in insufficient accuracy of strategy generation and ineffective capture of complex and changing interactive relationships.

Method used

By obtaining real-time environmental data, posture data, position data and game party data of the target game area, performing environmental modeling, constructing state space, action space and constraints, using Monte Carlo algorithm and deterministic strategy gradient algorithm to solve the objective function, generate a game counter strategy set, and determine the motion planning data of each agent based on the strategy set.

Benefits of technology

It significantly improves the accuracy and efficiency of the generation of multi-agent game counter strategy, can better adapt to complex and changeable environments and task requirements, and improves the overall performance and decision-making effect of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010265B_ABST
    Figure CN120010265B_ABST
Patent Text Reader

Abstract

The present application provides a game strategy planning method, system, electronic device and storage medium, which relates to the field of artificial intelligence technology. The method is applied to multi-agents, and the multi-agents include at least one unmanned aerial vehicle and at least one unmanned surface vehicle. The method constructs an accurate environmental model, state and action space by obtaining environmental data, attitude data, position data and game player data of the target game area in real time, and efficiently solves the objective function of the current task by using the Monte Carlo algorithm and the deterministic policy gradient algorithm, generates a game countermeasure strategy set, and determines the motion planning data for each agent to execute the current task based on the game countermeasure strategy set, and controls each agent to execute the task according to the plan, significantly improving the accuracy of generating game countermeasure strategies, the accuracy of task allocation and the task execution efficiency of multi-agents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a game strategy planning method, system, electronic device, and storage medium. Background Art

[0002] Multi-agent game strategy planning technology, with its dynamic environment adaptability, efficient information sharing mechanism, multi-objective optimization ability, excellent robustness and fault tolerance, has become an important tool for decision-makers to formulate optimal strategies in complex interactive environments. This technology has shown extensive application potential in multiple fields such as robot collaborative task coordination, game entertainment strategy layout, and intelligent transportation system optimization, significantly improving the working efficiency and response speed of the system.

[0003] However, most current multi-agent game strategy planning methods rely on rule-based methods, and these rules are often preset manually, lacking flexibility and real-time performance. Due to the inability to fully utilize real-time dynamic data, these methods have problems with insufficient accuracy in the game strategy generation process. Specifically, static rule configurations are difficult to capture the complex and changeable interaction relationships in the environment, resulting in deviations between the strategy planning results and actual requirements, affecting the overall performance and decision-making effect of the system. Therefore, how to fully utilize real-time and effective data in the multi-agent game strategy planning process to improve the accuracy and efficiency of strategy generation has become a key problem to be solved urgently. Summary of the Invention

[0004] Embodiments of this application provide a game strategy planning method, system, electronic device, and storage medium.

[0005] According to the first aspect of this application, a game strategy planning method is provided, which is applied to multi-agents. The multi-agents include at least one unmanned aerial vehicle and at least one unmanned boat. The method includes:

[0006] Obtain the regional data of the target game area, where the regional data includes environmental data, attitude data, position data, and game party data;

[0007] Perform environmental modeling based on the regional data to construct a state space, an action space, and constraint conditions;

[0008] Determine an objective function according to the state space, the action space, and the constraint conditions, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a game countermeasure strategy set;

[0009] Determine the task planning data of the current task according to the task data of the current task, the regional data, and the game countermeasure strategy set. The task planning data includes the motion planning data of each agent;

[0010] Control each agent to execute the current task according to its corresponding motion planning data.

[0011] According to an embodiment of the present application, the environmental data includes underwater data and meteorological data; correspondingly,

[0012] Obtain the regional data of the target game area, including:

[0013] Obtain the game party data and underwater data based on vision, lidar sensors and sonar sensors;

[0014] Measure the attitude data of multiple agents based on an inertial measurement unit, and the attitude data includes attitude, acceleration, angular velocity, roll and yaw;

[0015] Monitor the meteorological data of the target game area based on environmental monitoring sensors, and the meteorological data includes wind speed, wind direction and temperature and humidity;

[0016] Obtain position data based on a global navigation and positioning unit.

[0017] According to an embodiment of the present application, determining the objective function according to the state space, action space and constraint conditions, and using the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a game countermeasure strategy set, including:

[0018] Determine the objective function according to the state space, action space and constraint conditions:

[0019]

[0020] Wherein, is the parameter of the policy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the policy, is the distribution sampling of the state s according to the policy μ, represents the reward and punishment function, and the reward and punishment function is used to determine the reward after the state s executes the action a according to the reward mechanism;

[0021] Determine that the gradient of the objective function is:

[0022]

[0023] Wherein, is the expectation of the state s sampled from the experience replay buffer D, is the policy function with respect to gradient, is the probability distribution that the agent selects the action a according to the parameter in the state s, is the action value function The gradient with respect to action a and fix action a as the policy The selected action, is the expected cumulative reward obtained by taking action a in state s and continuing to interact following policy μ;

[0024] Based on the Monte Carlo algorithm, the deterministic policy gradient algorithm, and the gradient of the objective function, determine multiple policy parameters that maximize the objective function ;

[0025] According to the multiple policy parameters that maximize the objective function and determine multiple initial game countermeasure strategies;

[0026] Conduct multiple rounds of iterative game simulations on the multiple initial game countermeasure strategies, and determine the game countermeasure strategy set based on the results of the multiple rounds of iterative game simulations.

[0027] According to an embodiment of the present application, the determining the task planning data of the current task according to the task data of the current task, the area data, and the game countermeasure strategy set includes:

[0028] Establish an effectiveness characterization model according to the area data and the task data:

[0029]

[0030] wherein, n is the total number of indicators;

[0031] Establish a game multi-constraint model according to the effectiveness characterization model, the area data, the task data, and in combination with spatio-temporal hybrid constraints;

[0032] Construct a resource scheduling system according to the shortest job first algorithm;

[0033] Decompose the current task into task allocation, task scheduling, path planning, and trajectory tracking according to the task data of the current task, the game countermeasure strategy set, the effectiveness characterization model, the game multi-constraint model, and the resource scheduling system, generate a multi-level task solving framework, and perform task solving layer by layer to obtain the task planning data.

[0034] According to an embodiment of the present application, the method further includes:

[0035] Test and evaluate the task planning data, and if the test and evaluation are passed, determine to adopt the current task planning data;

[0036] Iterate the game countermeasure strategy set, the effectiveness characterization model, and the game multi-constraint model according to the test and evaluation results.

[0037] According to an embodiment of the present application, controlling each agent to execute the current task according to its corresponding motion planning data includes:

[0038] Controlling obstacle avoidance, cruising, and cooperative operation of each agent according to the motion planning data of each agent;

[0039] Controlling the motion parameters of each agent according to the motion planning data of each agent.

[0040] According to a second aspect of the present application, a game strategy planning system is provided, which is applied to multiple agents. The multiple agents include at least one unmanned aerial vehicle and at least one unmanned surface vehicle. The system includes:

[0041] A multi-agent environment perception module, configured to obtain regional data of a target game area, where the regional data includes environmental data, attitude data, position data, and game party data;

[0042] A first game countermeasure dataset generation module, configured to perform environmental modeling according to the regional data, and construct a state space, an action space, and constraint conditions;

[0043] A second game countermeasure dataset generation module, configured to determine an objective function according to the state space, the action space, and the constraint conditions, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a game countermeasure strategy set;

[0044] A multi-level planning task allocation module, configured to determine task planning data of the current task according to the task data of the current task, the regional data, and the game countermeasure strategy set, where the task planning data includes motion planning data of each agent;

[0045] A multi-agent motion control module, configured to control each agent to execute the current task according to its corresponding motion planning data.

[0046] According to an embodiment of the present application, the environmental data includes underwater data and meteorological data; correspondingly,

[0047] The multi-agent environment perception module includes:

[0048] A first acquisition sub-module, configured to acquire game party data and underwater data based on vision, lidar sensors, and sonar sensors;

[0049] A second acquisition sub-module, configured to measure the attitude data of the multi-agent based on an inertial measurement unit, where the attitude data includes attitude, acceleration, angular velocity, roll, and yaw;

[0050] A third acquisition sub-module, configured to monitor meteorological data of the target game area based on environmental monitoring sensors, where the meteorological data includes wind speed, wind direction, and temperature and humidity;

[0051] A fourth acquisition sub-module, configured to acquire position data based on a global navigation and positioning unit.

[0052] According to a third aspect of the present application, there is provided an electronic device, including:

[0053] At least one processor; and

[0054] A memory communicatively connected to the at least one processor; wherein,

[0055] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the present application.

[0056] According to a fourth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method described in the present application.

[0057] The game strategy planning method, system, electronic device and storage medium according to the embodiments of the present application, by acquiring environmental data, attitude data, position data and game player data of a target game area in real time, constructing an accurate environmental model and state and action spaces, and using the Monte Carlo algorithm and the deterministic policy gradient algorithm to efficiently solve the objective function of the current task, generating a game countermeasure strategy set, and determining the motion planning data for each agent to execute the current task based on the game countermeasure strategy set, and controlling each agent to execute the task according to the plan, significantly improves the accuracy of generating game countermeasure strategies for multiple agents, the accuracy of task allocation and the task execution efficiency.

[0058] It should be understood that the teachings of the present application do not necessarily achieve all the beneficial effects described above. Instead, specific technical solutions can achieve specific technical effects, and other embodiments of the present application can also achieve beneficial effects not mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] By referring to the drawings and reading the detailed description below, the above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood. In the drawings, several embodiments of the present application are shown in an exemplary but not restrictive manner, where:

[0060] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0061] Figure 1 Shows a schematic implementation flow diagram of the game strategy planning method provided by the embodiments of the present application;

[0062] Figure 2A schematic diagram showing the implementation flow of the regional data acquisition operation of the game strategy planning method provided in an embodiment of the present application is shown;

[0063] Figure 3 A schematic diagram showing the composition structure of the game strategy planning system provided by an embodiment of the present application is shown;

[0064] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0065] In order to make the purpose, features, and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0066] Figure 1 A schematic diagram of the implementation process of the game strategy planning method provided in an embodiment of the present application is shown.

[0067] refer to Figure 1 The embodiment of the present application provides a game strategy planning method, which is applied to a multi-agent, wherein the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, and the method includes:

[0068] Operation 101 : acquiring area data of a target game area, where the area data includes environment data, posture data, position data, and game player data.

[0069] When multiple agents work together, in order to effectively cope with complex and changing environments and task requirements, it is necessary to generate game strategies. In this process, the first step is to obtain regional data of the target game area. Among them, regional data includes environmental data (such as terrain, weather conditions, etc.), posture data (the current direction, speed and other motion states of the agent), position data (specific coordinates of the agent and the target point), and game party data (type, number, ability, etc. of other agents).

[0070] In one embodiment of the present application, the embodiment of the present application is mainly aimed at cross-domain collaborative operation scenarios between sea and air, such as cross-domain game scenarios between sea and air. Therefore, the multi-agents include drones and unmanned boats. According to actual needs, drones and unmanned boats can be configured as one or more.

[0071] Operation 102 : performing environment modeling according to the regional data, and constructing a state space, an action space and constraint conditions.

[0072] In order to generate accurate game strategies for the target game area subsequently, it is also necessary to accurately describe the target game area. The description of the target game area can be understood as environmental modeling, which expresses the actual states and actions of multiple agents through environmental modeling, generates the state space and action space of multiple agents, and sets corresponding constraint conditions.

[0073] Among them, the state space is used to define various states that an agent may be in the environment, such as position, speed, orientation, remaining energy, etc., and the action space is used to define various actions or behaviors that an agent can take, such as moving forward, backward, turning, accelerating, etc. At the same time, in order to ensure the rationality and feasibility of the game strategy, a series of constraint conditions also need to be set according to the environmental characteristics and task requirements, such as avoiding collisions, maintaining communication, energy consumption limits, etc.

[0074] In an embodiment of the present application, the constrained actions, action space, and constraint conditions are all generated based on the reinforcement learning theory.

[0075] In this way, through accurate environmental modeling and the definition of the state and action spaces, a solid foundation is provided for the subsequent generation of game strategies, ensuring that agents can make optimal decisions in a complex and changing environment and achieve the task objectives.

[0076] Operation 103: Determine the objective function according to the state space, action space, and constraint conditions, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a set of game countermeasure strategies.

[0077] After completing the environmental modeling of the target game area, it is necessary to configure an objective function that conforms to the characteristics of the target game area according to the state space, action space, and constraint conditions generated during the modeling process. Among them, the objective function aims to describe the state or result that an agent expects to achieve after taking actions in the environment.

[0078] And in order to generate accurate game countermeasure strategies, in the process of solving the objective function, the embodiments of the present application design a solution method that combines the Monte Carlo algorithm and the deterministic policy gradient algorithm.

[0079] First, use the Monte Carlo algorithm for preliminary solution, that is, randomly sample the state space and action space through the Monte Carlo algorithm, simulate the behavior of the agent in the environment, and calculate the corresponding objective function value. Through multiple samplings and calculations, an initial solution of the objective function is obtained.

[0080] After obtaining the initial solution of the Monte Carlo algorithm, use the deterministic policy gradient algorithm for further solution. By calculating the gradient of the objective function, find the optimal game countermeasure strategy that maximizes the objective function, that is, in the solution process, continuously adjust the policy parameters to make the objective function value gradually increase until convergence or the stop condition is met.

[0081] In the process of solving the objective function, a series of optimal or approximately optimal game countermeasure strategies can be obtained, and these game countermeasure strategies are formed into a game countermeasure strategy set.

[0082] Among them, the game countermeasure strategy set is used to show the optimal response game countermeasure strategies in various situations that the intelligent agent may encounter in the target game area. Preferably, it can be four game countermeasure strategies such as confrontation entanglement, tracking and reconnaissance, approaching and expelling, and dynamic penetration.

[0083] In this way, by determining the objective function and using the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve, an accurate game countermeasure strategy set is generated, which provides strong support for the decision-making of the intelligent agent in the game area.

[0084] Operation 104: According to the task data, regional data, and game countermeasure strategy set of the current task, determine the task planning data of the current task. The task planning data includes the motion planning data of each intelligent agent.

[0085] Based on the task data of the current task, one or more game countermeasure strategies that can most effectively achieve the task objective and at the same time take into account the limitations of the regional data can be determined from the game countermeasure strategy set through the analysis of the task data, the simulation of the game scenario, and the evaluation of the strategy effect.

[0086] After determining the game countermeasure strategy or combination of game countermeasure strategies, these game countermeasure strategies also need to be assigned to multiple intelligent agents participating in the game. Among them, the determined game countermeasure strategies can be assigned to the corresponding intelligent agents by evaluating the capabilities of the intelligent agents, matching the task requirements, and considering the cooperation between the intelligent agents, and the motion planning data of each intelligent agent is generated.

[0087] In an embodiment of the present application, the motion planning data of the intelligent agent may include the starting position, target position, path planning, speed control, etc. of the intelligent agent.

[0088] Operation 105: Control each intelligent agent to execute the current task according to its corresponding motion planning data.

[0089] After determining the motion planning data of each intelligent agent, control each corresponding intelligent agent to execute the corresponding actions according to the motion planning data, and the current game task can be completed.

[0090] In an embodiment of the present application, controlling each intelligent agent to execute the current task according to its corresponding motion planning data includes: controlling the obstacle avoidance, cruising, and cooperative operation of each intelligent agent according to the motion planning data of each intelligent agent; controlling the motion parameters of each intelligent agent according to the motion planning data of each intelligent agent.

[0091] For the scenario of cross - domain game between sea and air, the multi - agents in this application include both unmanned boats and unmanned aerial vehicles. The control of multi - agents can be regarded as controlling the state of each agent, performing motion control such as obstacle avoidance, cruising, and cooperative operation on the unmanned aerial vehicle and the unmanned boat according to the motion planning data of each agent, and controlling motion parameters such as the altitude, speed, angular velocity of the unmanned aerial vehicle and the speed, heading of the unmanned boat according to the motion planning data.

[0092] Thus, in the embodiment of this application, by obtaining environmental data, attitude data, position data, and game - player data of the target game area in real - time, constructing an accurate environmental model and state - action space, and efficiently solving the objective function of the current task using the Monte Carlo algorithm and the deterministic policy gradient algorithm, generating a game counter - measure strategy set, and determining the motion planning data for each agent to execute the current task based on the game counter - measure strategy set, and controlling each agent to execute the task according to the plan, the accuracy of generating game counter - measure strategies for multi - agents, the accuracy of task allocation, and the task execution efficiency are significantly improved.

[0093] Figure 2 The figure shows a schematic flow chart of the implementation of the regional data acquisition operation of the game strategy planning method provided by the embodiment of this application.

[0094] Reference Figure 2 , in an embodiment of this application, the environmental data includes underwater data and meteorological data. The above operation 201 of obtaining the regional data of the target game area includes:

[0095] Operation 201, obtaining game - player data and underwater data based on vision, lidar sensors, and sonar sensors;

[0096] Operation 202, measuring the attitude data of multi - agents based on an inertial measurement unit, where the attitude data includes attitude, acceleration, angular velocity, roll, and pitch;

[0097] Operation 203, monitoring the meteorological data of the target game area based on environmental monitoring sensors, where the meteorological data includes wind speed, wind direction, and temperature - humidity;

[0098] Operation 204, obtaining position data based on a global navigation and positioning unit.

[0099] During the cooperative operation of multi - agents, it is necessary to monitor the regional data in real - time during the operation process, such as environmental data, data of other agents (game - players), etc. Among them, the environmental data includes underwater data and meteorological data.

[0100] For the acquisition of regional data in the embodiment of this application, vision, lidar sensors, sonar sensors, inertial measurement units, environmental monitoring sensors, and global navigation and positioning units are configured.

[0101] Among them, based on visual, lidar, and sonar sensors, the situation of the target game area can be perceived, and data of the game players and underwater data can be obtained in real time. Based on the inertial measurement unit, the real-time attitude data of the unmanned aerial vehicle and the unmanned surface vehicle can be measured, such as attitude, acceleration, angular velocity, roll, and yaw. Based on the environmental monitoring sensors, the meteorological data of the surrounding environment can be detected, such as measuring wind speed and direction, temperature, and humidity. Based on the global navigation and positioning unit, positioning and navigation can be achieved, and the real-time position data of the unmanned aerial vehicle and the unmanned surface vehicle can be obtained to ensure precise route control.

[0102] In an embodiment of the present application, for the above operation 103, a target function is determined according to the state space, action space, and constraint conditions, and a game countermeasure strategy set is solved and generated by using the Monte Carlo algorithm and the deterministic policy gradient algorithm, including:

[0103] Determine the target function according to the state space, action space, and constraint conditions:

[0104]

[0105] Among them, is the parameter of the policy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the policy, is the distribution sampling of the state s according to the policy μ, represents the reward and punishment function, and the reward and punishment function is used to determine the reward after the state s executes the action a according to the reward mechanism;

[0106] Determine that the gradient of the target function is:

[0107]

[0108] Among them, is the expectation of the state s sampled from the experience replay buffer D, is the policy function with respect to gradient, is the probability distribution that the agent selects the action a according to the parameter under the state s, is the action value function with respect to the gradient of the action a, and fixes the action a as the action selected by the policy selected, is the expected cumulative reward obtained by taking the action a under the state s and continuing to interact following the policy μ;

[0109] Based on the Monte Carlo algorithm, the deterministic policy gradient algorithm, and the gradient of the target function, determine multiple policy parameters that maximize the target function ;

[0110] According to multiple policy parameters that maximize the objective function and , determine multiple initial game countermeasure strategies;

[0111] Conduct multiple rounds of iterative game simulations on the multiple initial game countermeasure strategies, and determine the game countermeasure strategy set based on the results of the multiple rounds of iterative game simulations.

[0112] Specifically, after environmental modeling, the state space, action space, and constraint conditions of the agent are determined. According to the state space, action space, and constraint conditions, a reward function and a punishment mechanism for the agent can be designed to determine the objective function that conforms to the target game area and the current task based on the reward function, punishment mechanism, and each state and action in the state space and action space. Among them, the objective function can be:

[0113] Among them, is the parameter of the policy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the policy, is the distribution sampling of state s according to policy μ, represents the reward and punishment function, and the reward and punishment function is used to determine the reward after state s executes action a according to the reward mechanism.

[0114] Based on the determined objective function, combining the Monte Carlo algorithm and the deterministic policy gradient algorithm, through multiple rounds of game simulations, using the policy combinations of different agents to conduct red-blue confrontation game simulation and deduction, the objective function can be accurately solved to obtain the final game countermeasure strategy set.

[0115] Among them, the solution of the objective function can be based on maximizing the gradient of the objective function, and the gradient of the objective function can be expressed as:

[0116]

[0117] is the expectation of state s sampled from the experience replay buffer D, is the policy function with respect to gradient, is the probability distribution of the agent choosing action a according to parameter under state s, is the action value function with respect to the gradient of action a, and fix action a as the action selected by policy , is the expected cumulative reward obtained by taking action a under state s and continuing to interact following policy μ.

[0118] By continuously adjusting the parameters of the strategy, the gradient of the objective function can be maximized, thereby maximizing the corresponding objective function. Among them, the experience replay buffer initializes the game countermeasure strategy based on historical data according to the Monte Carlo algorithm.

[0119] After solving the objective function, by performing game simulation on the obtained strategy, the final set of game countermeasure strategies can be obtained.

[0120] In an embodiment of the present application, for the above operation 104, according to the task data, regional data, and the set of game countermeasure strategies of the current task, determine the task planning data of the current task, including:

[0121] According to the regional data and task data, establish an effectiveness characterization model:

[0122]

[0123] Among them, , n is the total number of indicators;

[0124] According to the effectiveness characterization model, regional data, task data, and combined with spatio-temporal mixed constraints, establish a game multi-constraint model;

[0125] Construct a resource scheduling system according to the shortest job first algorithm;

[0126] According to the task data, set of game countermeasure strategies, effectiveness characterization model, game multi-constraint model, and resource scheduling system of the current task, decompose the current task into task allocation, task scheduling, path planning, and trajectory tracking, generate a multi-level task solving framework, and perform task solving layer by layer to obtain the task planning data.

[0127] Specifically, task planning needs to consider the energy consumption of the agent, various requirements of the task, and various constraint conditions of the target game area. Therefore, the effectiveness characterization model, game multi-constraint model, and resource scheduling system are required during the task planning process.

[0128] First of all, it is necessary to construct an effectiveness characterization model, and the effectiveness characterization model can be expressed as:

[0129] Among them, represents the i-th performance indicator, and n is the total number of indicators.

[0130] Based on the target game area and actual task requirements, various performance indicators for evaluating effectiveness can be established, and an effectiveness characterization model can be constructed according to the performance indicators. When performing task planning, the value under tasks with different threat levels and agents with different demand quantities can be evaluated based on the effectiveness characterization model.

[0131] After the construction of the performance characterization model, a game multi-constraint model for task execution is also constructed based on the performance characterization model. Specifically, a game multi-constraint model is established based on the established performance characterization model and combined with spatio-temporal motion mixing factors, so as to evaluate the task time, voyage distance, required fuel quantity, risk level, etc. required for the agent to reach the target location during task planning.

[0132] To achieve optimal task scheduling, the Shortest Job First (SJF) algorithm is also used to construct a resource scheduling system, so as to allocate resources to each agent using SJF during task planning, improve the efficiency of drones and unmanned boats during task execution, reduce the energy consumption of sea-air cross-domain task planning, and improve the motion ability of the unmanned system.

[0133] After the construction of the performance characterization model, the game multi-constraint model, and the resource scheduling system, the current task can be decomposed into task allocation, task scheduling, path planning, and trajectory tracking according to the task data, game countermeasure strategy set, performance characterization model, game multi-constraint model, and resource scheduling system of the current task, generate a multi-level task solving framework, and perform task solving layer by layer to obtain task planning data.

[0134] Specifically, the task data includes the total task objective, the total number of agents required for the task, complexity, and required communication status. Therefore, based on the analysis of regional data, combined with the scale of multiple agents, task complexity, environmental conditions, communication network status, etc., the task planning framework can be constrained. The information-coupled sea-air cross-domain task planning (current task) is decomposed into steps such as task allocation, task scheduling, path planning, and trajectory tracking through task decomposition to generate a multi-level task solving framework. And task solving is performed layer by layer through the performance characterization model, the game multi-constraint model, and the resource scheduling system to obtain the final task planning data.

[0135] Among them, the solution of the path planning and trajectory tracking steps is carried out using the sea-air cross-domain trajectory homotopy update algorithm to solve the non-cooperative and dynamic obstacle avoidance problems under multi-objective constraints, and ensure the smooth obstacle avoidance and motion stability of drones and unmanned boats within a fixed time limit.

[0136] In an embodiment of the present application, after generating the task planning data, the task planning data is also evaluated, and each model and strategy is iterated according to the evaluation results, specifically including: testing and evaluating the task planning data, and determining to adopt the current task planning data if the testing and evaluation are passed; iterating the game countermeasure strategy set, the performance characterization model, and the game multi-constraint model according to the testing and evaluation results.

[0137] Specifically, after generating the task planning data, simulation tests and evaluations are also carried out on the task planning data from aspects such as performance indicators and action risks. When the task planning data passes the evaluation, the task planning data is determined to be adopted to improve the feasibility, safety, and generalization of the task planning.

[0138] Furthermore, the evaluation results are also used to iterate the game countermeasure strategy set, the effectiveness representation model, and the game multi-constraint model to improve the generalization of the optimal task planning solution method.

[0139] Figure 3 The composition structure diagram of the game strategy planning system provided by the embodiment of the present application is shown.

[0140] Reference Figure 3 , based on the above game strategy planning method, the embodiment of the present application also provides a game strategy planning system, which is applied to multi-agents. The multi-agents include at least one unmanned aerial vehicle and at least one unmanned surface vehicle. The system includes:

[0141] The multi-agent environment perception module 301 is used to obtain the area data of the target game area. The area data includes environmental data, attitude data, position data, and game player data;

[0142] The first game countermeasure data set generation module 302 is used to perform environmental modeling according to the area data, and construct the state space, action space, and constraint conditions;

[0143] The second game countermeasure data set generation module 303 is used to determine the objective function according to the state space, action space, and constraint conditions, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate the game countermeasure strategy set;

[0144] The multi-level planning task allocation module 304 is used to determine the task planning data of the current task according to the task data, area data, and game countermeasure strategy set of the current task. The task planning data includes the motion planning data of each agent;

[0145] The multi-agent motion control module 305 is used to control each agent to execute the current task according to its corresponding motion planning data.

[0146] In an embodiment of the present application, the environmental data includes underwater data and meteorological data; correspondingly, the multi-agent environment perception module 301 includes:

[0147] The first acquisition sub-module is used to acquire the game player data and underwater data based on vision, lidar sensors, and sonar sensors;

[0148] The second acquisition sub-module is used to measure the attitude data of the multi-agent based on the inertial measurement unit. The attitude data includes attitude, acceleration, angular velocity, roll, and pitch;

[0149] A third acquisition sub-module, configured to monitor meteorological data of a target game area based on an environmental monitoring sensor, where the meteorological data includes wind speed, wind direction, temperature, and humidity;

[0150] A fourth acquisition sub-module, configured to acquire position data based on a global navigation and positioning unit.

[0151] It should be noted that the description of the system in the embodiments of the present application is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments, so details are not repeated here. For the technical details not described in the game strategy planning system provided in the embodiments of the present application, they can be understood according to Figures 1 to 2 the description of any one of the accompanying drawings.

[0152] According to an embodiment of the present application, the present application also provides an electronic device and a non-transitory computer-readable storage medium.

[0153] Figure 4 FIG. shows a schematic block diagram of an exemplary electronic device 40 that can be used to implement embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.

[0154] As Figure 4 shown, the electronic device 40 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 40 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0155] Multiple components in the electronic device 40 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the game strategy planning method. For example, in some embodiments, the game strategy planning method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the game strategy planning method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the game strategy planning method in any other suitable way (e.g., by means of firmware).

[0157] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0159] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0161] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0162] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0163] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitations are imposed herein.

[0164] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A game strategy planning method, characterized in that: Applied to a multi-agent, the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, the method includes: Acquire regional data of a target gaming area, wherein the regional data includes environment data, posture data, position data, and gaming party data; Perform environmental modeling based on regional data to construct state space, action space and constraints; Determine the objective function based on the state space, action space and constraints, and use the Monte Carlo algorithm and deterministic policy gradient algorithm to solve and generate the game counter-strategy set; Determine task planning data for the current task according to task data of the current task, the area data and the game countermeasure strategy set, wherein the task planning data includes motion planning data of each agent; Control each agent to perform the current task according to its corresponding motion planning data; Among them, the Monte Carlo algorithm and the deterministic policy gradient algorithm are used to solve and generate the game counter-strategy set, including: randomly sampling the state space and action space through the Monte Carlo algorithm, simulating the behavior of the intelligent agent and calculating the objective function value, and obtaining the initial solution through multiple sampling and calculation; based on the initial solution, the deterministic policy gradient algorithm is used to optimize the policy parameters and maximize the objective function until convergence or the stopping condition is reached; the optimal strategy in the policy parameter optimization process is extracted to construct the game counter-strategy set, specifically: Determine the objective function based on the state space, action space and constraints: in, is the parameter of the strategy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the strategy, is the distribution sampling of state s according to strategy μ, represents the reward and punishment function, which is used to determine the reward of state s after executing action a according to the reward mechanism; Determine the gradient of the objective function as: in, is the expectation of the state s sampled from the experience replay buffer D, is the policy function about The gradient of In state s, the agent follows the parameters The probability distribution of choosing action a, is the action value function The gradient of action a, and fix action a as the policy The action you choose, is the expected cumulative reward obtained by taking action a in state s and continuing to interact according to strategy μ; Based on the Monte Carlo algorithm, the deterministic policy gradient algorithm and the gradient of the objective function, multiple policy parameters are determined to maximize the objective function. ; According to multiple policy parameters that maximize the objective function and , determine multiple initial game counter-strategies; Multiple rounds of iterative game simulation are performed on the multiple initial game countermeasures, and a game countermeasure set is determined based on the results of the multiple rounds of iterative game simulation. The game countermeasure set is used to illustrate the optimal game countermeasure strategy for various situations that the intelligent agent may encounter in the target game area.

2. The method according to claim 1, characterized in that The environmental data includes underwater data and meteorological data; accordingly, Get the regional data of the target gaming area, including: Acquire player data and underwater data based on vision, lidar sensors, and sonar sensors; Measuring attitude data of multiple agents based on an inertial measurement unit, wherein the attitude data includes attitude, acceleration, angular velocity, roll and pitch; Monitoring the meteorological data of the target gaming area based on environmental monitoring sensors, wherein the meteorological data includes wind speed, wind direction, temperature and humidity; Acquire location data based on global navigation and positioning unit.

3. The method according to claim 1, characterized in that Determining the task planning data of the current task according to the task data of the current task, the area data and the game countermeasure strategy set includes: Based on the regional data and task data, an effectiveness characterization model is established: in, to Indicates the 1st to nth performance indicators, where n is the total number of indicators; Based on the effectiveness characterization model, regional data, task data, and combined with spatiotemporal mixed constraints, a game multi-constraint model is established; Build a resource scheduling system based on the shortest job first algorithm; According to the task data of the current task, the game countermeasure strategy set, the effectiveness characterization model, the game multi-constraint model and the resource scheduling system, the current task is decomposed into task allocation, task scheduling, path planning and trajectory tracking, and a multi-level task solving framework is generated. The task is solved layer by layer to obtain the task planning data.

4. The method according to claim 3, characterized in that The method further comprises: Test and evaluate the mission planning data, and if the test and evaluation pass, determine to adopt the current mission planning data; The game countermeasure strategy set, effectiveness characterization model and game multi-constraint model are iterated based on the test and evaluation results.

5. The method according to claim 1, characterized in that The controlling each agent to perform the current task according to its corresponding motion planning data includes: Control the obstacle avoidance, cruising and collaborative operation of each intelligent agent according to the motion planning data of each intelligent agent; According to the motion planning data of each agent, the motion parameters of each agent are controlled.

6. A game strategy planning system, characterized in that: The system is used to implement the method according to any one of claims 1 to 5, and is applied to a multi-agent, wherein the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, and the system includes: A multi-agent environment perception module is used to obtain regional data of a target game area, wherein the regional data includes environmental data, posture data, position data, and game player data; The first game countermeasure data set generation module is used to model the environment based on the regional data and construct the state space, action space and constraints; The second game countermeasure data set generation module is used to determine the objective function according to the state space, action space and constraints, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate the game countermeasure strategy set; A multi-level planning task allocation module, used to determine the task planning data of the current task according to the task data of the current task, the area data and the game counter-strategy set, wherein the task planning data includes the motion planning data of each intelligent agent; The multi-agent motion control module is used to control each agent to perform the current task according to its corresponding motion planning data; Among them, the Monte Carlo algorithm and the deterministic policy gradient algorithm are used to solve and generate a game counter-strategy set, including: randomly sampling the state space and action space through the Monte Carlo algorithm, simulating the behavior of the intelligent agent and calculating the objective function value, and obtaining the initial solution through multiple sampling and calculations; based on the initial solution, the deterministic policy gradient algorithm is used to optimize the policy parameters and maximize the objective function until convergence or the stopping condition is reached; the optimal strategy in the policy parameter optimization process is extracted to construct a game counter-strategy set.

7. The system according to claim 6, characterized in that The environmental data includes underwater data and meteorological data; accordingly, The multi-agent environment perception module includes: The first acquisition submodule acquires the player data and underwater data based on vision, lidar sensors and sonar sensors; A second acquisition submodule is used to measure the attitude data of the multi-agent based on the inertial measurement unit, wherein the attitude data includes attitude, acceleration, angular velocity, roll and pitch; The third acquisition submodule is used to monitor the meteorological data of the target gaming area based on the environmental monitoring sensor, wherein the meteorological data includes wind speed, wind direction, temperature and humidity; The fourth acquisition submodule is used to acquire position data based on the global navigation and positioning unit.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Active power distribution network optimization scheduling method, system and device and storage medium

    CN118017492A

  • Spacecraft cluster game hunting motion planning method based on multi-agent reinforcement learning

    CN118056756A