Game strategy planning method and system, electronic equipment and storage medium
By obtaining and modeling data of the multi-agent game environment in real time, using Monte Carlo algorithm and deterministic strategy gradient algorithm to generate a game counter strategy set, solving the problem of insufficient accuracy in strategy generation in the existing technology, and achieving more efficient multi-agent game strategy planning.
Patent Information
- Application Number
- CN202510457530.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing multi-agent game strategy planning method relies on static rules and cannot make full use of real-time dynamic data, resulting in insufficient accuracy of strategy generation, affecting system performance and decision-making effect.
By obtaining the environmental data, posture data, position data and game party data of the target game area, performing environmental modeling, constructing state space, action space and constraints, using Monte Carlo algorithm and deterministic strategy gradient algorithm to solve the objective function, generate a game counter strategy set, and determine the motion planning data of each agent based on the strategy set.
It significantly improves the accuracy and efficiency of multi-agent game counter strategy generation, improves the accuracy of task allocation and task execution efficiency.
Smart Images

Figure CN120010265A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a game strategy planning method, system, electronic device and storage medium. Background Art
[0002] Multi-agent game strategy planning technology, with its dynamic environment adaptability, efficient information sharing mechanism, multi-objective optimization capability, excellent robustness and fault tolerance, has become an important tool for decision makers to formulate optimal strategies in complex interactive environments. This technology has shown a wide range of application potential in many fields such as robot collaborative task coordination, game entertainment strategy layout, and intelligent transportation system optimization, significantly improving the system's work efficiency and response speed.
[0003] However, most current multi-agent game strategy planning methods rely on rule-based methods, which are often pre-set by humans and lack flexibility and real-time performance. Due to the inability to fully utilize real-time dynamic data, these methods have the problem of insufficient accuracy in the process of game strategy generation. Specifically, static rule configurations are difficult to capture the complex and changeable interactive relationships in the environment, resulting in deviations between strategy planning results and actual needs, affecting the overall performance and decision-making effect of the system. Therefore, how to make full use of real-time and effective data in the process of multi-agent game strategy planning to improve the accuracy and efficiency of strategy generation has become a key issue that needs to be solved urgently. Summary of the invention
[0004] Embodiments of the present application provide a game strategy planning method, system, electronic device, and storage medium.
[0005] According to a first aspect of the present application, a game strategy planning method is provided, which is applied to a multi-agent, wherein the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, and the method includes: Acquire regional data of a target gaming area, wherein the regional data includes environment data, posture data, position data, and gaming party data; Model the environment based on regional data and construct state space, action space and constraints; Determine the objective function based on the state space, action space and constraints, and use the Monte Carlo algorithm and deterministic policy gradient algorithm to solve and generate the game counter-strategy set; Determine task planning data for the current task according to task data of the current task, the area data and the game countermeasure strategy set, wherein the task planning data includes motion planning data of each agent; Control each agent to perform the current task according to its corresponding motion planning data.
[0006] According to an embodiment of the present application, the environmental data includes underwater data and meteorological data; accordingly, Get the regional data of the target gaming area, including: Acquire player data and underwater data based on vision, lidar sensors, and sonar sensors; Measuring attitude data of multiple agents based on an inertial measurement unit, wherein the attitude data includes attitude, acceleration, angular velocity, roll and pitch; Monitoring the meteorological data of the target gaming area based on environmental monitoring sensors, wherein the meteorological data includes wind speed, wind direction, temperature and humidity; Acquire location data based on global navigation and positioning unit.
[0007] According to an implementation manner of the present application, determining the objective function according to the state space, action space and constraints, and using the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a game countermeasure strategy set includes: Determine the objective function based on the state space, action space and constraints:
[0008] in, is the parameter of the strategy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the strategy, is the distribution sampling of state s according to strategy μ, represents the reward and punishment function, which is used to determine the reward of state s after executing action a according to the reward mechanism; Determine the gradient of the objective function as:
[0009] in, is the expectation of the state s sampled from the experience replay buffer D, is the policy function about The gradient of In state s, the agent follows the parameters The probability distribution of choosing action a, is the action value function The gradient of action a, and fix action a as the policy The action you choose, is the expected cumulative reward obtained by taking action a in state s and continuing to interact according to strategy μ; Based on the Monte Carlo algorithm, the deterministic policy gradient algorithm and the gradient of the objective function, multiple policy parameters are determined to maximize the objective function. ; According to multiple policy parameters that maximize the objective function and , determine multiple initial game counter-strategies; Perform multiple rounds of iterative game simulation on the multiple initial game countermeasure strategies, and determine a game countermeasure strategy set based on the results of the multiple rounds of iterative game simulation.
[0010] According to an implementation manner of the present application, determining the task planning data of the current task according to the task data of the current task, the area data and the game countermeasure strategy set includes: Based on the regional data and task data, an effectiveness characterization model is established:
[0011] in, , n is the total number of indicators; Based on the effectiveness characterization model, regional data, task data, and combined with spatiotemporal mixed constraints, a game multi-constraint model is established; Build a resource scheduling system based on the shortest job first algorithm; According to the task data of the current task, the game countermeasure strategy set, the effectiveness characterization model, the game multi-constraint model and the resource scheduling system, the current task is decomposed into task allocation, task scheduling, path planning and trajectory tracking, and a multi-level task solving framework is generated. The task is solved layer by layer to obtain the task planning data.
[0012] According to an embodiment of the present application, the method further includes: Test and evaluate the mission planning data, and if the test and evaluation pass, determine to adopt the current mission planning data; The game countermeasure strategy set, effectiveness characterization model, and game multi-constraint model are iterated based on the test and evaluation results.
[0013] According to an embodiment of the present application, the controlling each agent to perform the current task according to its corresponding motion planning data includes: Control the obstacle avoidance, cruising, and collaborative operations of each intelligent agent based on its motion planning data; According to the motion planning data of each agent, the motion parameters of each agent are controlled.
[0014] According to a second aspect of the present application, a game strategy planning system is provided, which is applied to a multi-agent, wherein the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, and the system includes: A multi-agent environment perception module is used to obtain regional data of a target game area, wherein the regional data includes environmental data, posture data, position data, and game player data; The first game countermeasure data set generation module is used to model the environment based on the regional data and construct the state space, action space and constraints; The second game countermeasure data set generation module is used to determine the objective function according to the state space, action space and constraints, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate the game countermeasure strategy set; A multi-level planning task allocation module, used to determine the task planning data of the current task according to the task data of the current task, the area data and the game counter-strategy set, wherein the task planning data includes the motion planning data of each intelligent agent; The multi-agent motion control module is used to control each agent to perform the current task according to its corresponding motion planning data.
[0015] According to an embodiment of the present application, the environmental data includes underwater data and meteorological data; accordingly, The multi-agent environment perception module includes: The first acquisition submodule acquires the player data and underwater data based on vision, lidar sensors and sonar sensors; A second acquisition submodule is used to measure the attitude data of the multi-agent based on the inertial measurement unit, wherein the attitude data includes attitude, acceleration, angular velocity, roll and pitch; The third acquisition submodule is used to monitor the meteorological data of the target gaming area based on the environmental monitoring sensor, wherein the meteorological data includes wind speed, wind direction, temperature and humidity; The fourth acquisition submodule is used to acquire position data based on the global navigation and positioning unit.
[0016] According to a third aspect of the present application, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present application.
[0017] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the present application.
[0018] The game strategy planning method, system, electronic device and storage medium of the embodiments of the present application acquire the environmental data, posture data, position data and game party data of the target game area in real time, construct an accurate environment model and state and action space, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to efficiently solve the objective function of the current task, generate a game counter-strategy set, and determine the motion planning data for each intelligent agent to perform the current task based on the game counter-strategy set, control each intelligent agent to perform the task according to the plan, and significantly improve the accuracy of multi-agent game counter-strategy generation, task allocation accuracy and task execution efficiency.
[0019] It should be understood that the teachings of the present application are not required to achieve all of the beneficial effects described above, but specific technical solutions can achieve specific technical effects, and other embodiments of the present application can also achieve beneficial effects not mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become readily understood. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, wherein: In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0021] Figure 1 A schematic diagram of the implementation process of the game strategy planning method provided in the embodiment of the present application is shown; Figure 2 A schematic diagram showing the implementation flow of the regional data acquisition operation of the game strategy planning method provided in an embodiment of the present application is shown; Figure 3 A schematic diagram showing the composition structure of the game strategy planning system provided by an embodiment of the present application is shown; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0022] In order to make the purpose, features, and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0023] Figure 1 A schematic diagram of the implementation process of the game strategy planning method provided in an embodiment of the present application is shown.
[0024] refer to Figure 1 The embodiment of the present application provides a game strategy planning method, which is applied to a multi-agent, wherein the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, and the method includes: Operation 101 : acquiring area data of a target game area, where the area data includes environment data, posture data, position data, and game player data.
[0025] When multiple agents work together, in order to effectively cope with complex and changing environments and task requirements, it is necessary to generate game strategies. In this process, the first step is to obtain regional data of the target game area. Among them, regional data includes environmental data (such as terrain, weather conditions, etc.), posture data (the current direction, speed and other motion states of the agent), position data (specific coordinates of the agent and the target point), and game party data (type, number, ability, etc. of other agents).
[0026] In one embodiment of the present application, the embodiment of the present application is mainly aimed at cross-domain collaborative operation scenarios between sea and air, such as cross-domain game scenarios between sea and air. Therefore, the multi-agents include drones and unmanned boats. According to actual needs, drones and unmanned boats can be configured as one or more.
[0027] Operation 102 : performing environment modeling according to the regional data, and constructing a state space, an action space and constraint conditions.
[0028] In order to generate accurate game strategies for the target game area in the future, it is also necessary to accurately describe the target game area. The description of the target game area can be understood as environmental modeling, which describes the actual state and actions of the multi-agents through environmental modeling, generates the state space and action space of the multi-agents, and sets corresponding constraints.
[0029] Among them, the state space is used to define the various states that the agent may be in in the environment, such as position, speed, orientation, remaining energy, etc., and the action space is used to define the various actions or behaviors that the agent can take, such as moving forward, backward, turning, accelerating, etc. At the same time, in order to ensure the rationality and feasibility of the game strategy, a series of constraints need to be set according to the environmental characteristics and task requirements, such as avoiding collisions, maintaining communication, and limiting energy consumption.
[0030] In one embodiment of the present application, the constrained actions, action spaces and constraint conditions are all generated based on reinforcement learning theory.
[0031] In this way, through accurate environmental modeling and the definition of state and action space, a solid foundation is provided for the generation of subsequent game strategies, ensuring that the intelligent agent can make optimal decisions and achieve task goals in a complex and changing environment.
[0032] Operation 103, determining the objective function according to the state space, action space and constraints, and using the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a game counter-strategy set.
[0033] After completing the environmental modeling of the target game area, it is necessary to configure an objective function that meets the characteristics of the target game area based on the state space, action space, and set constraints generated during the modeling process. The objective function is intended to describe the state or result that the agent expects to achieve after taking an action in the environment.
[0034] In order to generate an accurate game counter-strategy, in the process of solving the objective function, the embodiment of the present application designs a solution method that combines the Monte Carlo algorithm with the deterministic policy gradient algorithm.
[0035] First, the Monte Carlo algorithm is used for preliminary solution. That is, the state space and action space are randomly sampled by the Monte Carlo algorithm to simulate the behavior of the intelligent agent in the environment and calculate the corresponding objective function value. After multiple sampling and calculation, an initial solution of the objective function is obtained.
[0036] After obtaining the initial solution of the Monte Carlo algorithm, the deterministic policy gradient algorithm is used for further solution. By calculating the gradient of the objective function, the optimal game countermeasure strategy that maximizes the objective function is found. That is, in the solution process, the objective function value is gradually increased by continuously adjusting the policy parameters until convergence is reached or the stopping condition is met.
[0037] In the process of solving the objective function, a series of optimal or approximately optimal game counter-strategies can be obtained, and these game counter-strategies are organized into a game counter-strategy set.
[0038] Among them, the game counter-strategy set is used to show the optimal game counter-strategy for the intelligent agent in various situations that may be encountered in the target game area, and the preferred ones can be four game counter-strategies such as confrontation and entanglement, tracking and reconnaissance, close expulsion and dynamic penetration.
[0039] In this way, by determining the objective function and solving it using the Monte Carlo algorithm and the deterministic policy gradient algorithm, an accurate set of game counter-strategies is generated, which provides strong support for the agent's decision-making in the game area.
[0040] Operation 104 , determining the task planning data of the current task according to the task data, area data and game countermeasure strategy set of the current task, wherein the task planning data includes the motion planning data of each intelligent agent.
[0041] Based on the task data of the current task, one or more game countermeasures that can most effectively achieve the task objectives while taking into account the limitations of regional data can be determined from the game countermeasures strategy set through analysis of the task data, simulation of the game scenario, and evaluation of the strategy effect.
[0042] After determining the game counter-strategy or game counter-strategy combination, it is also necessary to assign these game counter-strategies to multiple agents participating in the game. The determined game counter-strategy can be assigned to the corresponding agent by evaluating the agent's capabilities, matching the task requirements, and considering the collaboration between agents, and generating motion planning data for each agent.
[0043] In one embodiment of the present application, the motion planning data of the intelligent agent may include the starting position, target position, path planning, speed control, etc. of the intelligent agent.
[0044] Operation 105 , controlling each agent to execute the current task according to its corresponding motion planning data.
[0045] After determining the motion planning data of each agent, each corresponding agent is controlled to perform corresponding actions according to the motion planning data to complete the current game task.
[0046] In one embodiment of the present application, each intelligent agent is controlled to perform the current task according to its corresponding motion planning data, including: controlling the obstacle avoidance, cruising, and collaborative operation of each intelligent agent according to the motion planning data of each intelligent agent; and controlling the motion parameters of each intelligent agent according to the motion planning data of each intelligent agent.
[0047] For the cross-domain game scenario of sea and air, the multi-agents of this application include both unmanned boats and drones. The control of multi-agents can be regarded as controlling the state of each agent, and performing motion control such as obstacle avoidance, cruising, and collaborative operation on drones and unmanned boats according to the motion planning data of each agent, and controlling the altitude, speed, angular velocity of the drones and the speed, heading and other motion parameters of the unmanned boats according to the motion planning data.
[0048] In this way, the embodiments of the present application acquire the environmental data, posture data, position data and game party data of the target game area in real time, construct an accurate environmental model and state and action space, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to efficiently solve the objective function of the current task, generate a game counter-strategy set, and determine the motion planning data for each intelligent agent to perform the current task based on the game counter-strategy set, and control each intelligent agent to perform the task according to the plan, which significantly improves the accuracy of multi-agent game counter-strategy generation, task allocation accuracy and task execution efficiency.
[0049] Figure 2A schematic diagram of the implementation flow of the regional data acquisition operation of the game strategy planning method provided in an embodiment of the present application is shown.
[0050] refer to Figure 2 In one embodiment of the present application, the environmental data includes underwater data and meteorological data. The above operation 201, obtaining the regional data of the target gaming area, includes: Operation 201, acquiring player data and underwater data based on vision, laser radar sensors and sonar sensors; Operation 202, measuring attitude data of the multi-agent based on an inertial measurement unit, where the attitude data includes attitude, acceleration, angular velocity, roll, and pitch; Operation 203, monitoring the meteorological data of the target gaming area based on the environmental monitoring sensor, where the meteorological data includes wind speed, wind direction, temperature and humidity; In operation 204, position data is acquired based on the global navigation and positioning unit.
[0051] In the process of multi-agent collaborative operation, it is necessary to monitor the regional data in the operation process in real time, such as environmental data, data of other agents (players), etc. Among them, environmental data includes underwater data and meteorological data.
[0052] In order to acquire regional data, the embodiment of the present application is equipped with vision and lidar sensors, sonar sensors, inertial measurement units, environmental monitoring sensors and global navigation and positioning units.
[0053] Among them, based on vision and lidar sensors and sonar sensors, it is possible to perceive the situation in the target game area and obtain the game party data and underwater data in real time. Based on the inertial measurement unit, the real-time attitude data of drones and unmanned boats can be measured, such as attitude, acceleration and angular velocity, roll and bow pitch, etc. Based on environmental monitoring sensors, it is possible to detect meteorological data of the surrounding environment, such as measuring wind speed and direction, temperature and humidity, etc. Based on the global navigation and positioning unit, positioning and navigation can be achieved, and the real-time position data of drones and unmanned boats can be obtained to ensure precise route control.
[0054] In one embodiment of the present application, the above operation 103 determines the objective function according to the state space, the action space and the constraints, and uses the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate a game countermeasure strategy set, including: Determine the objective function based on the state space, action space and constraints:
[0055] in, is the parameter of the strategy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the strategy, is the distribution sampling of state s according to strategy μ, represents the reward and punishment function, which is used to determine the reward of state s after executing action a according to the reward mechanism; Determine the gradient of the objective function as:
[0056] in, is the expectation of the state s sampled from the experience replay buffer D, is the policy function about The gradient of In state s, the agent follows the parameters The probability distribution of choosing action a, is the action value function The gradient of action a, and fix action a as the policy The action you choose, is the expected cumulative reward obtained by taking action a in state s and continuing to interact according to strategy μ; Based on the Monte Carlo algorithm, the deterministic policy gradient algorithm and the gradient of the objective function, multiple policy parameters are determined to maximize the objective function. ; According to multiple policy parameters that maximize the objective function and , determine multiple initial game counter-strategies; Multiple rounds of iterative game simulation are performed on multiple initial game countermeasures, and a set of game countermeasures is determined based on the results of the multiple rounds of iterative game simulation.
[0057] Specifically, after the environment is modeled, the state space, action space and constraints of the agent are determined. Based on the state space, action space and constraints, the reward function and penalty mechanism for the agent can be designed to determine the objective function that meets the target game area and this task based on the reward function, penalty mechanism and each state and action in the state space and action space. Among them, the objective function can be:
[0058] in, is the parameter of the strategy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the strategy, is the distribution sampling of state s according to strategy μ, Represents the reward and punishment function, which is used to determine the reward of state s after executing action a according to the reward mechanism.
[0059] Based on the determined objective function, combined with the Monte Carlo algorithm and the deterministic policy gradient algorithm, through multiple rounds of game simulation, the red-blue confrontation game simulation is carried out using the strategy combination of different intelligent agents. The objective function can be accurately solved and the final game counter-strategy set can be obtained.
[0060] Among them, the solution of the objective function can be based on maximizing the gradient of the objective function, and the gradient of the objective function can be expressed as:
[0061] is the expectation of the state s sampled from the experience replay buffer D, is the policy function about The gradient of In state s, the agent follows the parameters The probability distribution of choosing action a, is the action value function The gradient of action a, and fix action a as the policy The action you choose, It is the expected cumulative reward obtained by taking action a in state s and continuing to interact according to strategy μ.
[0062] By continuously adjusting the parameters of the strategy, the gradient of the objective function can be maximized, thereby maximizing the corresponding objective function. The experience replay buffer is initialized based on the Monte Carlo algorithm according to historical data to obtain the game countermeasure strategy.
[0063] After solving the objective function, the final set of game countermeasure strategies can be obtained by performing game simulation on the solved strategies.
[0064] In one embodiment of the present application, the above operation 104 determines the task planning data of the current task according to the task data, the area data and the game countermeasure strategy set of the current task, including: Based on the regional data and task data, an effectiveness characterization model is established:
[0065] in, , n is the total number of indicators; Based on the effectiveness characterization model, regional data, task data, and combined with spatiotemporal mixed constraints, a game multi-constraint model is established; Build a resource scheduling system based on the shortest job first algorithm; According to the task data of the current task, the game countermeasure strategy set, the effectiveness characterization model, the game multi-constraint model and the resource scheduling system, the current task is decomposed into task allocation, task scheduling, path planning and trajectory tracking, and a multi-level task solving framework is generated. The task is solved layer by layer to obtain the task planning data.
[0066] Specifically, task planning needs to take into account the energy consumption of the intelligent agent, various requirements of the task, and various constraints of the target game area. Therefore, performance characterization models, game multi-constraint models and resource scheduling systems are needed in the task planning process.
[0067] First, we need to build an effectiveness characterization model, which can be expressed as:
[0068] in, represents the i-th performance indicator, and n is the total number of indicators.
[0069] According to the target game area and the actual task requirements, various performance indicators for evaluating effectiveness can be established, and an effectiveness characterization model can be constructed based on the performance indicators. In task planning, the value of tasks with different threat levels and different required numbers of agents can be evaluated based on the effectiveness characterization model.
[0070] After the performance characterization model is constructed, the game multi-constraint model of task execution is also constructed based on the performance characterization model. Specifically, based on the established performance characterization model and combined with the mixed factors of space-time movement, the game multi-constraint model is established to evaluate the task time, flight distance, required fuel, risk level, etc. required for the intelligent agent to reach the target location during task planning.
[0071] In order to achieve optimal task scheduling, the shortest job first (SJF) algorithm is also used to build a resource scheduling system. When planning tasks, SJF is used to allocate resources to each intelligent agent, improve the efficiency of drones and unmanned boats in performing tasks, reduce the energy loss of cross-domain sea and air mission planning, and improve the mobility of unmanned systems.
[0072] After the performance characterization model, game multi-constraint model and resource scheduling system are constructed, the current task can be decomposed into task allocation, task scheduling, path planning and trajectory tracking according to the task data of the current task, the game counter-strategy set, the performance characterization model, the game multi-constraint model and the resource scheduling system, and a multi-level task solving framework can be generated. The task can be solved layer by layer to obtain the task planning data.
[0073] Specifically, the task data includes the total task goal, the total number of agents required for the task, the complexity, and the required communication status. Therefore, based on the analysis of regional data, the task planning framework can be constrained by combining the scale of multiple agents, task complexity, environmental conditions, and communication network status. Through task decomposition, the information-coupled sea and air cross-domain task planning (current task) is decomposed into task allocation, task scheduling, path planning, and trajectory tracking, generating a multi-level task solving framework. The task is solved layer by layer through the performance characterization model, the game multi-constraint model, and the resource scheduling system to obtain the final task planning data.
[0074] Among them, the path planning and trajectory tracking steps are solved using the sea-air cross-domain trajectory homology update algorithm to solve the non-cooperative and dynamic obstacle avoidance problems under multi-objective constraints, and to ensure the smooth obstacle avoidance and motion stability of UAVs and unmanned boats under fixed time constraints.
[0075] In one embodiment of the present application, after generating the task planning data, the task planning data is evaluated, and each model and strategy is iterated according to the evaluation results, specifically including: testing and evaluating the task planning data, and if the test and evaluation pass, determining to adopt the current task planning data; iterating the game countermeasure strategy set, effectiveness characterization model, and game multi-constraint model according to the test and evaluation results.
[0076] Specifically, after generating the mission planning data, simulation tests and evaluations are conducted on the mission planning data from the aspects of performance indicators, action risks, etc. When the mission planning data evaluation passes, the mission planning data will be adopted to improve the feasibility, safety and generalization of mission planning.
[0077] Furthermore, the evaluation results are used to iterate the game counter-strategy set, effectiveness characterization model, and game multi-constraint model to improve the generalization of the optimal task planning solution method.
[0078] Figure 3 A schematic diagram of the composition structure of the game strategy planning system provided in an embodiment of the present application is shown.
[0079] refer to Figure 3 Based on the above-mentioned game strategy planning method, the embodiment of the present application further provides a game strategy planning system, which is applied to a multi-agent, wherein the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, and the system includes: The multi-agent environment perception module 301 is used to obtain the regional data of the target game area, and the regional data includes environmental data, posture data, position data and game party data; The first game countermeasure data set generation module 302 is used to perform environment modeling based on regional data and construct state space, action space and constraint conditions; The second game countermeasure data set generation module 303 is used to determine the objective function according to the state space, action space and constraint conditions, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate the game countermeasure strategy set; The multi-level planning task allocation module 304 is used to determine the task planning data of the current task according to the task data, area data and game countermeasure strategy set of the current task, and the task planning data includes the motion planning data of each intelligent agent; The multi-agent motion control module 305 is used to control each agent to perform the current task according to its corresponding motion planning data.
[0080] In one embodiment of the present application, the environmental data includes underwater data and meteorological data; accordingly, the multi-agent environmental perception module 301 includes: The first acquisition submodule acquires the player data and underwater data based on vision, lidar sensors and sonar sensors; The second acquisition submodule is used to measure the attitude data of the multi-agent based on the inertial measurement unit, and the attitude data includes attitude, acceleration, angular velocity, roll and pitch; The third acquisition submodule is used to monitor the meteorological data of the target gaming area based on the environmental monitoring sensor, and the meteorological data includes wind speed, wind direction, temperature and humidity; The fourth acquisition submodule is used to acquire position data based on the global navigation and positioning unit.
[0081] It should be noted that the description of the system in the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated. Figure 1 to Figure 2 The present invention can be understood by referring to the description of any one of the accompanying drawings.
[0082] According to an embodiment of the present application, the present application also provides an electronic device and a non-transitory computer-readable storage medium.
[0083] Figure 4 A schematic block diagram of an example electronic device 40 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0084] like Figure 4 As shown, the electronic device 40 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 40 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0085] A number of components in the electronic device 40 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 40 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0086] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as the game strategy planning method. For example, in some embodiments, the game strategy planning method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 40 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the game strategy planning method described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the game strategy planning method in any other appropriate manner (e.g., by means of firmware).
[0087] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0088] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0089] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0091] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0092] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0093] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this application can be achieved, and this document is not limited here.
[0094] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A game strategy planning method, characterized in that: Applied to a multi-agent, the multi-agent includes at least one unmanned aerial vehicle and at least one unmanned boat, the method includes: Acquire regional data of a target gaming area, wherein the regional data includes environment data, posture data, position data, and gaming party data; Model the environment based on regional data and construct state space, action space and constraints; Determine the objective function based on the state space, action space and constraints, and use the Monte Carlo algorithm and deterministic policy gradient algorithm to solve and generate the game counter-strategy set; Determine task planning data for the current task according to task data of the current task, the area data and the game countermeasure strategy set, wherein the task planning data includes motion planning data of each agent; Control each agent to perform the current task according to its corresponding motion planning data; Among them, the Monte Carlo algorithm and the deterministic policy gradient algorithm are used to solve and generate a game counter-strategy set, including: randomly sampling the state space and action space through the Monte Carlo algorithm, simulating the behavior of the intelligent agent and calculating the objective function value, and obtaining the initial solution through multiple sampling and calculations; based on the initial solution, the deterministic policy gradient algorithm is used to optimize the policy parameters and maximize the objective function until convergence or the stopping condition is reached; the optimal strategy in the policy parameter optimization process is extracted to construct a game counter-strategy set.
2. The method according to claim 1, characterized in that The environmental data includes underwater data and meteorological data; accordingly, Get the regional data of the target gaming area, including: Acquire player data and underwater data based on vision, lidar sensors, and sonar sensors; Measuring attitude data of multiple agents based on an inertial measurement unit, wherein the attitude data includes attitude, acceleration, angular velocity, roll and pitch; Monitoring the meteorological data of the target gaming area based on environmental monitoring sensors, wherein the meteorological data includes wind speed, wind direction, temperature and humidity; Acquire location data based on global navigation and positioning unit.
3. The method according to claim 1, characterized in that The objective function is determined according to the state space, action space and constraints, and the Monte Carlo algorithm and the deterministic policy gradient algorithm are used to solve and generate the game counter-strategy set, including: Determine the objective function based on the state space, action space and constraints: in, is the parameter of the strategy, E is the expectation, s is the state in the state space, a is the action corresponding to the state in the action space, μ is the strategy, is the distribution sampling of state s according to strategy μ, represents the reward and punishment function, which is used to determine the reward of state s after executing action a according to the reward mechanism; Determine the gradient of the objective function as: in, is the expectation of the state s sampled from the experience replay buffer D, is the policy function about The gradient of In state s, the agent follows the parameters The probability distribution of choosing action a, is the action value function The gradient of action a, and fix action a as the policy The action you choose, is the expected cumulative reward obtained by taking action a in state s and continuing to interact according to strategy μ; Based on the Monte Carlo algorithm, the deterministic policy gradient algorithm and the gradient of the objective function, multiple policy parameters are determined to maximize the objective function. ; According to multiple policy parameters that maximize the objective function and , determine multiple initial game counter-strategies; Perform multiple rounds of iterative game simulation on the multiple initial game countermeasure strategies, and determine a game countermeasure strategy set based on the multiple rounds of iterative game simulation results.
4. The method according to claim 1, characterized in that: Determining the task planning data of the current task according to the task data of the current task, the area data and the game countermeasure strategy set includes: Based on the regional data and task data, an effectiveness characterization model is established: in, to Indicates the 1st to nth performance indicators, where n is the total number of indicators; Based on the effectiveness characterization model, regional data, task data, and combined with spatiotemporal mixed constraints, a game multi-constraint model is established; Build a resource scheduling system based on the shortest job first algorithm; According to the task data of the current task, the game countermeasure strategy set, the effectiveness characterization model, the game multi-constraint model and the resource scheduling system, the current task is decomposed into task allocation, task scheduling, path planning and trajectory tracking, and a multi-level task solving framework is generated. The task is solved layer by layer to obtain the task planning data.
5. The method according to claim 4, characterized in that The method further comprises: Test and evaluate the mission planning data, and if the test and evaluation pass, determine to adopt the current mission planning data; The game countermeasure strategy set, effectiveness characterization model, and game multi-constraint model are iterated based on the test and evaluation results.
6. The method according to claim 1, characterized in that The controlling each agent to perform the current task according to its corresponding motion planning data includes: Control the obstacle avoidance, cruising, and collaborative operations of each intelligent agent based on its motion planning data; According to the motion planning data of each agent, the motion parameters of each agent are controlled.
7. A game strategy planning system, characterized in that: Applied to multi-agents, the multi-agents include at least one unmanned aerial vehicle and at least one unmanned boat, and the system includes: A multi-agent environment perception module is used to obtain regional data of a target game area, wherein the regional data includes environmental data, posture data, position data, and game player data; The first game countermeasure data set generation module is used to model the environment based on the regional data and construct the state space, action space and constraints; The second game countermeasure data set generation module is used to determine the objective function according to the state space, action space and constraints, and use the Monte Carlo algorithm and the deterministic policy gradient algorithm to solve and generate the game countermeasure strategy set; A multi-level planning task allocation module, used to determine the task planning data of the current task according to the task data of the current task, the area data and the game counter-strategy set, wherein the task planning data includes the motion planning data of each intelligent agent; The multi-agent motion control module is used to control each agent to perform the current task according to its corresponding motion planning data; Among them, the Monte Carlo algorithm and the deterministic policy gradient algorithm are used to solve and generate a game counter-strategy set, including: randomly sampling the state space and action space through the Monte Carlo algorithm, simulating the behavior of the intelligent agent and calculating the objective function value, and obtaining the initial solution through multiple sampling and calculations; based on the initial solution, the deterministic policy gradient algorithm is used to optimize the policy parameters and maximize the objective function until convergence or the stopping condition is reached; the optimal strategy in the policy parameter optimization process is extracted to construct a game counter-strategy set.
8. The system according to claim 7, characterized in that The environmental data includes underwater data and meteorological data; accordingly, The multi-agent environment perception module includes: The first acquisition submodule acquires the player data and underwater data based on vision, lidar sensors and sonar sensors; A second acquisition submodule is used to measure the attitude data of the multi-agent based on the inertial measurement unit, wherein the attitude data includes attitude, acceleration, angular velocity, roll and pitch; The third acquisition submodule is used to monitor the meteorological data of the target gaming area based on the environmental monitoring sensor, wherein the meteorological data includes wind speed, wind direction, temperature and humidity; The fourth acquisition submodule is used to acquire position data based on the global navigation and positioning unit.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Mobile edge network intelligent resource allocation method capable of dividing tasks
CN113873022A
Active power distribution network optimization scheduling method, system and device and storage medium
CN118017492A
Spacecraft cluster game hunting motion planning method based on multi-agent reinforcement learning
CN118056756A
Cited By
Ship unmanned aerial vehicle game guidance and control method for channel ice condition detection
CN121995959A
Sea-air cross-domain multi-agent layered collaborative task planning system
CN122195097A