Wargame deduction strategy determination method and related product
By acquiring tactical objectives based on battlefield situation information, designing the action space of combat units and optimizing strategies, the problem of autonomous learning and collaborative decision-making in traditional wargaming strategies is solved, achieving higher precision and intelligence in wargaming.
Patent Information
- Application Number
- CN202511551265.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-10-28
AI Technical Summary
Traditional wargaming strategies lack self-learning and optimization capabilities, resulting in poor adaptability and an inability to effectively model the collaborative relationships between agents, thus affecting decision-making accuracy.
By acquiring tactical objectives based on current battlefield situation information, action space design is performed for multiple combat units, the target actions of each combat unit are determined, and a two-stage attention mechanism and decision network optimization strategy are used to ensure that each step serves the tactical objective. In combination with real-time and planned command design, terrain modeling and dynamic action masking mechanisms are introduced.
It improves the precision and task focus of wargaming strategies, enhances the accuracy of decision-making and the intelligence level of the system, and enables it to cope with diverse tactical tasks and emergencies.
Smart Images

Figure CN121009998A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a method for determining a war game strategy and related products. BACKGROUND
[0002] Traditional intelligent decision-making methods mainly include two types: rule-based decision-making method and single-agent reinforcement learning method. The rule-based decision-making method generates decisions based on predefined logical conditions, such as selecting a maneuvering path according to the terrain type or assigning an attack target according to the enemy forces. The single-agent reinforcement learning method optimizes the strategy through trial and error learning. However, both methods have certain defects: the rule-based decision-making method does not have the ability of autonomous learning and optimization, and the improvement of the strategy highly depends on the experience of domain experts, which cannot realize continuous evolution through interaction with the environment, resulting in poor adaptability and seriously affecting the accuracy of the decision; the single-agent reinforcement learning method cannot model the collaborative relationship between agents, which leads to difficulties in collaborative decision-making and reduces the accuracy of the decision.
[0003] Therefore, how to improve the accuracy of the war game strategy is a technical problem to be solved. SUMMARY
[0004] In view of the above problems, the present application provides a method for determining a war game strategy and related products, aiming to improve the accuracy of the war game strategy.
[0005] The embodiments of the present application disclose the following technical solutions: The first aspect of the present application provides a method for determining a war game strategy, which comprises: obtaining a tactical target corresponding to current battlefield situation information; designing an action space for each combat unit based on the tactical target, obtaining an action set corresponding to each combat unit, and determining a target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates an action in the action set corresponding to each combat unit corresponding to the current battlefield situation information; obtaining a war game strategy based on the target action of each combat unit.
[0006] Optionally, after obtaining the war game strategy based on the target action of each combat unit, the method further comprises: interacting with the current environment through the war game strategy to obtain a deduction result corresponding to the war game strategy and new battlefield situation information; judging whether the deduction result corresponding to the war game strategy meets a preset ending condition based on the new battlefield situation information to obtain a judgment result; If the determination result is yes, the two-stage attention mechanism is used to quantify the war game strategy to obtain a quantization result, and network parameters of a decision network are optimized based on the quantization result; the decision network is used to design an action space for each of a plurality of combat units based on the tactical target to obtain an action set corresponding to each combat unit, and determine a target action of each combat unit based on the action set corresponding to each combat unit. If the determination result is no, the step of obtaining the tactical target corresponding to the current battlefield situation information is returned until the preset ending condition is met.
[0007] Optionally, the two-stage attention mechanism is used to quantify the war game strategy to obtain a quantization result, comprising: Based on the tactical target, a plurality of target combat units corresponding to the tactical target are selected from a plurality of combat units; The contribution weight value of each target combat unit is calculated, and the action values of a plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain a quantization result; The action value of the target action indicates a cumulative reward value calculated after a first target action is executed according to the war game strategy under the current battlefield situation; the first target action is any one of the plurality of target actions.
[0008] Optionally, the plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain a quantization result, comprising: Based on the contribution weight value of each target combat unit, a hybrid network is used to weightedly fuse the action values of a plurality of target actions corresponding to the plurality of target combat units to obtain a quantization result.
[0009] Optionally, the action space of a plurality of combat units is designed based on the tactical target to obtain an action set corresponding to each combat unit, comprising: Based on the tactical target and the war game rules, the action space of a plurality of combat units is designed to obtain an action set corresponding to each combat unit; The action space design includes real-time instruction action space design and / or planned instruction action space design; the war game rules include at least one of terrain rules, stacking rules and line-of-sight rules.
[0010] Optionally, after the tactical target corresponding to the current battlefield situation information is obtained, the method further comprises: The tactical target is decomposed into a plurality of subtasks; A corresponding reward signal is set for each subtask to quantify the contribution of each combat unit to the completion of the corresponding subtask based on the reward signal of each subtask by using a credit assignment mechanism.
[0011] The second aspect of the present application provides a determination device of a war game strategy, the device comprising: A tactical target acquisition module is configured to acquire a tactical target corresponding to current battlefield situation information. A target action determination module is configured to design an action space for each combat unit based on the tactical target, to obtain an action set corresponding to each combat unit, and to determine a target action of each combat unit based on the action set corresponding to each combat unit, wherein the target action of each combat unit indicates an action corresponding to the current battlefield situation information in the action set corresponding to each combat unit. A war game strategy determination module is configured to obtain a war game strategy based on the target action of each combat unit.
[0012] Optionally, the device further comprises a judgment module. The judgment module is configured to interact with the current environment by using the war game strategy, to obtain a deduction result corresponding to the war game strategy and new battlefield situation information. Based on the new battlefield situation information, the judgment module is configured to judge whether the deduction result corresponding to the war game strategy meets a preset ending condition, to obtain a judgment result. If the judgment result is yes, the judgment module is configured to quantify the war game strategy by using a two-stage attention mechanism, to obtain a quantification result, and to optimize network parameters of a decision network based on the quantification result, wherein the decision network is configured to design an action space for each combat unit based on the tactical target, to obtain an action set corresponding to each combat unit, and to determine a target action of each combat unit based on the action set corresponding to each combat unit. If the judgment result is no, the judgment module is configured to return to the step of acquiring the tactical target corresponding to the current battlefield situation information until the preset ending condition is met.
[0013] The third aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and when the program is run by a processor, the determination method of the war game strategy is realized as any implementation manner of the first aspect.
[0014] The fourth aspect of the present application provides a processor for running a computer program, wherein the program is run to execute the determination method of the war game strategy as any implementation manner of the first aspect.
[0015] Compared with the prior art, the present application has the following beneficial effects: The method for determining wargaming strategies provided in this application involves acquiring tactical objectives corresponding to the current battlefield situation information; designing action spaces for multiple combat units based on the tactical objectives to obtain action sets for each combat unit; determining the target action for each combat unit based on its action sets; and identifying the action in each combat unit's action set that corresponds to the current battlefield situation information. Based on the target actions of each combat unit, a wargaming strategy is obtained. The determination of both the action sets and the target actions of the combat units closely revolve around the tactical objectives, ensuring that every step in determining the wargaming strategy serves the tactical objectives, improving task focus, and thus enhancing the accuracy of the wargaming strategy. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating a method for determining a wargaming strategy provided in this application embodiment; Figure 2 A flowchart illustrating another method for determining a wargaming strategy provided in this application embodiment; Figure 3 This is a schematic diagram of a device for determining a wargaming strategy, provided in an embodiment of this application. Detailed Implementation
[0018] As described earlier, current intelligent decision-making methods are mainly divided into two categories: rule-based decision-making methods and single-agent reinforcement learning methods. Rule-based decision-making methods generate decisions based on predefined logical conditions, such as selecting a maneuver path based on terrain type or allocating attack targets based on enemy troop strength. Single-agent reinforcement learning methods optimize strategies through trial and error. However, both types of methods have certain drawbacks: rule-based decision-making methods lack autonomous learning and optimization capabilities, and strategy improvement heavily relies on the experience of domain experts, making it impossible to achieve continuous evolution through interaction with the environment, resulting in poor adaptability and severely affecting the accuracy of decisions; single-agent reinforcement learning methods cannot model the cooperative relationships between agents, leading to difficulties in collaborative decision-making and reducing the accuracy of decisions.
[0019] In view of the above problems, this application proposes a method and related products for determining wargaming strategies. The method involves acquiring tactical objectives corresponding to the current battlefield situation information; designing action spaces for multiple combat units based on the tactical objectives to obtain action sets for each combat unit; and determining the target action for each combat unit based on its action sets. The target action for each combat unit indicates the action within its action set that corresponds to the current battlefield situation information. Based on the target actions of each combat unit, a wargaming strategy is obtained. The determination of both the action sets and the target actions of the combat units closely revolve around the tactical objectives, ensuring that every step in determining the wargaming strategy serves the tactical objectives, improving task focus, and thus enhancing the accuracy of the wargaming strategy.
[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0021] See Figure 1 The figure is a flowchart of a method for determining a wargaming strategy according to an embodiment of this application. Figure 1 As shown, the method includes the following steps: S101. Obtain the tactical objectives corresponding to the current battlefield situation information.
[0022] Battlefield situation information includes information such as terrain elevation, troop distribution, and the status of captured and controlled points.
[0023] The method of acquiring tactical objectives is not limited here. For example, tactical objectives can be formulated by the management level in a hierarchical decision-making framework based on the current battlefield situation information. Moreover, the tactical objectives here are macro-level tactical objectives, such as "seizing high ground" or "flanking maneuver".
[0024] S102. Based on the tactical objectives, the action space of multiple combat units is designed to obtain the action set corresponding to each combat unit, and the target action of each combat unit is determined based on the action set corresponding to each combat unit.
[0025] The target action of each combat unit refers to the action corresponding to the current battlefield situation information in the action set corresponding to each combat unit.
[0026] The action space design and the determination of target combat unit target actions all revolve around tactical objectives, which can ensure that every step in determining the wargaming strategy serves the tactical objectives and improve mission focus.
[0027] By defining the target actions of the target combat units, interference from ineffective actions is avoided, resources are concentrated on the target actions, and the accuracy of wargaming strategies is improved.
[0028] S103. Based on the target actions of each combat unit, a wargaming strategy is obtained.
[0029] The strategy involves integrating the target actions scattered across various combat units into a comprehensive and executable operational plan, known as wargaming. No restrictions are placed on the integration method here.
[0030] The method for determining wargaming strategies provided in this application involves acquiring tactical objectives corresponding to the current battlefield situation information; designing action spaces for multiple combat units based on the tactical objectives to obtain action sets for each combat unit; determining the target action for each combat unit based on the action sets; the target action for each combat unit indicates the action in its action set that corresponds to the current battlefield situation information; and obtaining the wargaming strategy based on the target actions of each combat unit. The determination of the action sets corresponding to combat units and the determination of the target actions of combat units are both closely aligned with the tactical objectives, ensuring that every step in determining the wargaming strategy serves the tactical objectives, improving task focus, and thus enhancing the accuracy of the wargaming strategy.
[0031] To further improve the method for determining wargaming strategies, based on the above embodiments, an additional step is added: by interacting with the current environment through the wargaming strategy, the simulation result corresponding to the wargaming strategy and new battlefield situation information are obtained; and based on the new battlefield situation information, it is determined whether the simulation result corresponding to the wargaming strategy meets the preset termination conditions, and the determination result is obtained.
[0032] See Figure 2 This figure is a flowchart of another method for determining a wargaming strategy provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps: S201. Obtain tactical objectives corresponding to the current battlefield situation information.
[0033] Battlefield situation information includes information such as terrain elevation, troop distribution, and the status of captured and controlled points.
[0034] The method of acquiring tactical objectives is not limited here. For example, tactical objectives can be formulated by the management level in a hierarchical decision-making framework based on the current battlefield situation information. Moreover, the tactical objectives here are macro-level tactical objectives, such as "seizing high ground" or "flanking maneuver".
[0035] S202. Based on the tactical objectives, the action space of multiple combat units is designed to obtain the action set corresponding to each combat unit, and the target action of each combat unit is determined based on the action set corresponding to each combat unit.
[0036] The target action of each combat unit refers to the action corresponding to the current battlefield situation information in the action set corresponding to each combat unit.
[0037] For example, the tactical objectives and observation information are input into a decision network to obtain multiple combat units, namely the 1st Infantry Company, the 2nd Tank Platoon, and the Artillery Group. The action set corresponding to the 1st Infantry Company is {concealed approach, fire suppression, launch an assault}, the action set corresponding to the 2nd Tank Platoon is {forward cover, pinpoint strike, retreat and resupply}, and the action set corresponding to the Artillery Group is {fire coverage of the high ground, smoke cover, cease fire}. The target action of the 1st Infantry Company is determined to be concealed approach, the target action of the 2nd Tank Platoon is forward cover, and the target action of the Artillery Group is fire coverage of the high ground.
[0038] In one feasible implementation, based on the tactical objective, action space design is performed for multiple combat units to obtain a set of actions corresponding to each combat unit, including: Based on the aforementioned tactical objectives and wargaming rules, action spaces are designed for multiple combat units to obtain the action set corresponding to each combat unit.
[0039] The action space design includes real-time command action space design and / or planned command action space design; the wargame simulation rules include at least one of terrain rules, stacking rules and line-of-sight rules.
[0040] Wargaming rules include stacking rules and line-of-sight rules. Observational information is usually local observational information, such as the positions of enemy and friendly forces within a local field of view.
[0041] The real-time command action space is designed to plan immediate response actions, including direct-fire and emergency maneuvers.
[0042] The planned instruction action space design is used to plan long-cycle response actions, including indirect artillery fire and marching route planning.
[0043] To ensure that the actions of combat units conform to the rules of wargaming, the system adopts a hierarchical processing mechanism for real-time / planned commands. This mechanism ensures the precise execution of commands such as direct fire and indirect artillery fire, thereby improving the accuracy of wargaming strategies.
[0044] Furthermore, the strategy determination process in wargaming incorporates a dynamic action masking mechanism. This mechanism automatically filters out invalid actions, which can be implemented through rule engines or neural network masks. For example, based on line-of-sight rules, it automatically masks firing commands for targets that cannot be observed; based on elevation difference penalties, it automatically masks invalid movement commands with elevation differences exceeding 20 meters. This dynamic action masking mechanism significantly improves the compliance and effectiveness of actions.
[0045] Furthermore, to make the wargaming strategies more closely resemble actual battlefield conditions, detailed terrain modeling rules were embedded in the state representation. For example, maneuver speed is halved in jungle terrain and when the elevation difference exceeds 20 meters. The embedding of terrain modeling rules allows the wargaming strategy determination process to fully consider the impact of terrain on maneuver speed, thereby generating more realistic wargaming strategies and improving their reliability.
[0046] S203. Based on the target actions of each combat unit, a wargaming strategy is obtained.
[0047] The strategy involves integrating the target actions scattered across various combat units into a comprehensive and executable operational plan, known as wargaming. No restrictions are placed on the integration method here.
[0048] S204. By interacting with the current environment through the wargaming strategy, the simulation results and new battlefield situation information corresponding to the wargaming strategy are obtained. Based on the new battlefield situation information, it is determined whether the simulation results corresponding to the wargaming strategy meet the preset termination conditions, and the determination result is obtained.
[0049] If the judgment result is yes, then the wargaming strategy is quantified using a two-stage attention mechanism to obtain the quantification result, and the network parameters of the decision network are optimized based on the quantification result; the decision network is used to design the action space for multiple combat units based on the tactical objective, to obtain the action set corresponding to each combat unit, and to determine the target action of each combat unit based on the action set corresponding to each combat unit.
[0050] If the judgment result is negative, the process returns to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset termination condition is met.
[0051] The preset termination condition is used to determine whether the current simulation has ended. There are no restrictions on the content of the preset termination condition; for example, it could mean successfully capturing the target position.
[0052] By making judgments, the battlefield situation was dynamically evolved in accordance with the execution of wargaming strategies, providing a foundation for subsequent strategy evaluation and network parameter optimization, and supporting the continuous evolution of the model. Based on the judgment results, corresponding measures were taken, constructing an iterative multi-stage decision-making loop, which improved the system's intelligence and versatility, enabling it to cope with diverse tactical tasks and contingencies.
[0053] The decision network is obtained after training, and its parameters are updated using reinforcement learning algorithms. These reinforcement learning algorithms include proximal policy optimization algorithms.
[0054] During the training of the decision network, a reward function is designed to guide the agent to take advantageous actions and avoid disadvantageous ones. Positive rewards include actions such as approaching control points and suppressing enemy units, while negative penalties include terrain penalties and violation penalties, such as illegal stacking and entering minefields. By setting sub-task rewards, dense reward signals can be provided before control points are captured, thereby alleviating the sparse reward problem and helping the agent learn and optimize policies more effectively. This reward mechanism guides the agent to gradually achieve higher-level goals by optimizing sub-task rewards, which helps to decompose complex global tasks into more manageable and optimizable sub-tasks, thereby improving learning efficiency and policy quality.
[0055] In one feasible implementation, a two-stage attention mechanism is used to quantify the wargaming strategy to obtain a quantification result, including: Based on the tactical objective, multiple target combat units corresponding to the tactical objective are selected from multiple combat units.
[0056] Calculate the contribution weight value of each target combat unit. Based on the contribution weight value of each target combat unit, the action values of multiple target actions corresponding to multiple target combat units are weighted and fused to obtain a quantitative result.
[0057] Wherein, the action value of the target action refers to the cumulative reward value calculated after the first target action is executed according to the wargaming strategy under the current battlefield situation; the first target action is any one of the plurality of target actions.
[0058] In the first phase, irrelevant combat units, such as logistics units far from the battlefield, were filtered out, and combat units that were more compatible with the current tactical objectives were selected as target combat units. This approach can reduce unnecessary computational burden and concentrate resources on processing key combat units.
[0059] In the second phase, the contributions of the selected target combat units are quantified, and the weight of each target combat unit is calculated. These weights reflect the actual contribution of each combat unit to the overall objective.
[0060] By weighted and fused together the action value of target combat units, the differences in contribution of different combat units to tactical objectives can be measured more accurately, thereby improving the accuracy of collaborative decision-making.
[0061] Action value is a quantitative result that predicts the cumulative tactical benefits that a combat unit can gain by performing a certain action under the current battlefield situation, and is used to guide the optimal action selection of the combat unit.
[0062] In one feasible implementation: Based on the contribution weight value of each target combat unit, a hybrid network is used to weight and fuse the action values of multiple target actions that correspond one-to-one with multiple target combat units to obtain a quantitative result.
[0063] Among them, using a hybrid network for weighted fusion can solve the evaluation bias caused by traditional linear superposition.
[0064] In one feasible implementation, after acquiring the tactical target corresponding to the current battlefield situation information, the method further includes: The tactical objective is broken down into multiple sub-tasks.
[0065] A corresponding reward signal is set for each sub-task, so as to use the credit allocation mechanism to quantify the contribution of each combat unit in completing the corresponding sub-task based on the reward signal of each sub-task.
[0066] There are no restrictions on the principles for setting reward signals in the credit allocation mechanism here. For example, reward signals can be set based on principles such as the difficulty or urgency of sub-tasks.
[0067] By decomposing tactical objectives into multiple sub-tasks and setting corresponding reward signals for each sub-task, the decision network can learn strategies under the guidance of clear sub-tasks, thereby improving learning efficiency and the quality of wargaming strategies.
[0068] This application provides another method for determining a wargaming strategy, which involves acquiring tactical objectives corresponding to the current battlefield situation information; designing action spaces for multiple combat units based on the tactical objectives to obtain action sets for each combat unit; determining the target action for each combat unit based on the action sets; the target action for each combat unit indicates the action in its action set that corresponds to the current battlefield situation information; and obtaining a wargaming strategy based on the target actions of each combat unit. The determination of the action sets corresponding to combat units and the determination of the target actions of combat units are both closely related to the tactical objectives, ensuring that every step in determining the wargaming strategy serves the tactical objectives, improving task focus, and thus enhancing the accuracy of the wargaming strategy.
[0069] Furthermore, by interacting with the current environment through the wargaming strategy, the simulation results corresponding to the strategy and new battlefield situation information are obtained. Based on the new battlefield situation information, it is determined whether the simulation results corresponding to the strategy meet the preset termination conditions, and a judgment result is obtained. Through this judgment, the battlefield situation is dynamically evolved as the wargaming strategy is executed, providing a foundation for subsequent strategy evaluation and network parameter optimization, and supporting continuous model optimization.
[0070] Based on the method for determining wargaming strategies described in the preceding embodiments, this application also provides a device for determining wargaming strategies. Figure 3 This is a schematic diagram of the device. Figure 3 As shown, the device for determining the wargaming strategy includes: The tactical target acquisition module 301 is used to acquire tactical targets corresponding to the current battlefield situation information.
[0071] The target action determination module 302 is used to design the action space for multiple combat units based on the tactical target, obtain the action set corresponding to each combat unit, and determine the target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates the action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information.
[0072] The wargaming strategy determination module 303 is used to obtain a wargaming strategy based on the target actions of each combat unit.
[0073] Optionally, the device further includes: a determination module; The judgment module is used to interact with the current environment through the wargaming strategy to obtain the simulation results and new battlefield situation information corresponding to the wargaming strategy.
[0074] Based on the new battlefield situation information, it is determined whether the simulation result corresponding to the wargaming strategy meets the preset termination condition, and the determination result is obtained.
[0075] If the judgment result is yes, then the wargaming strategy is quantified using a two-stage attention mechanism to obtain the quantification result, and the network parameters of the decision network are optimized based on the quantification result; the decision network is used to design the action space for multiple combat units based on the tactical objective, to obtain the action set corresponding to each combat unit, and to determine the target action of each combat unit based on the action set corresponding to each combat unit.
[0076] If the judgment result is negative, the process returns to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset termination condition is met.
[0077] Optionally, the quantification of the wargaming strategy using a two-stage attention mechanism to obtain the quantification result includes: Based on the tactical objective, multiple target combat units corresponding to the tactical objective are selected from multiple combat units.
[0078] Calculate the contribution weight value of each target combat unit. Based on the contribution weight value of each target combat unit, the action values of multiple target actions corresponding to multiple target combat units are weighted and fused to obtain a quantitative result.
[0079] The action value of the target action refers to the cumulative reward value calculated after the first target action is executed according to the wargaming strategy under the current battlefield situation; the first target action is any one of the plurality of target actions.
[0080] Optionally, based on the contribution weight value of each target combat unit, the action values of multiple target actions corresponding one-to-one with multiple target combat units are weighted and fused to obtain a quantitative result, including: Based on the contribution weight value of each target combat unit, a hybrid network is used to weight and fuse the action values of multiple target actions that correspond one-to-one with multiple target combat units to obtain a quantitative result.
[0081] Optionally, the step of designing the action space for multiple combat units based on the tactical objective to obtain the action set corresponding to each combat unit includes: Based on the aforementioned tactical objectives and wargaming rules, action spaces are designed for multiple combat units to obtain the action set corresponding to each combat unit.
[0082] The action space design includes real-time command action space design and / or planned command action space design; the wargame simulation rules include at least one of terrain rules, stacking rules and line-of-sight rules.
[0083] Optionally, after obtaining the tactical target corresponding to the current battlefield situation information, the method further includes: The tactical objective is broken down into multiple sub-tasks.
[0084] A corresponding reward signal is set for each sub-task, so as to use the credit allocation mechanism to quantify the contribution of each combat unit in completing the corresponding sub-task based on the reward signal of each sub-task.
[0085] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for determining a wargaming strategy as described in any of the method embodiments.
[0086] Furthermore, this application embodiment also provides a processor for running a computer program, which executes a method for determining a wargaming strategy as described in any of the foregoing method embodiments.
[0087] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. The components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment solution according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0088] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for determining a wargaming strategy, characterized in that, include: Acquire tactical objectives corresponding to the current battlefield situation information; Based on the tactical objectives, action space design is performed for multiple combat units to obtain the action set corresponding to each combat unit. Based on the action set corresponding to each combat unit, the target action of each combat unit is determined. The target action of each combat unit indicates the action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information. Based on the target actions of each combat unit, a wargaming strategy is derived.
2. The method according to claim 1, characterized in that, After obtaining the wargaming strategy based on the target actions of each combat unit, the method further includes: By interacting with the current environment through the wargaming strategy, the simulation results and new battlefield situation information corresponding to the wargaming strategy are obtained; Based on the new battlefield situation information, determine whether the wargaming strategy corresponding to the wargaming result meets the preset termination condition, and obtain the determination result. If the judgment result is yes, then the wargaming strategy is quantified using a two-stage attention mechanism to obtain the quantification result, and the network parameters of the decision network are optimized based on the quantification result; the decision network is used to design the action space for multiple combat units based on the tactical objective, to obtain the action set corresponding to each combat unit, and to determine the target action of each combat unit based on the action set corresponding to each combat unit. If the judgment result is negative, the process returns to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset termination condition is met.
3. The method according to claim 2, characterized in that, The quantification of the wargaming strategy using a two-stage attention mechanism, to obtain the quantification result, includes: Based on the tactical objective, select multiple target combat units corresponding to the tactical objective from multiple combat units; Calculate the contribution weight value of each target combat unit, and based on the contribution weight value of each target combat unit, weightedly fuse the action values of multiple target actions that correspond one-to-one with multiple target combat units to obtain a quantitative result; The action value of the target action refers to the cumulative reward value calculated after the first target action is executed according to the wargaming strategy under the current battlefield situation; the first target action is any one of the plurality of target actions.
4. The method according to claim 3, characterized in that, Based on the contribution weight value of each target combat unit, the action values of multiple target actions corresponding one-to-one with multiple target combat units are weighted and fused to obtain a quantitative result, including: Based on the contribution weight value of each target combat unit, a hybrid network is used to weight and fuse the action values of multiple target actions that correspond one-to-one with multiple target combat units to obtain a quantitative result.
5. The method according to claim 1, characterized in that, Based on the tactical objectives, the action space is designed for multiple combat units to obtain the action set corresponding to each combat unit, including: Based on the tactical objectives and wargaming rules, action spaces are designed for multiple combat units to obtain the action set corresponding to each combat unit. The action space design includes real-time command action space design and / or planned command action space design; the wargame simulation rules include at least one of terrain rules, stacking rules and line-of-sight rules.
6. The method according to claim 1, characterized in that, After obtaining the tactical targets corresponding to the current battlefield situation information, the process also includes: The tactical objective is broken down into multiple sub-tasks; A corresponding reward signal is set for each sub-task, so as to use the credit allocation mechanism to quantify the contribution of each combat unit in completing the corresponding sub-task based on the reward signal of each sub-task.
7. A device for determining a wargaming strategy, characterized in that, include: The tactical target acquisition module is used to acquire tactical targets corresponding to the current battlefield situation information; The target action determination module is used to design the action space for multiple combat units based on the tactical objective, obtain the action set corresponding to each combat unit, and determine the target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates the action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information. The wargaming strategy determination module is used to obtain the wargaming strategy based on the target actions of each combat unit.
8. The apparatus according to claim 7, characterized in that, The device further includes: a judgment module; The judgment module is used to interact with the current environment through the wargaming strategy to obtain the simulation results and new battlefield situation information corresponding to the wargaming strategy. Based on the new battlefield situation information, determine whether the wargaming strategy corresponding to the wargaming result meets the preset termination condition, and obtain the determination result. If the judgment result is yes, then the wargaming strategy is quantified using a two-stage attention mechanism to obtain the quantification result, and the network parameters of the decision network are optimized based on the quantification result; the decision network designs the action space for multiple combat units based on the tactical objective to obtain the action set corresponding to each combat unit, and determines the target action of each combat unit based on the action set corresponding to each combat unit. If the judgment result is negative, the process returns to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset termination condition is met.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for determining a wargaming strategy as described in any one of claims 1-6.
10. A processor, characterized in that, Used to run a computer program, which, when running, executes the method for determining a wargaming strategy as described in any one of claims 1-6.
Citation Information
Patent Citations
Knowledge-driven war game deduction intelligent decision-making method
CN113435598A
Intelligent war game deduction decision-making method based on deep reinforcement learning
CN116596343A
Wargame agent auxiliary decision-making method based on situation awareness interaction
CN118966356A
Intelligent war game deduction method based on reinforcement learning
CN120449646A
Apparatus and System to Counter Drones Using a Shoulder-Launched Aerodynamically Guided Missile
US20170307334A1