A method for determining a wargaming strategy and related products

By acquiring tactical objectives based on battlefield situation information, designing the action space of combat units and optimizing strategies, the adaptability and collaborative decision-making problems of traditional wargaming strategies are solved, thereby improving the accuracy and intelligence level of wargaming.

CN121009998BActive Publication Date: 2026-01-23BAIYANG TIMES (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511551265.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-01-23
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

Traditional wargaming strategies lack self-learning and optimization capabilities, resulting in poor adaptability. Furthermore, single-agent reinforcement learning cannot model the collaborative relationships between agents, thus reducing the accuracy of decision-making.

Method used

By acquiring tactical objectives based on current battlefield situation information, action space design is performed for multiple combat units, the target actions of each combat unit are determined, and a two-stage attention mechanism and decision network optimization strategy are used to ensure that each step serves the tactical objective, thereby improving the accuracy of wargaming strategies.

Benefits of technology

It improves the task focus and accuracy of wargaming strategies, enhances the system's intelligence and versatility, and enables it to cope with diverse tactical tasks and emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009998B_ABST
    Figure CN121009998B_ABST
Patent Text Reader

Abstract

The application discloses a kind of wargame strategy determination method and related products.The scheme, obtain the tactical target corresponding to current battlefield situation information;Based on the tactical target, the action space of multiple combat units is designed respectively, the action set corresponding to each combat unit is obtained, and based on the action set corresponding to each combat unit, the target action of each combat unit is determined;The target action of each combat unit indicates the action in the action set corresponding to each combat unit, corresponding to current battlefield situation information;Based on the target action of each combat unit, wargame strategy is obtained.Compared with the precision of wargame strategy generated by intelligent decision method in prior art is low, the application has obvious advantages.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, in particular to a method for determining a war game strategy and related products. BACKGROUND

[0002] Traditional intelligent decision-making methods mainly include two types: rule-based decision-making method and single-agent reinforcement learning method. The rule-based decision-making method generates decisions based on predefined logical conditions, such as selecting a maneuvering path according to terrain types or assigning an attack target according to enemy forces. The single-agent reinforcement learning method optimizes strategies through trial and error learning. However, both methods have certain defects: the rule-based decision-making method does not have the ability to learn and optimize autonomously, and the improvement of strategies highly depends on the experience of domain experts, which cannot realize continuous evolution through interaction with the environment, resulting in poor adaptability and seriously affecting the accuracy of decisions; the single-agent reinforcement learning method cannot model the collaborative relationship between agents, making collaborative decision-making difficult and reducing the accuracy of decisions.

[0003] Therefore, how to improve the accuracy of war game strategies is a technical problem to be solved. SUMMARY

[0004] In view of the above problems, the present application provides a method for determining a war game strategy and related products, aiming to improve the accuracy of war game strategies.

[0005] The embodiments of the present application disclose the following technical solutions:

[0006] The first aspect of the present application provides a method for determining a war game strategy, which comprises:

[0007] obtaining a tactical target corresponding to current battlefield situation information;

[0008] designing an action space for each combat unit based on the tactical target, obtaining an action set corresponding to each combat unit, and determining a target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates an action in the action set corresponding to each combat unit corresponding to the current battlefield situation information;

[0009] obtaining a war game strategy based on the target action of each combat unit.

[0010] Optionally, after obtaining the war game strategy based on the target action of each combat unit, the method further comprises:

[0011] interacting with the current environment through the war game strategy to obtain a deduction result corresponding to the war game strategy and new battlefield situation information;

[0012] determine whether the deduction result corresponding to the Kriegsspiel strategy satisfies a preset ending condition based on the new battlefield situation information, to obtain a determination result;

[0013] If the determination result is yes, the Kriegsspiel strategy is quantified by using a two-stage attention mechanism to obtain a quantization result, and network parameters of a decision network are optimized based on the quantization result; the decision network is used to design an action space for each combat unit based on the tactical target to obtain an action set corresponding to each combat unit, and determine a target action of each combat unit based on the action set corresponding to each combat unit.

[0014] If the determination result is no, return to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset ending condition is satisfied.

[0015] Optionally, the Kriegsspiel strategy is quantified by using the two-stage attention mechanism to obtain the quantization result, including:

[0016] Based on the tactical target, a plurality of target combat units corresponding to the tactical target are selected from a plurality of combat units;

[0017] The contribution weight value of each target combat unit is calculated, and the action values of a plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain the quantization result.

[0018] The action value of the target action indicates a cumulative reward value calculated after a first target action is performed according to the Kriegsspiel strategy under the current battlefield situation; the first target action is any one of the plurality of target actions.

[0019] Optionally, the action values of the plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain the quantization result, including:

[0020] Based on the contribution weight value of each target combat unit, a hybrid network is used to weight and fuse the action values of the plurality of target actions corresponding to the plurality of target combat units to obtain the quantization result.

[0021] Optionally, the action space of the plurality of combat units is designed based on the tactical target to obtain an action set corresponding to each combat unit, including:

[0022] Based on the tactical target and the Kriegsspiel rule, the action space of the plurality of combat units is designed to obtain an action set corresponding to each combat unit;

[0023] The action space design includes a real-time instruction action space design and / or a planned instruction action space design; and the wargame rule includes at least one of a terrain rule, a stacking rule, and a line-of-sight rule.

[0024] Optionally, after obtaining the tactical target corresponding to the current battlefield situation information, the method further includes:

[0025] decomposing the tactical target into a plurality of subtasks;

[0026] setting a corresponding reward signal for each subtask to quantify, by using a credit distribution mechanism, a contribution made by each combat unit to the corresponding subtask based on the reward signal of each subtask.

[0027] The second aspect of the present application provides a wargame strategy determination device, which includes:

[0028] a tactical target acquisition module configured to acquire a tactical target corresponding to current battlefield situation information;

[0029] a target action determination module configured to design an action space for each combat unit based on the tactical target, to obtain an action set corresponding to each combat unit, and to determine a target action of each combat unit based on the action set corresponding to each combat unit, wherein the target action of each combat unit indicates an action corresponding to the current battlefield situation information in the action set corresponding to each combat unit;

[0030] a wargame strategy determination module configured to obtain a wargame strategy based on the target action of each combat unit.

[0031] Optionally, the device further includes a judgment module.

[0032] The judgment module is configured to interact with a current environment by using the wargame strategy to obtain a wargame result corresponding to the wargame strategy and new battlefield situation information.

[0033] Based on the new battlefield situation information, the judgment module is configured to judge whether the wargame result corresponding to the wargame strategy satisfies a preset ending condition to obtain a judgment result.

[0034] If the judgment result is yes, the wargame strategy is quantified by using a two-stage attention mechanism to obtain a quantification result, and network parameters of a decision network are optimized based on the quantification result, wherein the decision network is configured to design an action space for each combat unit based on the tactical target, to obtain an action set corresponding to each combat unit, and to determine a target action of each combat unit based on the action set corresponding to each combat unit.

[0035] If the determination result is no, the step of obtaining the tactical target corresponding to the current battlefield situation information is returned until the preset ending condition is met.

[0036] The third aspect of the present application provides a computer readable storage medium, wherein a computer program is stored in the computer readable storage medium, and when the program is run by a processor, the method for determining a war game strategy is realized as any implementation manner of the first aspect.

[0037] The fourth aspect of the present application provides a processor for running a computer program, and the program performs the method for determining a war game strategy as any implementation manner of the first aspect when running.

[0038] Compared with the prior art, the present application has the following beneficial effects:

[0039] The method for determining a war game strategy provided by the present application obtains a tactical target corresponding to current battlefield situation information; based on the tactical target, action space design is performed on a plurality of combat units respectively to obtain an action set corresponding to each combat unit, and based on the action set corresponding to each combat unit, a target action of each combat unit is determined; the target action of each combat unit indicates an action in the action set corresponding to each combat unit and corresponding to the current battlefield situation information; based on the target action of each combat unit, a war game strategy is obtained. The determination of the action set corresponding to each combat unit and the determination of the target action of each combat unit are closely related to the tactical target, which ensures that each step of determining the war game strategy serves the tactical target, improves the task focusing degree, and further improves the precision of the war game strategy. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0041] Figure 1 A flowchart of a method for determining a war game strategy provided by an embodiment of the present application;

[0042] Figure 2 A flowchart of another method for determining a war game strategy provided by an embodiment of the present application;

[0043] Figure 3 A structural schematic diagram of a war game strategy determination device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0044] As described above, current intelligent decision-making methods are mainly divided into two categories: rule-based decision-making methods and single-agent reinforcement learning methods. Rule-based decision-making methods generate decisions based on predefined logical conditions, such as selecting a maneuver path according to terrain types or assigning attack targets according to enemy forces. Single-agent reinforcement learning methods optimize strategies through trial and error learning. However, both methods have certain defects: rule-based decision-making methods lack autonomous learning and optimization capabilities, and strategy improvement is highly dependent on the experience of domain experts, which cannot achieve continuous evolution through interaction with the environment, resulting in poor adaptability and seriously affecting the accuracy of decisions; single-agent reinforcement learning methods cannot model the collaborative relationship between agents, making collaborative decision-making difficult and reducing the precision of decisions.

[0045] In view of the above problems, the present application provides a method for determining a war game strategy and related products. The method includes the following steps: obtaining a tactical target corresponding to current battlefield situation information; based on the tactical target, respectively designing an action space for a plurality of combat units to obtain an action set corresponding to each combat unit, and determining a target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates an action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information; and obtaining a war game strategy based on the target action of each combat unit. The determination of the action set corresponding to each combat unit and the determination of the target action of each combat unit are closely related to the tactical target, ensuring that each step of determining the war game strategy serves the tactical target, improving the task focus, and thus improving the precision of the war game strategy.

[0046] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0047] Referring to Figure 1 , the figure is a flowchart of a method for determining a war game strategy provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0048] S101, obtaining a tactical target corresponding to current battlefield situation information.

[0049] The battlefield situation information includes information such as terrain elevation, force distribution, and capture point occupation status.

[0050] The manner of obtaining the tactical target is not limited here, for example, the management layer in the hierarchical decision-making framework can formulate the tactical target based on the current battlefield situation information. Moreover, the tactical target here is a macro tactical target, for example, "occupy the commanding heights" or "flanking feint".

[0051] In S102, action space design is performed on the multiple combat units respectively based on the tactical target, to obtain an action set corresponding to each combat unit, and a target action of each combat unit is determined based on the action set corresponding to each combat unit.

[0052] The target action of each combat unit indicates an action corresponding to the current battlefield situation information in the action set corresponding to each combat unit.

[0053] The action space design and the determination of the target action of the target combat unit are both centered on the tactical target, which ensures that each step of determining the wargaming strategy serves the tactical target and improves the task focus.

[0054] By determining the target action of the target combat unit, interference of invalid actions is avoided, resources are concentrated on the target action, and the precision of the wargaming strategy is improved.

[0055] In S103, a wargaming strategy is obtained based on the target action of each combat unit.

[0056] The target actions distributed in the multiple combat units are integrated into a complete combat scheme that is global and executable, i.e., the wargaming strategy. The integration manner is not limited here.

[0057] The method for determining the wargaming strategy provided in the embodiments of the present application obtains a tactical target corresponding to the current battlefield situation information, performs action space design on multiple combat units respectively based on the tactical target, to obtain an action set corresponding to each combat unit, and determines a target action of each combat unit based on the action set corresponding to each combat unit. The target action of each combat unit indicates an action corresponding to the current battlefield situation information in the action set corresponding to each combat unit. A wargaming strategy is obtained based on the target action of each combat unit. The determination of the action set corresponding to each combat unit and the determination of the target action of each combat unit are both centered on the tactical target, which ensures that each step of determining the wargaming strategy serves the tactical target, improves the task focus, and further improves the precision of the wargaming strategy.

[0058] To further improve the method for determining the wargame strategy, on the basis of the above embodiment, the step of interacting the wargame strategy with the current environment to obtain a deduction result corresponding to the wargame strategy and new battlefield situation information, and judging whether the deduction result corresponding to the wargame strategy meets a preset ending condition based on the new battlefield situation information to obtain a judgment result is added.

[0059] Referring to Figure 2 FIG. 6 is a flowchart of another method for determining a wargame strategy provided by an embodiment of the present application. As shown in Figure 2 the method comprises the following steps:

[0060] S201: Obtain a tactical target corresponding to current battlefield situation information.

[0061] The battlefield situation information includes information such as terrain elevation, troop distribution, and occupation state of a control point.

[0062] The manner of obtaining the tactical target is not limited here, for example, a management layer in a hierarchical decision-making framework can formulate a tactical target based on current battlefield situation information. Moreover, the tactical target here is a macroscopic tactical target, for example, "occupy a commanding point" or "flanking feint" and the like.

[0063] S202: Based on the tactical target, respectively design an action space for a plurality of combat units to obtain an action set corresponding to each combat unit, and based on the action set corresponding to each combat unit, determine a target action of each combat unit.

[0064] The target action of each combat unit indicates an action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information.

[0065] For example, input the tactical target and observation information into a decision network to obtain a plurality of combat units, which are respectively a first infantry company, a second tank platoon, and an artillery group. The action set corresponding to the first infantry company is {concealment approach, fire suppression, launch a surprise attack}, the action set corresponding to the second tank platoon is {forward out of cover, pinpoint strike, retreat for supply}, and the action set corresponding to the artillery group is {high ground fire coverage, smoke cover, stop firing}. The target action of the first infantry company is determined to be concealment approach, the target action of the second tank platoon is determined to be forward out of cover, and the target action of the artillery group is determined to be high ground fire coverage.

[0066] In an implementable embodiment, based on the tactical target, the action space is designed for a plurality of combat units to obtain an action set corresponding to each combat unit, comprising:

[0067] Based on the tactical target and the wargame rule, an action space is designed for each combat unit, and a corresponding action set of each combat unit is obtained.

[0068] The action space design includes real-time instruction action space design and / or planned instruction action space design; and the wargame rule includes at least one of a terrain rule, a stacking rule and a line-of-sight rule.

[0069] The wargame rule includes a stacking rule and a line-of-sight rule, etc. The observation information is usually local observation information, such as the positions of enemies and allies within a local field of view.

[0070] The real-time instruction action space design is used to plan immediate response actions, and the real-time instruction includes direct fire and emergency maneuver, etc.

[0071] The planned instruction action space design is used to plan long-period response actions, and the planned instruction includes indirect fire and march path planning, etc.

[0072] In order to ensure that the actions of the combat units comply with the rules of the wargame, the system adopts a hierarchical processing mechanism of real-time / planned instructions, which accurately executes instructions such as direct fire and indirect fire, and improves the accuracy of the wargame strategy.

[0073] In addition, the determination process of the wargame strategy also introduces a dynamic action shielding mechanism, which can automatically shield invalid actions, and can be realized by a rule engine or a neural network mask, etc. For example, according to the line-of-sight rule, shooting instructions that cannot observe the target are automatically shielded; and according to the elevation difference maneuver penalty, invalid movement instructions with an elevation difference of more than 20 meters are automatically shielded. The dynamic action shielding mechanism can significantly improve the compliance and effectiveness of the actions.

[0074] Further, in order to make the wargame strategy closer to the actual battlefield situation, detailed terrain modeling rules are embedded in the state representation. For example, the maneuver speed is halved in the jungle, and the maneuver speed is halved when the elevation difference exceeds 20 meters. The embedding of the terrain modeling rule enables the determination process of the wargame strategy to fully consider the influence of terrain on the maneuver speed, thereby generating a more realistic wargame strategy and improving the reliability of the wargame strategy.

[0075] S203, based on the target action of each combat unit, a wargame strategy is obtained.

[0076] The target actions dispersed in each combat unit are integrated into a global and executable complete combat scheme, i.e., a wargame strategy. The integration method is not limited here.

[0077] S204, interacting with the current environment through the war game deduction strategy to obtain a deduction result corresponding to the war game deduction strategy and new battlefield situation information, and determining whether the deduction result corresponding to the war game deduction strategy satisfies a preset ending condition based on the new battlefield situation information to obtain a determination result.

[0078] If the determination result is yes, the war game deduction strategy is quantified using a two-stage attention mechanism to obtain a quantification result, and network parameters of a decision network are optimized based on the quantification result; the decision network is used to design an action space for each combat unit based on the tactical target to obtain an action set corresponding to each combat unit, and determine a target action of each combat unit based on the action set corresponding to each combat unit.

[0079] If the determination result is no, the step of obtaining the tactical target corresponding to the current battlefield situation information is returned until the preset ending condition is satisfied.

[0080] The preset ending condition is used to determine whether the current deduction is ended. The content of the preset ending condition is not limited here, for example, successfully capturing a target position.

[0081] Through the determination, the battlefield situation is dynamically evolved with the execution of the war game deduction strategy, providing a basis for subsequent strategy evaluation and network parameter optimization, supporting the continuous evolution of the model. Based on the determination result, corresponding measures are taken to build an iterative multi-stage decision cycle, improve the intelligent level and universality of the system, and enable it to cope with diversified tactical tasks and unexpected situations.

[0082] The decision network is obtained after training, and the network parameters of the decision network are updated through a reinforcement learning algorithm. The reinforcement learning algorithm includes a proximal policy optimization algorithm.

[0083] During the training of the decision network, a reward function is designed to guide the agent to take beneficial actions and avoid detrimental actions. Positive rewards include actions such as approaching a control point and suppressing enemy units, while negative penalties include actions such as terrain penalties and rule violations, such as rule stacking and entering a minefield. By setting sub-task rewards, dense reward signals can be provided before the control point is captured, thereby alleviating the problem of sparse rewards and helping the agent to learn and optimize strategies more effectively. This reward mechanism guides the agent to gradually achieve high-level goals by optimizing the rewards of sub-tasks, which helps to decompose complex global tasks into more manageable and optimized sub-tasks, thereby improving learning efficiency and strategy quality.

[0084] In an implementable embodiment, the war game deduction strategy is quantified using a two-stage attention mechanism to obtain a quantification result, including:

[0085] Based on the tactical target, a plurality of target combat units corresponding to the tactical target are selected from a plurality of combat units.

[0086] A contribution weight value of each target combat unit is calculated, and action values of a plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain a quantitative result.

[0087] The action value of the target action indicates a cumulative reward value calculated after the first target action is performed according to the war game strategy under the current battlefield situation; and the first target action is any one of the plurality of target actions.

[0088] In the first stage, irrelevant combat units, such as logistics units far away from the battlefield, are filtered out, and combat units that match the current tactical target are screened out as target combat units. This way can reduce unnecessary computational burden and concentrate resources on processing key combat units.

[0089] In the second stage, the contribution of the screened target combat units is quantified, and the weights of the target combat units are calculated. These weights reflect the actual contribution of each combat unit to the global target.

[0090] Weighted fusion of the action values of the target actions of the target combat units can more accurately measure the contribution differences of different combat units to the tactical target, thereby improving the accuracy of collaborative decision-making.

[0091] The action value is a quantitative result of predicting the cumulative tactical income that a combat unit can bring by performing a certain action under the current battlefield situation, and is used to guide the optimal action selection of the combat unit.

[0092] In an implementable embodiment:

[0093] Based on the contribution weight value of each target combat unit, the action values of a plurality of target actions corresponding to the plurality of target combat units are weighted and fused by using a hybrid network to obtain a quantitative result.

[0094] The use of a hybrid network for weighted fusion can solve the evaluation deviation caused by traditional linear superposition.

[0095] In an implementable embodiment, after obtaining the tactical target corresponding to the current battlefield situation information, the method further includes:

[0096] The tactical target is decomposed into a plurality of subtasks.

[0097] A corresponding reward signal is set for each subtask to quantify the contribution of each combat unit to the completion of the corresponding subtask based on the reward signal of each subtask by using a credit assignment mechanism.

[0098] Here, the setting principle of the reward signal in the credit assignment mechanism is not limited, for example, the reward signal is set based on the difficulty or urgency of the subtask.

[0099] By decomposing the tactical target into multiple subtasks and setting a corresponding reward signal for each subtask, the decision network can perform policy learning under the guidance of the clear subtask, thereby improving the learning efficiency and the quality of the war game strategy.

[0100] Another method for determining a war game strategy provided by an embodiment of the present application includes obtaining a tactical target corresponding to current battlefield situation information; based on the tactical target, respectively designing an action space for multiple combat units to obtain an action set corresponding to each combat unit, and determining a target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates an action in the action set corresponding to each combat unit corresponding to the current battlefield situation information; based on the target action of each combat unit, a war game strategy is obtained. The determination of the action set corresponding to each combat unit and the determination of the target action of each combat unit are closely related to the tactical target, ensuring that each step of determining the war game strategy serves the tactical target, improving the task focus, and further improving the precision of the war game strategy.

[0101] Further, the war game strategy is interacted with the current environment to obtain a deduction result corresponding to the war game strategy and new battlefield situation information, and based on the new battlefield situation information, it is determined whether the deduction result corresponding to the war game strategy satisfies a preset ending condition to obtain a determination result. Through the determination, the battlefield situation dynamically evolves with the execution of the war game strategy, providing a basis for subsequent strategy evaluation and network parameter optimization, supporting continuous optimization of the model.

[0102] Based on the method for determining a war game strategy introduced in the foregoing embodiments, correspondingly, the present application also provides a war game strategy determination device. Figure 3 The structure diagram of the device is shown in FIG. 1. Figure 3 As shown in FIG. 1, the war game strategy determination device includes:

[0103] The tactical target acquisition module 301 is configured to obtain a tactical target corresponding to current battlefield situation information.

[0104] The target action determination module 302 is configured to perform action space design on each of the combat units based on the tactical target, to obtain an action set corresponding to each of the combat units, and to determine a target action of each of the combat units based on the action set corresponding to each of the combat units; the target action of each of the combat units indicates an action corresponding to current battlefield situation information in the action set corresponding to each of the combat units.

[0105] The wargame strategy determination module 303 is configured to obtain a wargame strategy based on the target action of each of the combat units.

[0106] Optionally, the apparatus further includes a judgment module.

[0107] The judgment module is configured to obtain a wargame result corresponding to the wargame strategy and new battlefield situation information by interacting the wargame strategy with a current environment.

[0108] Based on the new battlefield situation information, it is determined whether the wargame result corresponding to the wargame strategy satisfies a preset ending condition, to obtain a judgment result.

[0109] If the judgment result is yes, the wargame strategy is quantified by using a two-stage attention mechanism to obtain a quantification result, and network parameters of a decision network are optimized based on the quantification result; the decision network is configured to perform action space design on each of the combat units based on the tactical target, to obtain an action set corresponding to each of the combat units, and to determine a target action of each of the combat units based on the action set corresponding to each of the combat units.

[0110] If the judgment result is no, the step of obtaining the tactical target corresponding to the current battlefield situation information is returned until the preset ending condition is satisfied.

[0111] Optionally, the quantification of the wargame strategy by using the two-stage attention mechanism to obtain the quantification result includes:

[0112] Based on the tactical target, a plurality of target combat units corresponding to the tactical target are selected from the plurality of combat units.

[0113] A contribution weight value of each target combat unit is calculated, and action values of a plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain the quantification result.

[0114] The action value of the target action indicates a cumulative reward value calculated after a first target action is performed according to the wargame strategy under a current battlefield situation; the first target action is any one of the plurality of target actions.

[0115] Optionally, the action values of the plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit to obtain a quantitative result, including:

[0116] The action values of the plurality of target actions corresponding to the plurality of target combat units are weighted and fused based on the contribution weight value of each target combat unit by using a hybrid network to obtain a quantitative result.

[0117] Optionally, the action space of the plurality of combat units is designed based on the tactical target to obtain an action set corresponding to each combat unit, including:

[0118] The action space of the plurality of combat units is designed based on the tactical target and the wargame rule to obtain an action set corresponding to each combat unit.

[0119] The action space design includes real-time instruction action space design and / or planned instruction action space design; and the wargame rule includes at least one of a terrain rule, a stacking rule, and a line-of-sight rule.

[0120] Optionally, after the tactical target corresponding to the current battlefield situation information is obtained, the method further includes:

[0121] The tactical target is decomposed into a plurality of subtasks.

[0122] A corresponding reward signal is set for each subtask, so as to quantify the contribution of each combat unit to the completion of the corresponding subtask based on the reward signal of each subtask by using a credit distribution mechanism.

[0123] In addition, an embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the program is run by a processor, a wargame strategy determination method as any of the method embodiments is implemented.

[0124] In addition, an embodiment of the present application further provides a processor, which is used to run a computer program. When the program is run, a wargame strategy determination method as any of the method embodiments is implemented.

[0125] It should be noted that each of the embodiments described in the specification of the present application adopts a progressive mode, and the same or similar parts between the embodiments can be mutually referred to. Each of the embodiments focuses on the differences from other embodiments. In particular, the device embodiments are described more simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the part of the method embodiments. The device embodiments described above are only illustrative, and the units described as separate components can or can not be physically separated, and the components indicated as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiments according to the actual needs. Those skilled in the art can understand and implement it without creative labor.

[0126] The above describes only one specific embodiment of the present application, but the protection scope of the present application is not limited to this. Any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for determining a wargaming strategy, characterized in that, include: Acquire tactical objectives corresponding to the current battlefield situation information; Based on the tactical objectives, action space design is performed for multiple combat units to obtain the action set corresponding to each combat unit. Based on the action set corresponding to each combat unit, the target action of each combat unit is determined. The target action of each combat unit indicates the action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information. Based on the target actions of each combat unit, a wargaming strategy is derived. After obtaining the wargaming strategy based on the target actions of each combat unit, the method further includes: By interacting with the current environment through the wargaming strategy, the simulation results and new battlefield situation information corresponding to the wargaming strategy are obtained; Based on the new battlefield situation information, determine whether the wargaming strategy corresponding to the wargaming result meets the preset termination condition, and obtain the determination result. If the judgment result is yes, then the wargaming strategy is quantified using a two-stage attention mechanism to obtain the quantification result, and the network parameters of the decision network are optimized based on the quantification result; the decision network is used to design the action space for multiple combat units based on the tactical objective, to obtain the action set corresponding to each combat unit, and to determine the target action of each combat unit based on the action set corresponding to each combat unit. If the judgment result is negative, the process returns to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset termination condition is met.

2. The method according to claim 1, characterized in that, The quantification of the wargaming strategy using a two-stage attention mechanism, to obtain the quantification result, includes: Based on the tactical objective, select multiple target combat units corresponding to the tactical objective from multiple combat units; Calculate the contribution weight value of each target combat unit, and based on the contribution weight value of each target combat unit, weightedly fuse the action values ​​of multiple target actions that correspond one-to-one with multiple target combat units to obtain a quantitative result; The action value of the target action refers to the cumulative reward value calculated after the first target action is executed according to the wargaming strategy under the current battlefield situation; the first target action is any one of the plurality of target actions.

3. The method according to claim 2, characterized in that, Based on the contribution weight value of each target combat unit, the action values ​​of multiple target actions corresponding one-to-one with multiple target combat units are weighted and fused to obtain a quantitative result, including: Based on the contribution weight value of each target combat unit, a hybrid network is used to weight and fuse the action values ​​of multiple target actions that correspond one-to-one with multiple target combat units to obtain a quantitative result.

4. The method according to claim 1, characterized in that, Based on the tactical objectives, the action space is designed for multiple combat units to obtain the action set corresponding to each combat unit, including: Based on the tactical objectives and wargaming rules, action spaces are designed for multiple combat units to obtain the action set corresponding to each combat unit. The action space design includes real-time command action space design and / or planned command action space design; the wargame simulation rules include at least one of terrain rules, stacking rules and line-of-sight rules.

5. The method according to claim 1, characterized in that, After obtaining the tactical targets corresponding to the current battlefield situation information, the process also includes: The tactical objective is broken down into multiple sub-tasks; A corresponding reward signal is set for each sub-task, so as to use the credit allocation mechanism to quantify the contribution of each combat unit in completing the corresponding sub-task based on the reward signal of each sub-task.

6. A device for determining a wargaming strategy, characterized in that, include: The tactical target acquisition module is used to acquire tactical targets corresponding to the current battlefield situation information; The target action determination module is used to design the action space for multiple combat units based on the tactical objective, obtain the action set corresponding to each combat unit, and determine the target action of each combat unit based on the action set corresponding to each combat unit; the target action of each combat unit indicates the action in the action set corresponding to each combat unit that corresponds to the current battlefield situation information. The wargaming strategy determination module is used to obtain the wargaming strategy based on the target actions of each combat unit. The device further includes: a judgment module; The judgment module is used to interact with the current environment through the wargaming strategy to obtain the simulation results and new battlefield situation information corresponding to the wargaming strategy. Based on the new battlefield situation information, determine whether the wargaming strategy corresponding to the wargaming result meets the preset termination condition, and obtain the determination result. If the judgment result is yes, then the wargaming strategy is quantified using a two-stage attention mechanism to obtain the quantification result, and the network parameters of the decision network are optimized based on the quantification result; the decision network designs the action space for multiple combat units based on the tactical objective to obtain the action set corresponding to each combat unit, and determines the target action of each combat unit based on the action set corresponding to each combat unit. If the judgment result is negative, the process returns to the step of obtaining the tactical target corresponding to the current battlefield situation information until the preset termination condition is met.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for determining a wargaming strategy as described in any one of claims 1-5.

8. A processor, characterized in that, Used to run a computer program, which, when running, executes the method for determining a wargaming strategy as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Knowledge-driven war game deduction intelligent decision-making method

    CN113435598A

  • Wargame agent auxiliary decision-making method based on situation awareness interaction

    CN118966356A