A method, device, and program product for decision-making in unmanned swarm combat.
By employing a game-theoretic, hierarchical reinforcement learning architecture in unmanned swarm combat, and combining information interaction between long-term planning and short-term decision-making layers, the problem of low decision-making efficiency in unmanned swarm combat is solved, achieving consistency between long-term strategy and short-term actions, and globally optimal decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies make it difficult to establish effective information exchange between long-term planning and short-term decision-making in unmanned swarm combat, resulting in low decision-making efficiency and inconsistency between long-term planning and short-term actions.
A pre-defined game-enhanced hierarchical reinforcement learning architecture is adopted. Combat mission instructions are issued through the long-term planning layer. Combined with the short-term combat status data of the unmanned swarm, the deep reinforcement learning algorithm is used to output real-time tactical actions, and the long-term plan is adjusted by evaluating the global mission progress.
It enables information exchange between long-term planning and short-term decision-making in unmanned swarm combat, improves decision-making efficiency, ensures consistency between long-term strategic goals and short-term actions, and achieves optimal overall decision-making results.
Smart Images

Figure CN121348781B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned swarm control technology, and in particular to a method, device and program product for unmanned swarm combat decision-making. Background Technology
[0002] With the maturation of unmanned equipment technology and the increasing complexity of combat scenarios, achieving efficient decision-making has become one of the core research directions. However, existing technologies struggle to establish an effective information exchange mechanism between long-term planning and short-term decision-making. On the one hand, during the long-term planning phase, when formulating missions and strategies for the combat phase, it is often impossible to dynamically obtain detailed information on real-time changes in the battlefield, making it difficult to flexibly adjust the planning direction according to the current combat progress. This makes the planning easily disconnected from actual battlefield needs, failing to provide accurate and suitable guidance for short-term decision-making.
[0003] On the other hand, when generating immediate tactical actions, short-term decision-making relies heavily on local, real-time battlefield data and lacks effective perception of overall long-term planning goals and overall mission progress. This may result in short-term actions that only meet local combat needs but are inconsistent with long-term strategic goals, failing to achieve optimal decision-making results globally. Summary of the Invention
[0004] The embodiments of the present invention provide a method, apparatus and program product for decision-making in unmanned swarm combat, which aims to solve the problem that the existing technology is difficult to establish information interaction between long-term planning and short-term decision-making, resulting in low efficiency of decision-making in unmanned swarm combat.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a decision-making method for unmanned swarm combat, comprising the following steps:
[0006] The long-term planning layer of the hierarchical reinforcement learning architecture, enhanced by pre-defined game theory, issues operational mission instructions for the current operational phase; these operational mission instructions include mission priority, resource allocation instructions, and / or adversarial instructions.
[0007] According to the operational mission instructions, each unmanned combat unit in the unmanned swarm is deployed to conduct real-time detection of the combat target to obtain short-term operational status data of the unmanned swarm; the short-term operational status data includes obstacle information, target damage status, location of friendly units and / or enemy threat assessment.
[0008] The short-term combat status data is input into the short-term decision layer of a predetermined game-enhanced hierarchical reinforcement learning architecture, and the short-term decision layer outputs the real-time tactical actions of each unmanned combat unit through a predetermined deep reinforcement learning algorithm.
[0009] After completing the current operational phase's operational mission instructions by executing the aforementioned real-time tactical actions, the global mission progress status is assessed. The global mission progress status includes mission completion progress, remaining resources, changes in enemy deployment and / or game model, and operational time.
[0010] The global task progress status is transmitted to the long-term planning layer, which then adjusts the long-term planning actions using a predetermined reinforcement learning algorithm based on game theory-based enhanced strategy gradients.
[0011] Furthermore, the deep reinforcement learning algorithm includes a predetermined short-term decision objective function, which is calculated using the following formula:
[0012] ,
[0013] ,
[0014] In the formula, Indicating short-term strategy The expected reward for performing a given immediate tactical action; Let it be the expected function; Indicating short-term strategy strategy entropy, The weighting coefficients for the policy entropy; Indicates a short-term reward; A successful dodge scores a point. To avoid the weighting coefficient of successful scores; Indicates the target strike effect score. The weighting coefficient for the target strike effect score; This represents the penalty value for task failure. The weighting coefficient for the task failure penalty value; This represents the simulated value of enemy interference. The weighting coefficients are the simulated values of enemy interference.
[0015] Furthermore, the short-term decision objective function updates its parameters using the following formula:
[0016] ,
[0017] ,
[0018] In the formula, It is a short-term decision objective function Relative to short-term policy network parameters The gradient; Indicating short-term strategy In short-term combat status Immediate tactical actions The logarithmic probability gradient; The dominant function; Let be the state value function, representing the expected estimate of the overall value of the short-term operational state s based on the current strategy; For short-term combat status Immediate tactical actions The value function; Discount factor; Indicates the next short-term operational status Maximize the value of real-time tactical actions The value of.
[0019] Furthermore, the reinforcement learning algorithm based on game-theoretic enhancement policy gradient includes a predetermined long-term policy objective function, which is calculated using the following formula:
[0020] ,
[0021] ,
[0022] In the formula, For long-term strategy The expected return for performing a given long-term planned action; For long-term total rewards; For cluster task completion rate, This is a weighting coefficient for the cluster task completion rate; Due to resource scheduling deviation, This represents the weighting coefficient for resource scheduling deviation; To gain an advantage in the game and score points. The weighting coefficients for the game advantage score; This represents the cluster loss value. The weighting coefficients for cluster loss values.
[0023] Furthermore, the long-term strategy objective function also updates its parameters using the following formula:
[0024] ,
[0025] In the formula, It is the long-term decision objective function Relative to long-term policy network parameters The gradient; Indicating long-term strategy In the global task progress status Long-term planning actions The logarithmic probability gradient; In the postwar state Long-term planning actions The advantage function.
[0026] Furthermore, the short-term rewards include path offset, tactical navigation rewards, action penalties, target recognition rewards, and / or enemy interference penalties.
[0027] Furthermore, the path offset is calculated using the following formula:
[0028] ,
[0029] In the formula, Indicates the path offset; Indicates the closest distance to an enemy target / obstacle; Indicates the distance to the obstacle; Indicates the distance deviating from the tactical route; This indicates a deviation in heading.
[0030] Furthermore, the short-term decision-making layer uploads the short-term operational status data to the long-term planning layer every 'a' seconds; the long-term planning layer issues the operational mission instructions to the short-term decision-making layer every 'b' minutes.
[0031] Secondly, the present invention provides an unmanned swarm combat decision-making device, including a memory and a processor, wherein the memory stores at least one program, and the at least one program is executed by the processor to implement the unmanned swarm combat decision-making method as described above.
[0032] Thirdly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the unmanned swarm combat decision-making method as described above.
[0033] The above technical solution has the following technical effects:
[0034] The invention addresses the problem of low decision-making efficiency in unmanned swarm combat operations caused by the difficulty in establishing information exchange between long-term planning and short-term decision-making in existing technologies. It involves issuing operational mission instructions for the current operational phase through a predetermined long-term planning layer; deploying unmanned combat units within the swarm to conduct real-time detection of operational targets to obtain short-term operational status data; inputting this short-term operational status data into a predetermined short-term decision-making layer, which then outputs immediate tactical actions for each unmanned combat unit; and evaluating the overall mission progress status after executing these immediate tactical actions, transmitting the data to the long-term planning layer for adjustments to long-term planned actions. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating a decision-making method for unmanned swarm combat according to an embodiment of the present invention.
[0036] Figure 2 This is a schematic diagram of the structure of an unmanned swarm combat decision-making device according to an embodiment of the present invention. Detailed Implementation
[0037] To further illustrate the various embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, primarily used to illustrate the embodiments and to explain the operating principles of the embodiments in conjunction with the relevant descriptions in the specification. With reference to these drawings, those skilled in the art should be able to understand other possible implementations and the advantages of the present invention. Components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0038] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0039] Example 1:
[0040] Figure 1 This is a flowchart illustrating a decision-making method for unmanned swarm combat according to an embodiment of the present invention, as shown below. Figure 1 As shown, this embodiment constructs a hierarchical reinforcement learning architecture for game-theoretic enhancement, comprising a short-term decision-making layer responsible for individual immediate local decisions and a long-term planning layer that performs long-term planning based on the overall combat situation and the enemy's game model, including the following steps:
[0041] The operational mission instructions for the current operational phase are issued through a predetermined long-term planning layer based on the overall operational objectives; in one specific implementation, the operational mission instructions include mission priorities such as priority destruction of air defense systems, resource allocation instructions, and / or countermeasure instructions.
[0042] According to the operational mission instructions, each unmanned combat unit in the unmanned swarm is deployed to conduct real-time detection of the combat target to obtain short-term operational status data of the unmanned swarm; in one specific implementation, the unmanned combat unit refers to unmanned equipment with independent mission execution capabilities, such as reconnaissance and strike UAVs, unmanned combat vehicles, etc., equipped with perception, decision-making and execution modules, supporting autonomous / cooperative operations, and conducting real-time detection of the combat target through real-time sensors: lidar, millimeter-wave radar, optoelectronic pods, Beidou / GNSS positioning modules. In one specific implementation, short-term combat status data includes obstacle information (presented as a numerical combination of "coordinates + type + size", such as "(X:1234.5, Y:6789.0, Z:100.2), concrete wall, length 15m × width 3m × height 5m"), target damage status (presented as a combination of percentage values and text descriptions, such as "75%, enemy tank tracks damaged, main gun unable to fire", with a value range of 0%-100%, where 0% indicates the target is intact and 100% indicates the target is completely destroyed, and the text description supplements key damaged parts and functional failures), and the location of friendly units (presented as a numerical form of real-time coordinates + status codes, with the coordinate format the same as the obstacle information, and the status codes identified by 2-digit numbers (such as 01 - normal combat, 02 - low battery, 03 - slightly damaged), such as "(X:1240.3, Y:6792.8, ... Z:101.5), Status 01") and / or enemy threat assessment (presented as a combination of threat level (level 1-5, level 1 being the lowest and level 5 the highest) + threat type + distance, for example, "threat level 4, enemy air defense missile system, distance 3.2km", where the distance unit is kilometers and the threat type is specified through preset keywords (such as air defense missile, tank, infantry fighting vehicle));
[0043] Short-term combat status data is input into a predetermined short-term decision-making layer, which then outputs real-time tactical actions for each unmanned combat unit through a predetermined deep reinforcement learning algorithm.
[0044] In one specific implementation, the deep reinforcement learning algorithm includes a predetermined short-term decision objective function, which is calculated using the following formula:
[0045] ,
[0046] ,
[0047] In the formula, Indicating short-term strategy The expected reward for performing a given immediate tactical action; Let it be the expected function; Indicating short-term strategy The strategy entropy is used to encourage units to explore diverse tactical actions in combat and to inject noise to simulate enemy interference. In one specific implementation, the initial short-term policy network parameters are sampled using a Gaussian distribution for immediate actions such as turning angles and attack commands, ensuring that actions are within tactical limits, such as the maximum turning angle. The weighting coefficients for the policy entropy; Indicates a short-term reward; A successful dodge scores a point. To avoid the weighting coefficient of successful scores; Indicates the target strike effect score. The weighting coefficient for the target strike effect score; This represents the penalty value for task failure. The weighting coefficient for the task failure penalty value; This represents the simulated value of enemy interference. The weighting coefficients are the simulated values of enemy interference.
[0048] In one specific implementation, short-term rewards include path offset, tactical navigation rewards such as the proportion of advance towards the target point, action penalties such as energy consumption per step, target recognition rewards such as 1 for accurately locking onto an enemy target and -1 for misjudging, and / or enemy interference penalties.
[0049] In one specific implementation, the path offset is calculated using the following formula:
[0050] ,
[0051] In the formula, Indicates the path offset; Indicates the closest distance to an enemy target / obstacle; Indicates the distance to the obstacle; Indicates the distance deviating from the tactical route; This indicates a deviation in heading.
[0052] In one specific implementation, the short-term decision objective function also updates its parameters using the following formula:
[0053] ,
[0054] ,
[0055] In the formula, It is a short-term decision objective function Relative to short-term policy network parameters The gradient; Indicating short-term strategy In short-term combat status Immediate tactical actions The logarithmic probability gradient; The advantage function measures the advantage of choosing action a in state s compared to the average value. Let be the state value function, representing the expected estimate of the overall value of the short-term operational state s based on the current strategy; For short-term combat status Immediate tactical actions The value function; Discount factor; Indicates the next short-term operational status Maximize the value of real-time tactical actions The value of this is determined by integrating the rewards of current tactical actions with expected future rewards, and considering game-theoretic interference, thus optimizing the value assessment of unit actions. In one specific implementation, every 100 steps, 32 combat samples are extracted from the experience pool, such as "enemy fire detected → evasion action → successful evasion," and the Q-function parameters are updated using gradient descent.
[0056] After executing immediate tactical actions to complete the current operational phase's mission instructions, the overall mission progress status is assessed. In one specific implementation, the overall mission progress status includes: mission completion progress (presented as a percentage value + task node completion status, e.g., "60%, completed 2 nodes: 'destroy enemy forward radar station' and 'establish temporary communication hub', with 1 node remaining: 'cover ground force advance'), calculated as "total weight of completed nodes / total weight of the mission × 100%", with each task node having a preset weight, such as a critical node weight of 0.5 and a general node weight of 0.2), and remaining resources (presented as a combination of "resource type + remaining quantity / total quantity + unit", e.g., "missiles: 12 / 20, fuel: 500L / 800L, ammunition: 300 rounds / "500 rounds", resource types include weapons, equipment, energy, consumables, etc., and the unit is set according to resource characteristics (e.g., rounds, units, fires), enemy deployment changes (presented in the form of "enemy unit type + original coordinates + new coordinates + quantity change" combined with text, such as "enemy infantry fighting vehicles: original (X:1250.1, Y:6800.5) → new (X:1260.3, Y:6810.2), quantity 8 vehicles → 6 vehicles (2 vehicles destroyed)") and / or game model (presented in the form of "enemy and friendly strategy payoff matrix + strategy probability distribution", where rows in the payoff matrix represent friendly strategies (e.g., 01 - frontal assault, 02 - flanking maneuver), columns represent enemy strategies (e.g., A - defense, B - retreat), and cell values are the payoff values under the corresponding strategy combination (e.g., +50 indicates friendly advantage, -30 indicates enemy advantage); the strategy probability distribution is presented as a percentage value, such as "friendly strategy 01 Probability 40%, 02 probability 60%, enemy strategy A probability 70%, B probability 30%) and combat time (presented as a time value in the form of "current stage time / planned time", in the unit of "hour:minute:second", for example "01:25:30 / 02:00:00").
[0057] The global task progress status is transmitted to the long-term planning layer, which then adjusts the long-term planning actions using a pre-defined reinforcement learning algorithm based on game theory to enhance policy gradients.
[0058] In one specific implementation, the reinforcement learning algorithm based on game theory to enhance policy gradients includes a predetermined long-term policy objective function, which is calculated using the following formula:
[0059] ,
[0060] ,
[0061] In the formula, For long-term strategy The expected return for performing a given long-term planned action; For long-term total rewards; For cluster task completion rate, This is a weighting coefficient for the cluster task completion rate; Due to resource scheduling deviation, This represents the weighting coefficient for resource scheduling deviation; To score the advantage in the game, the advantage quantification value (which can be positive or negative, with a positive value indicating that our side has the advantage and a negative value indicating that the enemy has the advantage) is calculated through game models (such as game payoff matrix and strategy confrontation deduction) during the confrontation between the unmanned swarm and the enemy. The weighting coefficients for the game advantage score. >0, used to strengthen the guidance of the game theory dimension on long-term strategies; This is the cluster loss value, a normalized value that comprehensively considers indicators such as the number of unmanned combat units destroyed and the degree of functional degradation (value range). (A larger value indicates a more severe cluster loss). The weighting coefficients for cluster loss values. The greater the loss, the more the total reward will be deducted.
[0062] In one specific implementation, the long-run policy objective function updates its parameters using the following formula:
[0063] ,
[0064] In the formula, It is the long-term decision objective function Relative to long-term policy network parameters The gradient; Indicating long-term strategy In the global task progress status Long-term planning actions The logarithmic probability gradient; In the postwar state Long-term planning actions The advantage function is determined by backpropagation of total reward and game-theoretic enhancements to optimize global resource allocation, such as ammunition resupply priorities, and phased mission planning, such as batch strike orders. In one specific implementation, parameters are updated every 30 minutes after each operational phase is completed, and long-term strategies are dynamically adjusted based on damage progress and threat assessments from short-term decision-making feedback.
[0065] In one specific implementation, the short-term decision-making layer uploads short-term operational status data to the long-term planning layer every 'a' seconds, such as 10 seconds; the long-term planning layer issues the operational mission instructions to the short-term decision-making layer every 'b' minutes, such as 5 minutes.
[0066] Example 2:
[0067] Figure 2 This is a schematic diagram of the structure of an unmanned swarm combat decision-making device according to an embodiment of the present invention, as shown below. Figure 2 As shown, the device includes a processor 201, a memory 202, a bus 203, and a computer program stored in the memory 202 and executable on the processor 201. The processor 201 includes one or more processing cores. The memory 202 is connected to the processor 201 via the bus 203. The memory 202 is used to store program instructions. When the processor executes the computer program, it implements the steps in the above-described method embodiment of Embodiment 1 of the present invention.
[0068] Furthermore, as an executable solution, the unmanned swarm combat decision-making device can be a computer unit, which can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer unit may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above-described computer unit structure is merely an example and does not constitute a limitation on the computer unit. It may include more or fewer components, or combine certain components, or use different components. For example, the computer unit may also include input / output devices, network access devices, buses, etc., and this embodiment of the invention does not limit this.
[0069] Furthermore, as an executable solution, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor is the control center of the computer unit, connecting various parts of the entire computer unit via various interfaces and lines.
[0070] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer unit by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0071] Example 3:
[0072] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the unmanned swarm combat decision-making method as described above.
[0073] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.
Claims
1. A decision-making method for unmanned swarm combat, characterized in that, Includes the following steps: The long-term planning layer of the hierarchical reinforcement learning architecture, enhanced by pre-defined game theory, issues operational mission instructions for the current operational phase; these operational mission instructions include mission priority, resource allocation instructions, and / or adversarial instructions. According to the operational mission instructions, each unmanned combat unit in the unmanned swarm is deployed to conduct real-time detection of the combat target to obtain short-term operational status data of the unmanned swarm; the short-term operational status data includes obstacle information, target damage status, location of friendly units and / or enemy threat assessment. The short-term combat status data is input into the short-term decision layer of a predetermined game-theoretic enhanced hierarchical reinforcement learning architecture. The short-term decision layer outputs the real-time tactical actions of each unmanned combat unit through a predetermined deep reinforcement learning algorithm. The deep reinforcement learning algorithm includes a predetermined short-term decision objective function, which is calculated using the following formula: , , In the formula, Indicating short-term strategy The expected reward for performing a given immediate tactical action; Let it be the expected function; Indicating short-term strategy strategy entropy, The weighting coefficients for the policy entropy; Indicates a short-term reward; A successful dodge scores a point. To avoid the weighting coefficient of successful scores; Indicates the target strike effect score. The weighting coefficient for the target strike effect score; This represents the penalty value for task failure. The weighting coefficient for the task failure penalty value; This represents the simulated value of enemy interference. The weighting coefficients for the simulated enemy interference values; After completing the current operational phase's operational mission instructions by executing the aforementioned real-time tactical actions, the global mission progress status is assessed. The global mission progress status includes mission completion progress, remaining resources, changes in enemy deployment and / or game model, and operational time. The global task progress status is transmitted to the long-term planning layer, which adjusts its long-term planning actions using a predetermined reinforcement learning algorithm based on game-theoretic enhanced policy gradients. The reinforcement learning algorithm based on game-theoretic enhanced policy gradients includes a predetermined long-term policy objective function, which is calculated using the following formula: , , In the formula, For long-term strategy The expected return of performing a given long-term planned action. Let it be the expected function; For long-term total rewards; For cluster task completion rate, This is a weighting coefficient for the cluster task completion rate; Due to resource scheduling deviation, This represents the weighting coefficient for resource scheduling deviation; To gain an advantage in the game and score points. The weighting coefficients for the game advantage score; This represents the cluster loss value. The weighting coefficients for cluster loss values.
2. The unmanned swarm combat decision-making method according to claim 1, characterized in that, The short-term decision objective function also updates its parameters using the following formula: , , In the formula, It is the short-term decision objective function Relative to short-term policy network parameters The gradient; Indicating short-term strategy In short-term combat status Immediate tactical actions The logarithmic probability gradient; The advantage function measures the advantage of choosing action a in state s compared to the average value. Let be the state value function, representing the expected estimate of the overall value of the short-term operational state s based on the current strategy; For short-term combat status Immediate tactical actions The value function; Discount factor; Indicates the next short-term operational status Maximize the value of real-time tactical actions The value of.
3. The unmanned swarm combat decision-making method according to claim 1, characterized in that, The long-term strategy objective function also updates its parameters using the following formula: , In the formula, It is the long-term decision objective function Relative to long-term policy network parameters gradient, Let it be the expected function; Indicating long-term strategy In the global task progress status Long-term planning actions The logarithmic probability gradient; In the postwar state Long-term planning actions The advantage function.
4. The unmanned swarm combat decision-making method according to claim 1, characterized in that, The short-term rewards include path offset, tactical navigation rewards, action penalties, target recognition rewards, and / or enemy interference penalties.
5. The unmanned swarm combat decision-making method according to claim 4, characterized in that, The path offset is calculated using the following formula: , In the formula, Indicates the path offset; Indicates the closest distance to an enemy target / obstacle; Indicates the distance to the obstacle; Indicates the distance deviating from the tactical route; This indicates a deviation in heading.
6. The unmanned swarm combat decision-making method according to claim 1, characterized in that, The short-term decision-making layer uploads the short-term operational status data to the long-term planning layer every 'a' seconds; the long-term planning layer issues the operational mission instructions to the short-term decision-making layer every 'b' minutes.
7. A decision-making device for unmanned swarm combat, characterized in that, It includes a memory and a processor, wherein the memory stores at least one program, which is executed by the processor to implement the unmanned swarm combat decision-making method as described in any one of claims 1 to 6.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the unmanned swarm combat decision-making method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Construction method of cooperative combat decision-making agent for ground unmanned equipment
CN119514637A