Aerial formation strike distribution method and related products

By acquiring air formation situational information and using decision-making models of hybrid networks and multi-agent networks to generate precise strike allocation control commands, the problem of insufficient flexibility of traditional methods in complex battlefield environments is solved, and efficient autonomous decision-making and precision strikes of air formations in highly dynamic environments are realized.

CN120973024APending Publication Date: 2025-11-18BAIYANG TIMES (BEIJING) TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511423448.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional methods for air formation target decision-making and allocation are difficult to meet the requirements of real-time performance and flexibility in complex and ever-changing battlefield environments. In particular, in modern warfare with its high dynamics and increased uncertainty, existing methods face challenges such as poor flexibility and high computational difficulty.

Method used

By acquiring situational information of the air formation, a trained decision model is used to generate strike allocation control commands based on threat data. The decision model includes a hybrid network and a multi-agent network. A proximal optimization algorithm is used to train the agents to learn collaboratively, ultimately generating accurate strike allocation control commands.

Benefits of technology

It significantly enhances the autonomous decision-making capability and real-time mission response of air formations in dynamic environments, improves the accuracy and flexibility of target allocation, and optimizes the computational efficiency of the decision-making model while reducing the overall computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973024A_ABST
    Figure CN120973024A_ABST
Patent Text Reader

Abstract

The invention discloses an air formation strike distribution method and related products. The method comprises the steps of obtaining situation information of a target area where an air formation is located; the air formation comprises a first formation or a second formation; determining threat data based on the situation information; the threat data comprises threat values of a plurality of sub-regions and threat values of a plurality of objects; using the trained decision model to obtain a strike distribution control instruction based on the threat data and the situation information; the trained decision model comprises a first formation decision model or a second formation decision model; each of the first formation decision model and the second formation decision model comprises a hybrid network and a multi-agent network; the multi-agent network is obtained by training based on a near-end optimization algorithm; the multi-agent network comprises a plurality of agents; each agent comprises a strategy network, a learning network and an evaluation network; the air formation is controlled to strike based on the strike distribution control instruction, the calculation efficiency of the decision model is optimized, and the overall operation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cooperative formation control, in particular to a strike allocation method for an aerial formation and related products. BACKGROUND

[0002] With the development of artificial intelligence technology, its application in military command and decision-making is increasing, especially in the use of aerial formations in urban combat environments. In such scenarios, aerial formations need to analyze threats, judge enemy intentions, and quickly decide on the best strike plan, such as direct annihilation, pre-emption, or interception, to achieve rapid victory. However, in the face of complex and changing task targets and battlefield environments, traditional target decision allocation methods cannot meet the real-time and flexibility requirements.

[0003] The current commonly used target decision allocation methods include classical methods based on problem solving, knowledge-driven methods, and data-driven methods. Classical methods based on problem solving solve the optimal strategy by establishing mathematical models, which are suitable for small-scale problems, but as the problem size increases, the demand for computing resources increases dramatically, and the difficulty of operation increases. Knowledge-driven methods rely on expert experience to build state-action models, which have good interpretability, but when faced with complex problems, the workload of knowledge extraction and modeling is huge, and the flexibility is insufficient. Data-driven methods use supervised learning, unsupervised learning, and reinforcement learning to improve strategies, among which reinforcement learning is particularly suitable for handling decision-making problems in dynamic environments, but its flexibility is limited by the changes in opponent strategies and the authenticity of the simulation environment. In summary, although the above methods are effective under certain conditions, they generally have poor flexibility and high difficulty in operation in the face of modern warfare with high dynamics and increasing uncertainty in the battlefield environment. SUMMARY

[0004] Based on the above problems, the present application provides a strike allocation method for an aerial formation and related products, aiming to improve the flexibility of target strike allocation and reduce the difficulty of operation.

[0005] The embodiments of the present application disclose the following technical solutions:

[0006] The first aspect of the present application provides a strike allocation method for an aerial formation, which comprises:

[0007] obtaining situation information of a target area where the aerial formation is located; the aerial formation comprises a first formation or a second formation; the first formation and the second formation each comprise a plurality of aircrafts;

[0008] determining threat data based on the situation information; the threat data comprises threat values of a plurality of sub-regions and threat values of a plurality of objects; the target area comprises a plurality of sub-regions; the sub-region comprises a plurality of objects;

[0009] obtain, by using a trained decision model, a strike allocation control instruction based on the threat data and the situation information; the trained decision model comprises a first formation decision model or a second formation decision model; the first formation decision model and the second formation decision model both comprise a hybrid network and a multi-agent network; the multi-agent network is trained based on a proximal optimization algorithm; the multi-agent network comprises a plurality of agents; each agent comprises a policy network, a learning network and an evaluation network;

[0010] control the aerial formation to carry out the strike based on the strike allocation control instruction.

[0011] The second aspect of the present application provides an aerial formation strike allocation device, which comprises:

[0012] an acquisition module, configured to acquire situation information of a target region where an aerial formation is located; the aerial formation comprises a first formation or a second formation; the first formation and the second formation both comprise a plurality of aircrafts;

[0013] a threat data determination module, configured to determine threat data based on the situation information; the threat data comprises threat values of a plurality of sub-regions and threat values of a plurality of objects; the target region comprises a plurality of sub-regions; the sub-regions comprise a plurality of objects;

[0014] a strike allocation control instruction determination module, configured to obtain, by using a trained decision model, a strike allocation control instruction based on the threat data and the situation information; the trained decision model comprises a first formation decision model or a second formation decision model; the first formation decision model and the second formation decision model both comprise a hybrid network and a multi-agent network; the multi-agent network is trained based on a proximal optimization algorithm; the multi-agent network comprises a plurality of agents; each agent comprises a policy network, a learning network and an evaluation network;

[0015] a control module, configured to control the aerial formation to carry out the strike based on the strike allocation control instruction.

[0016] The third aspect of the present application provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor; the processor executes the computer program to implement the aerial formation strike allocation method provided in the first aspect.

[0017] Compared with the prior art, the present application has the following beneficial effects:

[0018] The application obtains situation information of a target area where an air formation is located, determines threat data based on the situation information, obtains a strike allocation control instruction based on the threat data and the situation information by using a trained decision model, and controls the air formation to strike based on the strike allocation control instruction. By obtaining the situation information of the air formation including multiple aircrafts in the target area in real time, accurately identifying and quantifying the threat data of multiple sub-regions and multiple objects based on the information, and then intelligently generating a strike allocation control instruction by using a decision model of a fusion hybrid network and a multi-agent network, the multi-agent network is trained by a proximal policy optimization algorithm, which effectively improves the collaborative decision-making efficiency in a complex battlefield environment. Finally, the air formation is controlled to implement precise strikes based on the generated strike allocation control instruction, which not only significantly enhances the autonomous decision-making ability and task response real-time of the formation in a dynamic environment, but also improves the accuracy and flexibility of the strike target allocation, optimizes the calculation efficiency of the decision model, and reduces the overall computational complexity. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor under the premise of not paying creative labor.

[0020] Figure 1 A flowchart of a strike allocation method of an air formation provided by an embodiment of the present application;

[0021] Figure 2 A training flowchart of a decision model provided by an embodiment of the present application;

[0022] Figure 3 A training flowchart using a proximal optimization algorithm provided by an embodiment of the present application;

[0023] Figure 4 A structure diagram of a strike allocation device of an air formation provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] As described earlier, with the development of artificial intelligence (AI) technology, military command and decision-making are gradually incorporating AI to address dynamic battlefield environments. In urban warfare, air formations need to quickly analyze threats, assess enemy intentions, and conduct precision strikes after reconnaissance, such as decapitation strikes or intercepting escaping targets. This places extremely high demands on the real-time performance and accuracy of command and decision-making, creating a complex multi-target decision allocation problem. Traditional methods, including classical mathematical optimization, knowledge-driven reasoning, and data-driven learning, while effective in small-scale scenarios, generally suffer from high computational complexity, poor flexibility, difficulty in knowledge acquisition, or reliance on large amounts of labeled data, making them unsuitable for large-scale, highly dynamic combat needs. Especially in areas such as formation coordination and dynamic adjustment of target priorities, existing methods lack sufficient flexibility.

[0025] In view of the above problems, this application provides a method and related products for generating strike allocation for aerial formations. The method includes: acquiring situational information of the target area where the aerial formation is located; determining threat data based on the situational information; obtaining strike allocation control commands based on the threat data and the situational information using a trained decision model; and controlling the aerial formation to carry out strikes based on the strike allocation control commands. This application acquires situational information of multi-aircraft aerial formations in real time, accurately identifies and quantifies the threat level of each sub-region and target object, constructs a decision model integrating hybrid networks and multi-agent networks, and intelligently generates strike allocation commands. The multi-agent network is trained using the Proximal Policy Optimization (PPO) algorithm, enabling each agent to collaboratively learn the optimal strategy under complex battlefield conditions, significantly improving decision-making efficiency and robustness. Precision strikes are carried out based on commands, effectively enhancing the autonomous coordination capability of the formation in highly dynamic environments and improving the accuracy and flexibility of target allocation. Simultaneously, model optimization reduces computational complexity, improving real-time performance and scalability.

[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0027] Figure 1 A flowchart illustrating a strike allocation method for an air formation provided in this application embodiment is shown below. Figure 1 As shown, a strike allocation method for an air formation includes:

[0028] S101: Obtain situational information about the target area where the air formation is located.

[0029] This application does not limit the definition of an aerial formation. Exemplarily, the aerial formation includes a first formation or a second formation, wherein each of the first and second formations contains multiple aircraft. In other words, the aerial formation consists of a group of controllable aircraft, enabling precise command and control of each aircraft, thereby improving the organization and responsiveness of the overall combat operations.

[0030] The first and second formations are in a confrontational relationship. Therefore, the situational information in this application needs to obtain the states of both the first and second formations. For example, if an air formation includes the first formation, then the enemy is the second formation. To effectively command and control the aircraft in the first formation, the states of both the first and second formations must be obtained simultaneously. Only then can the strike allocation control command for the first formation be determined based on the states of the first and second formations. However, this application does not limit the specific data of the situational information. For example, the situational information includes: the simulated entity attribute characteristics, statistical characteristics, and spatial situational characteristics of the first and second formations. The simulated entity attribute characteristics of the first and second formations consist of the entity's inherent attributes, values ​​assigned during scene initialization, and dynamic state values, including: formation, longitude, latitude, altitude, pitch angle, roll angle, yaw angle, speed, damage level, load, protection capability, radar cross section (RCS) value, strike capability, etc. Statistical characteristics are reports of overall changes caused by a specific event in the first formation during a combat scenario. Examples include: reconnaissance reports generated by reconnaissance actions, strike reports generated by attack actions, and fuel consumption resulting from maneuvers. These characteristics include: simulation propulsion time, types of second formation units detected by reconnaissance, number of second formation units detected by reconnaissance, ammunition consumption of the first formation, remaining fuel of the first formation, remaining ground-attack missiles of the first formation, number of destroyed second formation units, number of destroyed first formation units, and number of destroyed second formation units. Spatial situational awareness can describe changes in the deployment and response of both the first and second formations based on geographic environmental data. Real-time acquisition of overall air formation situational awareness information, including both the first and second formations, supports comprehensive perception of the coordinated status of multiple aircraft.

[0031] S102: Determine threat data based on the situation information.

[0032] The threat data includes threat values ​​for multiple sub-regions and threat values ​​for multiple objects; the target region includes multiple sub-regions; and each sub-region includes multiple objects.

[0033] This application does not limit the method for determining threat data. For example, it can use a large language model to analyze situational information and determine threat data, such as a mobile air defense system detected in sub-area A2, with a threat value of 0.9; a strong electromagnetic interference source present in sub-area A3, with a threat value of 0.7; a suspected command center in sub-area A4, with a threat value of 0.85; and threat values ​​for other areas below 0.3. Each threat object is labeled with its type, location, movement trend, and priority. Based on this situational information, the threat level of each sub-area and each object can be accurately identified and quantified, threat data can be determined, and the accuracy of battlefield awareness can be improved.

[0034] S103: Using the trained decision model, based on the threat data and the situation information, obtain strike allocation control instructions.

[0035] This application does not limit the strike allocation control command, such as strike allocation control command including target selection, maneuvering or attack (e.g., maneuvering toward the target to attack, then irradiating or striking).

[0036] The trained decision-making model proposed in this application includes either a first formation decision-making model or a second formation decision-making model. The first and second formation decision-making models are primarily determined by the air formation; that is, if the air formation includes the first formation, the trained decision-making model includes the first formation decision-making model; if the air formation includes the second formation, the trained decision-making model includes the second formation decision-making model. Both the first and second formation decision-making models include a hybrid network and a multi-agent network. The multi-agent network is trained based on a proximal optimization algorithm. The multi-agent network includes multiple agents; each agent includes a policy network, a learning network, and an evaluation network to support efficient local collaborative decision-making and global coordination. In short, by equipping each formation with a dedicated decision-making model, the accuracy and execution effect of strike allocation commands can be significantly improved, thereby enhancing overall combat effectiveness. This not only ensures the effectiveness and relevance of commands but also improves the ability to cope with changes in complex battlefield environments. This application does not limit the network structure of the hybrid network; for example, the hybrid network includes two fully connected modules and one supernetwork module. Multi-agent networks are used to output the actions (i.e., the strike commands of each aircraft) and values ​​of multiple agents, while hybrid networks calculate the total value of the overall situation of the multi-agent network based on the actions and values ​​of each agent.

[0037] The above describes the trained decision model from the perspective of network structure. To further illustrate the trained decision model, it can also be described from the perspective of hierarchy. For example, the trained decision model mainly includes an input part, an intermediate part, and an output part. The input part mainly uses a multilayer perceptron to transform various entities into a vector, which can be an embedding process. The intermediate process is relatively simple, that is, the current embedding process is input through a deep gated recurrent unit network. The output part includes two parts: one part is to output the value of the action through the multilayer perceptron, and the other part includes to output the vertical movement type through the residual perceptron, to output the horizontal movement type through the residual perceptron, and to first output whether to attack through the residual perceptron, then superimpose the enemy situation, and output the selected enemy target through the pointer network to obtain the attack allocation control command.

[0038] In practical applications, the trained decision-making model is trained by continuously trying and failing in a simulation environment and adjusting the rewards based on the adversarial effect. The trained decision-making model directs the behavior of the air formation in real time, and can cope with a variety of threats and enemy behaviors and strategies. It prioritizes the task of decapitating the enemy's leader, while also taking into account the tasks of annihilating ground forces and occupying core positions, achieving a good balance between mission completion and combat losses.

[0039] S104: Control the air formation to carry out an attack based on the strike allocation control command.

[0040] Based on the generated strike allocation control commands, each formation is precisely scheduled to execute strike missions, significantly improving the autonomous decision-making level, mission response speed, and resource allocation rationality of multiple formations. This not only enhances operational flexibility and robustness in highly confrontational and dynamic environments but also reduces overall computational complexity, avoids the "curse of dimensionality" of centralized decision-making, and improves scalability and fault tolerance.

[0041] The above describes the main technical solution of this application. Further implementations of the main technical solution are now introduced. Details are as follows:

[0042] Regarding S102's determination of threat data based on the situational information, this application provides an optional embodiment:

[0043] Based on the situational information, the static and dynamic attributes of the target are distinguished, and a threat indicator system is constructed based on the static and dynamic attributes of the target.

[0044] The static attributes of the target in this application may include the target's basic attributes, target type (such as radar vehicle, air defense vehicle, command post, tank, etc.), and defense level; the dynamic attributes of the target may include the target's location, target's direction, target's movement speed, rate of change of heading, electromagnetic radiation intensity, and behavior pattern (such as patrol, silence, assembly); the constructed threat index system may include the target's static attributes (such as target type weight (such as radar vehicle = 4, tank = 1) and protection level (high / medium / low)) and the target's dynamic attributes (such as movement trend (approaching formation = +0.3), communication activity (strong = 0.8), and behavior anomaly (deviation from normal path = 0.6)).

[0045] The threat indicator system is preprocessed to obtain the processed threat indicator system.

[0046] Preprocessing the threat indicator system can facilitate subsequent operations, but this application does not limit the preprocessing method. The preprocessing may include dimensionless processing or score-based processing.

[0047] For example, the range method is used to perform dimensionless processing, mapping all indicators uniformly to the [0, 1] interval; and the categorical variables (such as type) are scored and encoded (such as radar vehicle → 4 points, air defense vehicle → 3 points) to form a standardized processed threat indicator system, which is convenient for weighted fusion.

[0048] Threat values ​​for multiple sub-regions are obtained based on the processed threat index system.

[0049] This application does not limit the method for determining the threat value of a sub-region. For example, the threat value of multiple sub-regions can be obtained by using a deep belief network based on the processed threat index system, or the threat value of each sub-region can be calculated by using a weighted summation model based on the processed threat index system. Example results: Sub-region A2 has a radar vehicle and high electromagnetic activity → threat value = 0.91; Sub-region A3 has a mobile air defense unit → threat value = 0.76; Sub-region A4 is suspected to be a command node → threat value = 0.83.

[0050] The threat value of each object in each sub-region is obtained by analyzing the situational information using Bayesian inference methods.

[0051] Threat data is determined based on the threat values ​​of the multiple sub-regions and the threat value of each object in each sub-region.

[0052] To present the threat data generated in this application more intuitively, threat information can also be visualized through text briefings or graphics. For example, an electronic map overlaid with a heatmap can be used to display the threat value distribution of each sub-region, with different colors indicating threat levels (e.g., red for high threat, yellow for medium threat, and green for low threat). Icons can be used to mark the location, type, and real-time threat score of high-threat targets. Structured text briefings can also be generated, listing key targets and their attributes sorted by threat level. This combination of text and graphics significantly improves the readability, spatial awareness, and decision support efficiency of threat data, enabling commanders to quickly grasp the battlefield situation and make timely and accurate operational responses.

[0053] This application also provides a specific application embodiment based on the premise that the air formation includes a first formation and the enemy includes a second formation:

[0054] Through T f l =S r / V t Calculate the approach time T of the second formation target fl Among them: S r Indicates the distance from the second formation to the target; V t This represents the speed value of the second formation. It also calculates the target behavior intent threat value K of the second formation. a (For example, the threat value for air attack is 4, the threat value for reconnaissance is 3, the threat value for cover is 2, the threat value for feint attack is 1, and the threat value for retreat is 0), the target threat value K of the second formation. b (For example, the threat value of the radar vehicle is 4, the threat value of the air defense vehicle is 3, and the threat value of the individual RPG is 2). Based on the predicted targets, determine the importance value K of the first formation's target. r (For example, the importance coefficient of an individual helicopter is 6, the importance coefficient of a medium-sized UAV is 3, and the importance coefficient of a small UAV is 1).

[0055] Therefore, the threat data f(T) of a single target unit in the second formation can be calculated using the threat data calculation formula. f l ,K a ,K b ,K r The threat data calculation formula is as follows:

[0056] ;

[0057] Where: p represents the probability that the target is real; w1, w2, w3 and w4 are all weight parameters that determine the degree of influence of the first formation target, predicted behavior, predicted attack target, arrival time and probability on the threat level; k represents a constant used to adjust the influence of the real probability of the first formation target on the threat level.

[0058] After the threat data is calculated, it can be classified into levels. Taking into account factors such as the practicality of defensive operations, the feasibility of model processing, and the commander's thinking habits, the target threat level is divided into three levels: low threat (e.g., threat data greater than or equal to 0 and less than 0.5), medium threat (e.g., threat data greater than or equal to 0.5 and less than 0.8), and high threat (e.g., threat data greater than or equal to 0.8 and less than 1).

[0059] The above describes the method for determining threat data in detail. Now, the method for determining strike allocation control commands will be described in more detail. Specifically, regarding S103, using the trained decision model based on the threat data and the situation information, this application provides an optional embodiment:

[0060] The threat data and the situation information are normalized to obtain normalized threat data and normalized situation information.

[0061] The acquired threat data and situational information are preprocessed to obtain threat data and situational information that meet the input requirements of the trained decision model. The simulated entity attribute features, statistical features, and threat data in the situational information can be normalized by dividing the current value by the maximum value to obtain values ​​in the [0, 1] interval. The spatial situational features in the situational information can be used to discretize the map features, using a 256×256×3 matrix. The map coordinate system is divided according to 1×1 kilometers, and the attribute values ​​such as unit combat effectiveness and mobility within the grid are normalized.

[0062] The normalized threat data and the normalized situation information are input into the trained decision model to obtain the strike allocation control command.

[0063] By normalizing threat data and situational information, data from different sources and at different scales can be standardized and transformed, enabling these data to be more effectively applied to subsequent analysis and decision-making processes. The normalized threat data and situational information meet the input requirements of the trained decision model, ensuring the effective fusion of multi-source heterogeneous data and enhancing the model's ability to handle complex battlefield environments. By scaling the acquired threat data and situational information to the [0, 1] interval using the method of dividing the current value by the maximum value, a unified scale representation of the data is achieved, facilitating direct comparison across different dimensions and improving the accuracy and efficiency of data analysis. The spatial features in the situational information are used to discretize the map features, and a 256×... 256×The 3D matrix representation not only preserves geospatial information but also considers the spatial distribution characteristics of key attributes such as unit combat effectiveness and mobility. This grid-based representation facilitates refined management of battlefield resources and optimizes strike allocation strategies. Inputting normalized threat data and situational information into a trained decision-making model yields more scientific and rational strike allocation control commands. This not only improves the accuracy and response speed of combat command but also enhances the ability to counter complex and dynamic environmental changes.

[0064] In summary, by normalizing threat data and situational information, the applicability of the data and the effectiveness of the decision support system are significantly improved, providing strong technical support for precision strikes.

[0065] While ensuring the validity of the data input to the trained decision model, it is also necessary to guarantee the accuracy of the trained decision model. Since the trained decision model in this application includes either a first formation decision model or a second formation decision model, and the first formation decision model is for the first formation while the second formation decision model is for the second formation, and there is an adversarial relationship between the first and second formations, in order to ensure the effectiveness of the trained decision model, this application first trains the first formation decision model and the second formation decision model separately during the training process. Then, a game-theoretic approach is used to perform a second round of network parameter adjustments on the first and second formation decision models after the first adjustment, ensuring the effectiveness of the trained decision model constructed from the final first or second formation decision model. For the specific training process, this application provides an optional implementation:

[0066] Figure 2 A flowchart of the training process of a decision model provided in this application embodiment is shown below. Figure 2 As shown, it includes the following steps:

[0067] S201: The network parameters of the first target agent are trained and updated using a proximal optimization algorithm to obtain the updated network parameters of the first target agent.

[0068] Wherein, the first target agent is any agent in the multi-agent network of the first formation decision model.

[0069] When training and updating the network parameters of a decision model, a proximal policy optimization algorithm can be used to iteratively optimize the parameters of each network in the model. However, considering that the multi-agent network of the decision model in this application consists of multiple agents that share the same set of network parameters (i.e., a parameter sharing mechanism), this application proposes an efficient parameter update strategy: selecting only one agent in the multi-agent network and using its experience data from interacting with the environment to train and update the network parameters; subsequently, the updated parameters are synchronously applied to all other agents. This fully utilizes the policy consistency advantage under the parameter sharing mechanism, significantly reducing redundant computation while ensuring the collaborative capabilities of each agent, and avoiding redundant gradient update processes performed individually for each agent. This not only effectively reduces the computational overhead and communication costs during training but also improves the model's convergence speed and training stability.

[0070] S202: Based on the updated network parameters of the first target agent, update the network parameters of the agents other than the first target agent in the first formation decision model to obtain the updated first formation decision model.

[0071] Since the network parameters of this application are shared, when the number of agents changes, it is only necessary to increase the number of input dimensions.

[0072] Since the trained decision model is mainly used in adversarial scenarios, after updating the network parameters of the first formation decision model using the proximal optimization algorithm, it is necessary to further adjust the network parameters of the updated first formation decision model in an adversarial manner, namely S203-S209.

[0073] S203: Obtain the decision model and historical situation information based on the first training data.

[0074] The only difference between historical situation information and the situation information in S101 is the acquisition time. The situation information in S101 is acquired in real time, while historical situation information is data from any point in history. The decision model based on the first training data can be understood as a neural network model trained using traditional training methods, that is, trained using training data.

[0075] S204: Input the historical situation information into the decision model based on the first training data and the updated first formation decision model respectively to obtain the first historical strike allocation control command and the second historical strike allocation control command.

[0076] The first historical strike allocation control command is obtained using the decision model based on the first training data; the second historical strike allocation control command is obtained using the updated first formation decision model.

[0077] S205: Control the first simulated air formation and the second simulated air formation to engage in combat and obtain the first combat result.

[0078] The first simulated air formation is controlled using the first historical strike allocation control command; the second simulated air formation is controlled using the second historical strike allocation control command.

[0079] S206: Does the preset number of iterations meet the requirement?

[0080] This application does not limit the preset number of iterations, but can set it according to the actual situation, such as the preset number of iterations including 5,000.

[0081] S207: If the preset number of iterations is met, the first formation decision model after the first adjustment is obtained, and the network parameters of the first formation decision model and the second formation decision model after the first adjustment are adjusted.

[0082] S208: If the preset number of iterations is not met, and the first adversarial result indicates that the first simulated air formation has won, then return to S201.

[0083] S209: If the preset number of iterations is not met, and the first adversarial result indicates that the second simulated air formation has won, then the network parameters of the decision model based on the first training data are adjusted, and the process returns to S204.

[0084] Adjusting the network parameters of the updated first formation decision model using a decision model derived from historical training data can be simply understood as using a well-trained and stable "teacher model" (i.e., the global decision model trained in the traditional way) to guide and optimize the parameter update process of the "student model" (i.e., the updated first formation decision model). This mechanism is essentially a training strategy combining model transfer learning and knowledge distillation. Specifically, in a multi-agent collaborative decision-making system, although the agents in the first formation learn autonomously through online interaction or reinforcement learning methods such as PPO, their initial strategies may be unstable and their exploration efficiency low.

[0085] Introducing a pre-trained decision-making model as a reference allows the output of a mature model to serve as prior knowledge, guiding the agent to converge faster. Adding a regularization term during parameter updates prevents deviations from known effective strategies, improving training stability. General decision-making capabilities can be transferred to specific formation models, making them particularly suitable for scenarios involving new formation additions or sudden environmental changes. Therefore, the method described in this embodiment not only accelerates the learning process of the first formation decision-making model but also inherits the rationality and robustness of the globally optimal strategy while ensuring personalized decision-making capabilities, effectively balancing exploration and utilization, and enhancing the overall system's intelligence and practical adaptability.

[0086] Regarding the network parameter adjustment of the first formation decision model and the second formation decision model after the first adjustment in S207, this application provides an optional embodiment:

[0087] The network parameters of the second target agent are trained and updated using a proximal optimization algorithm to obtain the updated network parameters of the second target agent. The second target agent is any agent in the multi-agent network of the second formation decision model.

[0088] Based on the updated network parameters of the second target agent, the network parameters of agents other than the second target agent in the second formation decision model are updated to obtain the updated second formation decision model.

[0089] Obtain a decision model based on the second training data.

[0090] The training methods for the decision model based on the second training data and the decision model based on the first training data are the same. The difference is that the decision model based on the first training data is trained on the historical data of the second formation, while the decision model based on the second training data is trained on the historical data of the first formation.

[0091] The historical situation information is input into the decision model based on the second training data and the updated second formation decision model, respectively, to obtain the third historical strike allocation control command and the fourth historical strike allocation control command; the third historical strike allocation control command is obtained using the decision model based on the second training data; the fourth historical strike allocation control command is obtained using the updated second formation decision model.

[0092] The third and fourth simulated air formations are controlled to engage in combat, resulting in a second combat outcome. The third simulated air formation is controlled using the third historical strike allocation control command, and the fourth simulated air formation is controlled using the fourth historical strike allocation control command.

[0093] If the second adversarial result indicates that the third simulated air formation has won, then the process returns to the step of training and updating the network parameters of the second target agent using a near-end optimization algorithm to obtain the updated network parameters of the second target agent, until the preset number of iterations is met, to obtain the second formation decision model after the first adjustment, and then the network parameters of the first adjusted first formation decision model and the first adjusted second formation decision model are adjusted using a game theory method.

[0094] If the second confrontation result indicates that the fourth simulated air formation has won, then the decision model based on the second training data is adjusted, and the process of inputting the historical situation information into the decision model based on the second training data and the updated second formation decision model is returned to obtain the third historical strike allocation control command and the fourth historical strike allocation control command, until the preset number of iterations is met, and the first adjusted second formation decision model is obtained. Then, the network parameters of the first adjusted first formation decision model and the first adjusted second formation decision model are adjusted using a game theory method.

[0095] This embodiment mainly utilizes Figure 2 The training method shown yields the second formation decision model after the first adjustment, thereby achieving the goal of adjusting the network parameters of the first and second formation decision models after the first adjustment using a game theory approach. In short, it achieves the goal of agent models adversarially against each other, and this application also provides corresponding optional embodiments:

[0096] The historical situation information is input into the first formation decision model and the second formation decision model after the first adjustment, respectively, to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command; the fifth historical strike allocation control command is obtained using the first formation decision model after the first adjustment; the sixth historical strike allocation control command is obtained using the second formation decision model after the first adjustment.

[0097] The fifth and sixth simulated air formations are controlled to engage in combat, resulting in a third combat outcome. The fifth simulated air formation is controlled using the fifth historical strike allocation control command, and the sixth simulated air formation is controlled using the sixth historical strike allocation control command.

[0098] If the third confrontation result indicates that the fifth simulated air formation has won, then the network parameters of the second formation decision model after the first adjustment are trained and updated using a near-end optimization algorithm. The process of inputting the historical situation information into the first formation decision model and the second formation decision model after the first adjustment is returned to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command is repeated until the preset number of iterations is met, and the final second formation decision model is obtained.

[0099] If the third confrontation result indicates that the sixth simulated air formation has won, then the network parameters of the first formation decision model after the first adjustment are trained and updated using a near-end optimization algorithm. The process of inputting the historical situation information into the first formation decision model and the second formation decision model after the first adjustment, respectively, to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command is repeated until the preset number of iterations is met, and the final first formation decision model is obtained.

[0100] The trained decision model is constructed based on the final first formation decision model or the final second formation decision model.

[0101] This application significantly improves the robustness and combat adaptability of the formation decision-making model by introducing an adversarial iterative training mechanism based on historical situational information. Specifically, historical situational information is input into the first and second formation decision-making models after their first adjustment, generating the fifth and sixth historical strike allocation control commands. These commands then drive the corresponding fifth and sixth simulated air formations to engage in virtual combat, yielding the third combat result. Based on the outcome of the combat, the lagging model is dynamically selected for reinforcement learning updates, and this process is iterated until a preset number of training rounds is reached.

[0102] Through a closed-loop mechanism of "instruction generation - simulated confrontation - feedback optimization," the decision-making model continuously learns and evolves in ongoing games with similar strategies, approaching a better strategy. Regardless of which formation wins, the model capabilities of the losing side are specifically strengthened, forming a healthy competition mechanism and preventing the model from getting stuck in local optima or outdated strategies. Diverse historical situational data is used as input to enhance the model's adaptability to complex and dynamic battlefield environments and its decision-making robustness. Combined with the stable gradient update characteristics of the PPO algorithm, the convergence speed is accelerated while ensuring training stability. The decision-making models of the first and second formations are trained and optimized separately, facilitating subsequent expansion to more types of formations or heterogeneous platforms (such as manned / unmanned collaboration).

[0103] Finally, based on the final decision-making model of the first formation or the final decision-making model of the second formation after iterative optimization, a trained decision-making model is constructed to ensure that it possesses highly intelligent and adaptable collaborative combat decision-making capabilities. This effectively solves the problems of poor model flexibility and unstable performance in actual combat under the traditional static training mode, and provides reliable technical support for autonomous collaborative command of multiple formations in complex battlefield environments.

[0104] Regarding the use of a proximal optimization algorithm in S201 to train and update the network parameters of the first target agent, resulting in the updated network parameters of the first target agent, this application provides an optional embodiment:

[0105] Figure 3 The training flowchart using the proximal optimization algorithm provided in the embodiments of this application is as follows: Figure 3 As shown, it includes the following steps:

[0106] S301: Expand the experience pool based on the policy network in the first target agent to obtain the expanded experience pool.

[0107] The expanded experience pool includes data combinations from multiple moments; these data combinations include environmental information, actions, and rewards.

[0108] When conducting target strikes, air formations must manage coordination between different entities, confrontation with the enemy, elimination of enemy leaders, and engagement with ground forces. Therefore, target strikes by air formations are a hybrid game problem. First, there is the issue of partial observability, meaning that the observation range of a single aircraft is limited, making it impossible to make optimal decisions based solely on global observation. Second, there is environmental instability, meaning that while individual aircraft are making decisions, other aircraft are also making autonomous decisions. Changes in the situational state are related to the coordinated actions of the air formation; therefore, when the strategies of other aircraft change, existing local strategies become ineffective, preventing the formation from achieving a Nash equilibrium. Third, there is the issue of individual goal consistency, requiring the goals of each agent to be adjusted to achieve the optimal global reward, rather than the optimal local reward. Therefore, it is necessary to set reasonable individual and overall goals and phased reward functions.

[0109] The reward function considers both individual and overall rewards. Individual rewards represent the combat objectives of a single agent-controlled fighter jet cluster, while overall rewards represent the entire air formation's successful completion of its mission. While considering both individual and overall rewards, it's crucial to maintain a balance between them, encouraging individual growth while guiding training towards the overall objective. Secondly, rewards that promote cooperation among multiple agents need to be introduced between individual and overall rewards. For example, rewards from other entities can be incorporated into the individual reward of a single agent, encouraging stronger collaboration and achieving a win-win outcome. Finally, while various detailed reward designs aim to guide agents to learn effective tactics more quickly, overly detailed designs may deviate from the original training objectives. Therefore, after a certain training stage, the simplest and most direct win / loss indicators should be used as reward signals, allowing agents more freedom to explore optimal strategies for victory.

[0110] The policy gradient algorithm does not backpropagate through error. Instead, it selects a behavior based on observation information and backpropagates it directly. It uses rewards to directly increase or decrease the probability of the selected behavior of each fighter group. Good behavior will increase the probability of being selected next time, while bad behavior will decrease the probability of being selected next time.

[0111] The reward function in this application includes an attack reward (e.g., calculated using the variable DISTANCE_REWARD_COEF, where DISTANCE_REWARD_COEF represents the distance from the enemy's core position compared to the previous state), a strike reward (e.g., obtained when a laser illuminates or strikes a target using the variable ATTACK_REWARD_COEF), a target damage reward (e.g., obtained each time the enemy inflicts damage on a new unit using the variable DAMAGE_REWARD_COEF), a self-damage reward (e.g., obtained when the self is damaged using the variable DEAD_REWARD), and an end reward (e.g., obtained when a phase / match ends using the variable FINAL_REWARD, where FINAL_REWARD = ALIVE_COEF × number of helicopters + NO_DAMAGE_REWARD × weighted sum of enemy unit types + FINISHED_COEF).

[0112] Policy gradient algorithms are highly sensitive to step size during the update process. Too small a step size leads to slow learning, while too large a step size results in excessive differences between the old and new policies. This causes the next sampling for interaction with the environment to deviate, leading to a vicious cycle where the policy is updated to the wrong position, resulting in a catastrophic performance degradation. Therefore, TrustRegion Policy Optimization (TRPO) uses importance sampling and KL divergence to limit the differences in the distribution of the old and new policies. The objective function is then... It becomes:

[0113] ;

[0114] in, The parameters representing the old strategy, a represents the penalty coefficient. t s represents the action at time t. t Represents the environmental information at time t, KL[] represents the KL divergence calculation function, and E t Expressing hope, Indicates in s t Choose a below t The new probability, Indicates in s t Choose a below t The old probability. To further constrain the change between the old and new policies, the PPO algorithm prunes the ratio of the old and new policies to obtain the objective function of the policy:

[0115] ;

[0116] in, Indicates will Distribution ratio limited to Inside, This indicates the range of change in the upper and lower limits of control. When a t When the value is greater than 0, it indicates that the current action is better than the average, increasing the probability of that action, but the update step size is limited to... ; when a t When the value is less than 0, it indicates that the current action is worse than the average action. Therefore, the probability of this action is reduced, and the step size is updated. Cut off at the point.

[0117] S302: Determine the environmental information at the current moment based on the expanded experience pool.

[0118] S303: Input the environmental information at the current moment into the evaluation network of the first target agent to obtain the value at the current moment.

[0119] S304: Calculate the step discount reward based on the value at the current moment.

[0120] S305: Input the environmental information from the data combination at each moment in the expanded experience pool into the evaluation network of the first target agent to obtain the value at multiple moments.

[0121] S306: A first advantage estimate is obtained based on the step discount reward and the value of the multiple moments.

[0122] The high intensity and frequent situational changes in aerial formation coordinated target allocation and strike operations lead to excessive fluctuations in reward values. This results in reward values ​​failing to accurately reflect the actual merits of each decision, and directly using reward values ​​to guide strategy updates causes significant fluctuations in strategy quality. To more accurately estimate the relative merits of each decision and guide more stable strategy updates, a generalized advantage estimation method is used. This method combines a series of actual reward values ​​and a trained value network to evaluate the situational state, and uses the evaluation values ​​to calculate the actual merits of the decisions made in each state. This application does not limit the calculation method of the advantage estimate; for example, the advantage estimate can be calculated using the following formula:

[0123] ;

[0124] in, Indicates in s t down, a t The degree of good or bad relative to the mean, that is, the deviation of a random variable from the mean; Indicates in s t Next, choose a. t The subsequent action value function; Indicates in s t The state-value function is used. Employing the advantage function helps improve learning efficiency and makes policy learning more stable. It also helps reduce variance and mitigate overfitting.

[0125] The objective function and gradient update rule of the policy gradient combined with the advantage function as follows:

[0126] ;

[0127] ;

[0128] S307: Calculate the first loss value based on the first advantage estimate.

[0129] The PPO algorithm is also based on a policy network and an evaluation network architecture. The policy network is used for action generation, mapping environmental information into a probability distribution of actions. The evaluation network, a state-value network, is used to evaluate the quality of actions. The PPO algorithm improves reward acquisition by optimizing the policy network; that is, it samples the action based on interaction with the environment and adjusts the probability of actions using reward information. If an action yields a good value function, its probability is increased; if it yields a bad value function, its probability is decreased. Therefore, updating the policy incurs a loss. for:

[0130] ;

[0131] in, 'Represents a policy network,' ' represents network parameters, f(s) t a t ) indicates in s t Below a t The assessment.

[0132] S308: Update the network parameters of the evaluation network of the first target agent based on the first loss value through backpropagation.

[0133] S309: Input the environmental information from the data combination at each moment in the expanded experience pool into the policy network and learning network of the first target agent to obtain the first normal distribution and the second normal distribution.

[0134] S310: Obtain importance weights for actions based on the data combinations at each time step in the first normal distribution, the second normal distribution, and the expanded experience pool.

[0135] S311: Calculate the second advantage estimate based on the importance weights.

[0136] S312: Calculate the second loss value based on the second advantage estimate.

[0137] S313: Update the network parameters of the policy network in the first target agent based on the second loss value through backpropagation.

[0138] S314: Update the network parameters of the learning network in the first target agent using the updated network parameters of the policy network in the first target agent to obtain the updated network parameters of the first target agent.

[0139] This application utilizes action-environment interactions generated by a policy network to continuously collect multi-temporal data combinations containing environmental information, actions, and rewards, forming an expanded experience pool. This provides rich and diverse learning samples for subsequent training, enhancing the model's generalization ability. An evaluation network assesses the value of the current and historical moments, and, combined with step-discount rewards, calculates an accurate first-advantage estimate, clarifying the direction of policy updates. The first-advantage estimate is used to construct the loss function of the evaluation network, and its parameters are optimized through backpropagation to improve the accuracy of value estimation. The normal distribution of the policy network and learning network outputs is introduced, and an importance sampling technique is used to calculate a second-advantage estimate, which is then used to generate a second loss value for updating the policy network. This ensures the stability of the policy improvement process and avoids performance fluctuations caused by excessive updates. By synchronizing the updated policy network parameters to the learning network, the alignment of the two in policy distribution is maintained, enhancing the stability of the training process and effectively solving problems such as training instability, low sample utilization, and large policy update bias in traditional reinforcement learning methods.

[0140] Regarding S301, which expands the experience pool based on the policy network in the first target agent to obtain an expanded experience pool, this application provides an optional embodiment:

[0141] Let the value of t be 1.

[0142] The environmental information at time t is obtained and input into the policy network of the first target agent to obtain the Gaussian distribution at time t.

[0143] Obtain environmental information s1 at time t=1 from the simulation environment or sensors, such as: UAV coordinates (x=10km, y=5km); enemy radar activity status (e.g., sub-region A2 is activated); 2 munitions remaining; formation coordination status is in the standby zone.

[0144] Input s1 into the policy network of the first target agent and output a Gaussian distribution in the action space.

[0145] The action at time t is obtained by randomly sampling the Gaussian distribution at time t.

[0146] A specific action a1 is randomly sampled from the Gaussian distribution, such as "fly towards sub-region A2 with a yaw of ±15°". a1 is then sent to the environment for execution.

[0147] The action at time t is interacted with the environment to obtain the reward at time t and the environment information at time t+1.

[0148] The environmental information at time t, the action at time t, and the reward at time t are stored in the experience pool as a data combination.

[0149] Increment the value of t by 1, return to obtain the environmental information at time t, and input the environmental information at time t into the policy network of the first target agent to obtain the Gaussian distribution at time t. Continue until the number of data combinations in the experience pool meets the preset threshold, then end the expansion and obtain the expanded experience pool.

[0150] This application establishes a digital simulation environment to continuously simulate the confrontation between the first and second formations. Based on the confrontation results, the model is dynamically adjusted and trained to ultimately obtain a target collaborative strike allocation strategy with a higher win rate. Compared to traditional models based on constraint solving or empirical knowledge, this not only shortens decision-making time and improves the win rate but also demonstrates better flexibility and adaptability. Furthermore, this application employs a two-stage multi-agent model training method and optimizes the reward function during training, significantly improving iterative convergence speed and learning efficiency. Especially when dealing with multiple mission objectives in urban warfare, in addition to considering the performance factors of the air formation itself, it also fully incorporates complex factors such as enemy troop deployment and operational intentions, overcoming the limitations of traditional control methods and demonstrating broader applicability and robustness. Facing new adversarial environments, it only requires retraining-evaluation-optimization, ensuring the model's efficient application and adaptability in different scenarios.

[0151] Figure 4 A structural diagram of an air formation strike allocation device provided in this application embodiment is shown below. Figure 4 As shown, based on the strike allocation method for an air formation provided in the preceding embodiments, this application also provides a strike allocation device for an air formation, comprising:

[0152] The acquisition module is used to acquire situational information about the target area where the aerial formation is located; the aerial formation includes a first formation or a second formation; both the first formation and the second formation include multiple aircraft.

[0153] The threat data determination module is used to determine threat data based on the situation information; the threat data includes threat values ​​of multiple sub-regions and threat values ​​of multiple objects; the target area includes multiple sub-regions; and each sub-region includes multiple objects.

[0154] The strike allocation control command determination module is used to obtain strike allocation control commands based on the threat data and the situation information using a trained decision model; the trained decision model includes a first formation decision model or a second formation decision model; both the first formation decision model and the second formation decision model include a hybrid network and a multi-agent network; the multi-agent network is trained based on a proximal optimization algorithm; the multi-agent network includes multiple agents; each agent includes a policy network, a learning network, and an evaluation network.

[0155] The control module is used to control the air formation to carry out strikes based on the strike allocation control command.

[0156] As an optional embodiment, the threat data determination module specifically includes:

[0157] The threat indicator system determination unit is used to distinguish between the static attributes and dynamic attributes of the target based on the situation information, and to construct a threat indicator system based on the static attributes and dynamic attributes of the target.

[0158] The preprocessing unit is used to preprocess the threat index system to obtain the processed threat index system; the preprocessing includes dimensionless processing or score-based processing.

[0159] The sub-region threat value determination unit is used to obtain the threat values ​​of multiple sub-regions based on the processed threat index system.

[0160] The threat value determination unit is used to obtain the threat value of each object in each sub-region based on the situational information analysis using Bayesian inference methods.

[0161] The threat data determination unit is used to determine threat data based on the threat values ​​of the multiple sub-regions and the threat value of each object in each sub-region.

[0162] As an optional embodiment, the strike allocation control command determination module specifically includes:

[0163] The normalization unit is used to normalize the threat data and the situation information to obtain normalized threat data and normalized situation information.

[0164] The instruction generation unit is used to input the normalized threat data and the normalized situation information into the trained decision model to obtain the strike allocation control instruction.

[0165] As an optional embodiment, the device further includes:

[0166] The first training module is used to train and update the network parameters of the first target agent using a proximal optimization algorithm to obtain the updated network parameters of the first target agent; the first target agent is any agent in the multi-agent network of the first formation decision model.

[0167] The first update template is used to update the network parameters of agents other than the first target agent in the first formation decision model based on the updated network parameters of the first target agent, so as to obtain the updated first formation decision model.

[0168] The historical situation information determination module is used to acquire the decision model and historical situation information based on the first training data.

[0169] The decision model application module is used to input the historical situation information into the decision model based on the first training data and the updated first formation decision model, respectively, to obtain the first historical strike allocation control command and the second historical strike allocation control command; the first historical strike allocation control command is obtained using the decision model based on the first training data; the second historical strike allocation control command is obtained using the updated first formation decision model.

[0170] The simulated combat module is used to control the first simulated air formation and the second simulated air formation to engage in combat and obtain the first combat result; the first simulated air formation is controlled using the first historical strike allocation control command; the second simulated air formation is controlled using the second historical strike allocation control command.

[0171] The first judgment module is used to return to the first training module if the first adversarial result indicates that the first simulated air formation has won, until the preset number of iterations is met, to obtain the first formation decision model after the first adjustment, and to execute the network parameter adjustment unit.

[0172] The second judgment module is used to adjust the network parameters of the decision model based on the first training data if the first adversarial result indicates that the second simulated air formation has won, and return to the decision model application module until the preset number of iterations is met, so as to obtain the first formation decision model after the first adjustment and execute the network parameter adjustment unit.

[0173] As an optional embodiment, the network parameter adjustment unit specifically includes:

[0174] The second training subunit is used to train and update the network parameters of the second target agent using a proximal optimization algorithm to obtain the updated network parameters of the second target agent; the second target agent is any agent in the multi-agent network of the second formation decision model.

[0175] The second update subunit is used to update the network parameters of agents other than the second target agent in the second formation decision model based on the updated network parameters of the second target agent, so as to obtain the updated second formation decision model.

[0176] Obtain sub-units to acquire decision models based on the second training data.

[0177] The decision model application subunit is used to input the historical situation information into the decision model based on the second training data and the updated second formation decision model, respectively, to obtain the third historical strike allocation control command and the fourth historical strike allocation control command; the third historical strike allocation control command is obtained using the decision model based on the second training data; the fourth historical strike allocation control command is obtained using the updated second formation decision model.

[0178] The simulated combat subunit is used to control the third and fourth simulated air formations to engage in combat and obtain a second combat result; the third simulated air formation is controlled using the third historical strike allocation control command; the fourth simulated air formation is controlled using the fourth historical strike allocation control command.

[0179] The first judgment subunit is used to return to the second training subunit if the second adversarial result indicates that the third simulated air formation has won, until the preset number of iterations is met, to obtain the second formation decision model after the first adjustment, and to use a game theory method to adjust the network parameters of the first formation decision model and the second formation decision model after the first adjustment.

[0180] The second judgment subunit is used to adjust the decision model based on the second training data if the second adversarial result indicates that the fourth simulated air formation has won, and return to the decision model application subunit until the preset number of iterations is met, to obtain the second formation decision model after the first adjustment, and to use a game theory method to adjust the network parameters of the first formation decision model and the second formation decision model after the first adjustment.

[0181] As an optional embodiment, the step of using a game theory method to adjust the network parameters of the first formation decision model and the second formation decision model after the first adjustment specifically includes:

[0182] The historical situation information is input into the first formation decision model and the second formation decision model after the first adjustment, respectively, to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command. The fifth historical strike allocation control command is obtained using the first formation decision model after the first adjustment; the sixth historical strike allocation control command is obtained using the second formation decision model after the first adjustment. The fifth and sixth simulated air formations are controlled to engage in combat, resulting in a third combat outcome. The fifth simulated air formation is controlled using the fifth historical strike allocation control command; the sixth simulated air formation is controlled using the sixth historical strike allocation control command. If the third combat outcome indicates that the fifth simulated air formation has won, the network parameters of the second formation decision model after the first adjustment are trained and updated using a near-end optimization algorithm, and the results are returned. The historical situation information is input into the first formation decision model and the second formation decision model after the first adjustment, respectively, to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command, until the preset number of iterations is met, and the final second formation decision model is obtained; if the third confrontation result indicates that the sixth simulated air formation wins, then the network parameters of the first formation decision model after the first adjustment are trained and updated using a near-end optimization algorithm, and the process of inputting the historical situation information into the first formation decision model and the second formation decision model after the first adjustment, respectively, to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command is returned, until the preset number of iterations is met, and the final first formation decision model is obtained; a trained decision model is constructed based on the final first formation decision model or the final second formation decision model.

[0183] As an optional embodiment, the first training module specifically includes:

[0184] An expansion unit is used to expand the experience pool based on the policy network of the first target agent, obtaining an expanded experience pool; the expanded experience pool includes data combinations from multiple time points; the data combinations include environmental information, actions, and rewards. An environmental information determination unit is used to determine the environmental information of the current time point based on the expanded experience pool. A current time point value determination unit is used to input the environmental information of the current time point into the evaluation network of the first target agent to obtain the value of the current time point. A step-discount reward determination unit is used to calculate a step-discount reward based on the value of the current time point. A multi-time point value determination unit is used to input the environmental information from the data combinations of each time point in the expanded experience pool into the evaluation network of the first target agent to obtain the value of multiple time points. A first advantage estimate determination unit is used to calculate a first advantage estimate based on the step-discount reward and the values ​​of the multiple time points. A first loss value determination unit is used to calculate a first loss value based on the first advantage estimate. A first backpropagation unit is used to update the network parameters of the evaluation network of the first target agent based on the first loss value through backpropagation. A normal distribution determination unit is used to input environmental information from the data combinations at each time step in the expanded experience pool into the policy network and learning network of the first target agent, obtaining a first normal distribution and a second normal distribution. An importance weight determination unit is used to obtain importance weights based on the first normal distribution, the second normal distribution, and the actions in the data combinations at each time step in the expanded experience pool. A second advantage estimate determination unit is used to calculate a second advantage estimate based on the importance weights. A second loss value determination unit is used to calculate a second loss value based on the second advantage estimate. A second backpropagation unit is used to update the network parameters of the policy network in the first target agent based on the second loss value through backpropagation. An updated network parameter determination unit is used to update the network parameters of the learning network in the first target agent using the updated network parameters of the policy network in the first target agent, obtaining the updated network parameters of the first target agent.

[0185] As an optional embodiment, the expansion unit specifically includes:

[0186] Set a sub-unit to set the value of t to 1.

[0187] The Gaussian distribution determines the sub-unit, which is used to acquire environmental information at time t and input the environmental information at time t into the policy network of the first target agent to obtain the Gaussian distribution at time t.

[0188] The action determination subunit is used to randomly sample the Gaussian distribution at time t to obtain the action at time t.

[0189] The environmental information determination subunit is used to interact with the environment at time t to obtain the reward at time t and the environmental information at time t+1.

[0190] The storage sub-unit is used to store the environmental information at time t, the action at time t, and the reward at time t in the form of a data combination into the experience pool.

[0191] An iterative judgment subunit is used to increment the value of t by 1 and return the Gaussian distribution determination subunit until the number of data combinations in the experience pool meets a preset threshold, at which point the expansion ends and the expanded experience pool is obtained.

[0192] This application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a strike allocation method for air formations.

[0193] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a strike allocation method for air formations.

[0194] This application provides a computer program product, including a computer program that, when executed by a processor, implements a strike allocation method for air formations.

[0195] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and equipment embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0196] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for attack allocation in an aerial formation, characterized in that, The strike allocation method for the air formation includes: Acquire situational information about the target area where the aerial formation is located; the aerial formation includes a first formation or a second formation; both the first formation and the second formation include multiple aircraft; Threat data is determined based on the situational information; the threat data includes threat values ​​for multiple sub-regions and threat values ​​for multiple objects; the target area includes multiple sub-regions; each sub-region includes multiple objects; Based on the threat data and situational information, a strike allocation control command is obtained using a trained decision model. The trained decision model includes a first formation decision model or a second formation decision model. Both the first and second formation decision models include a hybrid network and a multi-agent network. The multi-agent network is trained based on a proximal optimization algorithm. The multi-agent network includes multiple agents. Each agent includes a policy network, a learning network, and an evaluation network. The air formation is controlled to carry out strikes based on the strike allocation control command.

2. The strike allocation method for air formations according to claim 1, characterized in that, The determination of threat data based on the situational information specifically includes: Based on the situational information, the static and dynamic attributes of the target are distinguished, and a threat indicator system is constructed based on the static and dynamic attributes of the target. The threat index system is preprocessed to obtain a processed threat index system; the preprocessing includes dimensionless processing or score-based processing. Based on the processed threat index system, threat values ​​for multiple sub-regions are obtained; The threat value of each object in each sub-region is obtained by analyzing the situational information using Bayesian inference methods. The threat data is determined based on the threat values ​​of the multiple sub-regions and the threat value of each object in each sub-region.

3. The strike allocation method for air formations according to claim 1, characterized in that, The step of using the trained decision model to obtain strike allocation control instructions based on the threat data and the situation information specifically includes: The threat data and the situation information are normalized to obtain normalized threat data and normalized situation information; The normalized threat data and the normalized situation information are input into the trained decision model to obtain the strike allocation control command.

4. The strike allocation method for air formations according to claim 1, characterized in that, The method further includes: The network parameters of the first target agent are trained and updated using a proximal optimization algorithm to obtain the updated network parameters of the first target agent; the first target agent is any agent in the multi-agent network of the first formation decision model. Based on the updated network parameters of the first target agent, the network parameters of the agents other than the first target agent in the first formation decision model are updated to obtain the updated first formation decision model. Acquire decision-making models and historical situation information based on the first training data; The historical situation information is input into the decision model based on the first training data and the updated first formation decision model, respectively, to obtain the first historical strike allocation control command and the second historical strike allocation control command; the first historical strike allocation control command is obtained using the decision model based on the first training data; the second historical strike allocation control command is obtained using the updated first formation decision model. The first simulated air formation and the second simulated air formation are controlled to engage in combat, resulting in a first combat outcome. The first simulated air formation is controlled using the first historical strike allocation control command, and the second simulated air formation is controlled using the second historical strike allocation control command. If the first adversarial result indicates that the first simulated air formation has won, then return to the step of training and updating the network parameters of the first target agent using the near-end optimization algorithm to obtain the updated network parameters of the first target agent, until the preset number of iterations is met, to obtain the first formation decision model after the first adjustment, and adjust the network parameters of the first formation decision model and the second formation decision model after the first adjustment; If the first confrontation result indicates that the second simulated air formation has won, then the network parameters of the decision model based on the first training data are adjusted, and the process of inputting the historical situation information into the decision model based on the first training data and the updated first formation decision model is returned to obtain the first historical strike allocation control command and the second historical strike allocation control command is obtained, until the preset number of iterations is met, the first adjusted first formation decision model is obtained, and the network parameters of the first adjusted first formation decision model and the second formation decision model are adjusted.

5. The strike allocation method for air formations according to claim 4, characterized in that, The adjustment of network parameters for the first formation decision model and the second formation decision model after the first adjustment specifically includes: The network parameters of the second target agent are trained and updated using a proximal optimization algorithm to obtain the updated network parameters of the second target agent; the second target agent is any agent in the multi-agent network of the second formation decision model. Based on the updated network parameters of the second target agent, the network parameters of the agents other than the second target agent in the second formation decision model are updated to obtain the updated second formation decision model. Obtain a decision model based on the second training data; The historical situation information is input into the decision model based on the second training data and the updated second formation decision model, respectively, to obtain the third historical strike allocation control command and the fourth historical strike allocation control command; the third historical strike allocation control command is obtained using the decision model based on the second training data; the fourth historical strike allocation control command is obtained using the updated second formation decision model. The third and fourth simulated air formations are controlled to engage in combat, resulting in a second combat outcome. The third simulated air formation is controlled using the third historical strike allocation control command, and the fourth simulated air formation is controlled using the fourth historical strike allocation control command. If the second adversarial result indicates that the third simulated air formation has won, then return to the step of training and updating the network parameters of the second target agent using the near-end optimization algorithm to obtain the updated network parameters of the second target agent, until the preset number of iterations is met, to obtain the second formation decision model after the first adjustment, and use a game theory method to adjust the network parameters of the first formation decision model and the second formation decision model after the first adjustment; If the second confrontation result indicates that the fourth simulated air formation has won, then the decision model based on the second training data is adjusted, and the process of inputting the historical situation information into the decision model based on the second training data and the updated second formation decision model is returned to obtain the third historical strike allocation control command and the fourth historical strike allocation control command, until the preset number of iterations is met, and the first adjusted second formation decision model is obtained. Then, the network parameters of the first adjusted first formation decision model and the first adjusted second formation decision model are adjusted using a game theory method.

6. The strike allocation method for air formations according to claim 5, characterized in that, The process of adjusting the network parameters of the first formation decision model and the second formation decision model after the first adjustment using a game theory method specifically includes: The historical situation information is input into the first formation decision model and the second formation decision model after the first adjustment, respectively, to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command; the fifth historical strike allocation control command is obtained using the first formation decision model after the first adjustment; the sixth historical strike allocation control command is obtained using the second formation decision model after the first adjustment. The fifth and sixth simulated air formations were controlled to engage in combat, resulting in a third combat outcome. The fifth simulated air formation was controlled using the fifth historical strike allocation control command, and the sixth simulated air formation was controlled using the sixth historical strike allocation control command. If the third confrontation result indicates that the fifth simulated air formation has won, then the network parameters of the second formation decision model after the first adjustment are trained and updated using a near-end optimization algorithm, and the steps of inputting the historical situation information into the first formation decision model and the second formation decision model after the first adjustment are returned to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command are obtained, until the preset number of iterations is met, and the final second formation decision model is obtained. If the third confrontation result indicates that the sixth simulated air formation has won, then the network parameters of the first formation decision model after the first adjustment are trained and updated using the near-end optimization algorithm, and the steps of inputting the historical situation information into the first formation decision model and the second formation decision model after the first adjustment respectively to obtain the fifth historical strike allocation control command and the sixth historical strike allocation control command are returned until the preset number of iterations is met to obtain the final first formation decision model. The trained decision model is constructed based on the final first formation decision model or the final second formation decision model.

7. The strike allocation method for air formations according to claim 4, characterized in that, The step of training and updating the network parameters of the first target agent using a proximal optimization algorithm to obtain the updated network parameters of the first target agent specifically includes: The experience pool is expanded based on the policy network in the first target agent to obtain an expanded experience pool; the expanded experience pool includes data combinations from multiple time points; the data combinations include environmental information, actions, and rewards; The environmental information at the current moment is determined based on the expanded experience pool. The environmental information at the current moment is input into the evaluation network of the first target agent to obtain the value at the current moment; The step discount reward is calculated based on the value at the current moment. The environmental information from the data combination at each moment in the expanded experience pool is input into the evaluation network of the first target agent to obtain the value at multiple moments; A first advantage estimate is obtained based on the step discount reward and the value of the multiple moments; The first loss value is calculated based on the first advantage estimate; The network parameters of the evaluation network for the first target agent are updated based on the first loss value through backpropagation. The environmental information from the data combination at each moment in the expanded experience pool is input into the policy network and learning network of the first target agent to obtain the first normal distribution and the second normal distribution; Importance weights are derived from the actions in the data combinations at each time step in the first normal distribution, the second normal distribution, and the expanded experience pool. The second advantage estimate is calculated based on the aforementioned importance weights; The second loss value is calculated based on the second advantage estimate; The network parameters of the policy network in the first target agent are updated based on the second loss value through backpropagation; The network parameters of the learning network in the first target agent are updated by updating the network parameters of the policy network in the first target agent, thus obtaining the updated network parameters of the first target agent.

8. The strike allocation method for air formations according to claim 7, characterized in that, The expansion of the experience pool based on the policy network in the first target agent to obtain the expanded experience pool specifically includes: Let the value of t be 1; Obtain the environmental information at time t and input it into the policy network of the first target agent to obtain the Gaussian distribution at time t; The action at time t is obtained by randomly sampling the Gaussian distribution at time t; The action at time t is interacted with the environment to obtain the reward at time t and the environment information at time t+1. The environmental information at time t, the action at time t, and the reward at time t are stored in the experience pool as a data combination. Increment the value of t by 1, return to obtain the environmental information at time t, and input the environmental information at time t into the policy network of the first target agent to obtain the Gaussian distribution at time t. Continue until the number of data combinations in the experience pool meets the preset threshold, then end the expansion and obtain the expanded experience pool.

9. A strike distribution device for an aerial formation, characterized in that, The strike distribution device for the air formation includes: The acquisition module is used to acquire situational information about the target area where the aerial formation is located; the aerial formation includes a first formation or a second formation; both the first formation and the second formation include multiple aircraft; A threat data determination module is used to determine threat data based on the situational information; the threat data includes threat values ​​of multiple sub-regions and threat values ​​of multiple objects; the target area includes multiple sub-regions; and each sub-region includes multiple objects. The strike allocation control command determination module is used to obtain strike allocation control commands based on the threat data and the situation information using a trained decision model; the trained decision model includes a first formation decision model or a second formation decision model; both the first formation decision model and the second formation decision model include a hybrid network and a multi-agent network; the multi-agent network is trained based on a proximal optimization algorithm; the multi-agent network includes multiple agents; each agent includes a policy network, a learning network, and an evaluation network; The control module is used to control the air formation to carry out strikes based on the strike allocation control command.

10. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the strike allocation method for air formations as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Formation instruction determination method, system and device and storage medium

    CN115826627A

  • Heterogeneous multi-unmanned aerial vehicle strike decision-making method and device based on dynamic Bayesian network

    CN117113216A

  • Decision learning method and device of command agent, equipment and medium

    CN118095339A

  • Multi-subject pursuit optimal strategy method based on threat degree reinforcement learning algorithm

    CN118276438A

  • Multi-unmanned aerial vehicle collaborative adversarial learning method based on reinforcement learning

    CN119443202A