Target matching method and device for unmanned aerial vehicle cluster, terminal and storage medium

By building a drone formation model and using deep offline reinforcement learning and network flow algorithms, the problem of low target matching efficiency in drone autonomous decision-making is solved, and efficient and real-time goal matching decisions are achieved.

CN120335496AActive Publication Date: 2025-07-18HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510806197.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing autonomous decision-making technology of drone needs to rely on a large amount of data to assist decision-making, with low target matching efficiency and low real-time performance.

Method used

The drone formation model is built and trained based on deep offline reinforcement learning, the target advantage and threat of the drone relative to the attack target are calculated, and the optimal matching result is determined using network flow algorithm.

Benefits of technology

It improves the efficiency and real-time performance of target matching of drone clusters, and can converge the optimal drone formation with fewer training samples to achieve efficient target matching decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335496A_ABST
    Figure CN120335496A_ABST
Patent Text Reader

Abstract

The invention provides a target matching method and device for an unmanned aerial vehicle cluster, a terminal and a storage medium, and relates to the technical field of unmanned aerial vehicles, and the method comprises the steps: constructing an unmanned aerial vehicle formation model, and carrying out the training of the unmanned aerial vehicle formation model based on deep offline reinforcement learning, so as to deduce an optimal unmanned aerial vehicle formation; when the unmanned aerial vehicle cluster is controlled to form the optimal unmanned aerial vehicle formation, the target dominance degree of the unmanned aerial vehicle relative to the attack target and the target threat degree of the attack target to the unmanned aerial vehicle are calculated; determining a pairing decision value between the unmanned aerial vehicle and the attack target according to the target dominance degree and the target threat degree, and constructing a weight network based on the pairing decision value; and determining an optimal matching result of the unmanned aerial vehicle matching attack target under the weight network by using a network flow algorithm. The formation strategy in unmanned aerial vehicle cooperative combat is optimized based on the formation model of deep offline reinforcement learning, global optimization of target distribution is realized based on the target matching strategy of the network flow algorithm, and high-efficiency and high-real-time target matching is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a method, device, terminal and storage medium for target matching of an unmanned aerial vehicle cluster. Background Art

[0002] With the progress of aviation technology and the traction of military struggle requirements, since multi-aircraft cooperative air combat has higher combat capabilities and combat efficiency compared with single-aircraft air combat, multi-aircraft cooperative air combat has become the main form and development trend of future air combat.

[0003] In a battlefield environment with complex environmental changes and rapid situation changes, how to implement efficient and flexible autonomous decision-making reasoning based on uncertain situation information and determine a reliable task cooperative execution method is crucial for ensuring the safety of the cluster and improving combat effectiveness. The decision-making problem is described as an optimization game problem of the target under strong non-linear constraints. For this optimization game problem, currently, one or more methods based on expert knowledge, intelligent optimization, artificial intelligence, and game theory are usually used to achieve autonomous decision-making of unmanned aerial vehicles, but all four have certain defects: (1) Autonomous decision-making based on expert knowledge: It has poor adaptability to the environment, requires a large amount of data sets to assist decision-making, and the decision-making accuracy drops significantly when the relevant templates are lacking in the expert knowledge base.

[0004] (2) Autonomous decision-making based on intelligent optimization: When there are more optimization targets, the search efficiency decreases, and it is easy to fall into local optimal solutions and does not have global optimization capabilities.

[0005] (3) Autonomous decision-making based on artificial intelligence: It is difficult to design the reward function, takes a long time to learn from scratch, and requires a large number of samples for training.

[0006] (4) Autonomous decision-making based on game theory: It is difficult to model based on game theory and cannot provide an analysis for a highly dynamic battlefield.

[0007] In summary, the existing unmanned aerial vehicle autonomous decision-making technology needs to rely on a large amount of data to assist decision-making, and has low target matching efficiency and low real-time performance. Therefore, how to provide a solution to the above technical problems is an issue that those skilled in the art need to solve currently. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a method, device, terminal and storage medium for target matching of an unmanned aerial vehicle cluster aiming at the above-mentioned defects of the existing technology, and aims to solve the problems that the existing unmanned aerial vehicle autonomous decision-making technology needs to rely on a large amount of data to assist decision-making, and has low target matching efficiency and low real-time performance.

[0009] The technical solution adopted by the present invention to solve the technical problem is as follows: A target matching method for an unmanned aerial vehicle (UAV) cluster, wherein the method includes: Construct a UAV formation model, and train the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation; When controlling the UAV cluster to form the optimal UAV formation, calculate the target superiority of the UAVs in the UAV cluster relative to the attack target and the target threat degree of the attack target to the UAVs; Determine the pairing decision value between the UAVs and the attack target according to the target superiority and the target threat degree; Construct a weight network based on the pairing decision value, and use the network flow algorithm to determine the optimal matching result of the UAVs matching the attack target under the weight network.

[0010] In one implementation, the constructing the UAV formation model includes: Construct a UAV formation model based on the Actor-Critic network; Wherein, the training the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation includes: In a preset simulation environment, use the initial UAV formation generated based on the expert strategy to conduct multiple confrontations to obtain the corresponding confrontation results; Determine the corresponding state, action and score according to the confrontation results to obtain the training samples for offline reinforcement learning; Based on the preset double-delay deep deterministic policy gradient algorithm and the behavior cloning algorithm, and use the training samples to train the UAV formation model to infer the optimal UAV formation.

[0011] In one implementation, the calculating the target superiority of the UAVs in the UAV cluster relative to the attack target and the target threat degree of the attack target to the UAVs includes: Calculate the angular superiority, distance superiority and speed superiority of the UAVs relative to the attack target; Determine the target superiority of the UAVs relative to the attack target based on the angular superiority, the distance superiority and the speed superiority; Calculate the angular threat degree, distance threat degree and speed threat degree of the attack target to the UAVs; Determine the target threat degree of the attack target to the UAVs based on the angular threat degree, the distance threat degree and the speed threat degree.

[0012] In one implementation, the calculating the angular superiority, distance superiority and speed superiority of the UAVs relative to the attack target includes: Calculate the angular superiority degree, distance superiority degree, and speed superiority degree of the UAV relative to the attack target by using the angular superiority function, distance superiority function, and speed superiority function, respectively; Among them, the angular superiority function is: ; The distance superiority function is: ; ; ; The speed superiority function is: ; Among them, is the angular superiority degree, is the distance superiority degree, is the speed superiority degree, is the target azimuth angle, is the target approach angle, is the relative distance between the UAV and the attack target, is the maximum launch distance of the missile carried on the UAV, is the minimum launch distance of the missile carried on the UAV, is the UAV speed, is the attack target speed.

[0013] In one implementation, calculating the angular threat degree, distance threat degree, and speed threat degree of the attack target to the UAV includes: Calculate the angular threat degree, distance threat degree, and speed threat degree of the UAV relative to the attack target by using the angular threat function, distance threat function, and speed threat function, respectively; Among them, the angular threat function is: ; The distance threat function is: ; The speed threat function is: ; Among them, is the angular threat degree, is the distance threat degree, is the speed threat degree, is the target azimuth angle, is the target approach angle, is the relative distance between the UAV and the attack target, is the maximum launch distance of the missile carried on the UAV, is the attack distance of the attack target, is the maximum tracking distance of the UAV, is the UAV speed, is the attack target speed.

[0014] In one implementation manner, determining the pairing decision value between the unmanned aerial vehicle and the attack target according to the target superiority degree and the target threat degree includes: Calculating the target superiority degree and the target threat degree by using a decision function to determine the pairing decision value between the unmanned aerial vehicle and the attack target; The decision function is: ; Wherein, is the pairing decision value, is the first preference weight, is the second preference weight, is the target superiority degree, is the target threat degree.

[0015] In one implementation manner, constructing the weight network based on the pairing decision value includes: Taking the pairing decision value between the unmanned aerial vehicle and the attack target as the weight of the edge, and establishing a corresponding weighted bipartite graph to obtain the weight network; Wherein, the method for determining the optimal matching result of the unmanned aerial vehicle matching the attack target under the weight network by using the network flow algorithm includes: Constructing a network flow model based on the weight network; wherein, the edge capacity is 1, and the cost is the negative of the pairing decision value; Using the minimum cost maximum flow algorithm and solving the minimum cost maximum flow based on the network flow model to obtain the optimal matching result of the unmanned aerial vehicle matching the attack target.

[0016] The present invention also discloses a target matching device for an unmanned aerial vehicle cluster, wherein the device includes: A formation model construction module, configured to construct an unmanned aerial vehicle formation model; A formation model training module, configured to train the unmanned aerial vehicle formation model based on deep offline reinforcement learning to infer the optimal unmanned aerial vehicle formation; A matching index calculation module, configured to calculate the target superiority degree of the unmanned aerial vehicle in the unmanned aerial vehicle cluster relative to the attack target and the target threat degree of the attack target to the unmanned aerial vehicle when controlling the unmanned aerial vehicle cluster to form the optimal unmanned aerial vehicle formation; A decision value determination module, configured to determine the pairing decision value between the unmanned aerial vehicle and the attack target according to the target superiority degree and the target threat degree; A matching determination module, configured to construct a weight network based on the pairing decision value, and use a network flow algorithm to determine the optimal matching result of the unmanned aerial vehicle matching the attack target under the weight network.

[0017] The present invention also discloses a terminal, which includes: a memory, a processor, and a target matching program of a drone cluster stored on the memory and executable on the processor. When the target matching program of the drone cluster is executed by the processor, the steps of the target matching method of the drone cluster as described above are implemented.

[0018] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the target matching method of the drone cluster as described above.

[0019] A target matching method, device, terminal, and storage medium for a drone cluster provided by the present invention. The target matching method for the drone cluster includes: constructing a drone formation model, and training the drone formation model based on deep offline reinforcement learning to infer an optimal drone formation; when controlling the drone cluster to form the optimal drone formation, calculating the target advantage degree of the drones in the drone cluster relative to the attack target and the target threat degree of the attack target to the drones; determining a pairing decision value between the drones and the attack target according to the target advantage degree and the target threat degree; constructing a weight network based on the pairing decision value, and using a network flow algorithm to determine an optimal matching result of the drones matching the attack target under the weight network. It can be seen that for the target matching problem in the autonomous decision-making of drones, the present invention first performs pre-arrangement of the drone formation based on deep offline reinforcement learning, and can train and converge to an optimal drone formation with fewer training samples. Then, when the drone cluster forms the optimal drone formation, the pairing decision value between the drones and the attack target is calculated, and then a weight network is constructed based on the pairing decision value. Under the weight network, a network flow algorithm is used for target matching, which can improve the efficiency and real-time performance of target matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a preferred embodiment of the target matching method for a drone cluster in the present invention; Figure 2 is a schematic diagram of a workflow for establishing an optimal drone formation by a specific deep offline reinforcement learning in the present invention; Figure 3 is a schematic diagram of the architecture of the BC+TD3 algorithm in the present invention; Figure 4 is a schematic diagram of the situation of drones and attack targets in the present invention; Figure 5 is a schematic diagram of a specific network flow model disclosed in the present invention; Figure 6 is a flowchart of a specific target matching method for a drone cluster disclosed in the present invention; Figure 7 It is a functional principle block diagram of a preferred embodiment of the target matching device for an unmanned aerial vehicle (UAV) cluster in the present invention; Figure 8 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. Specific embodiments

[0021] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further elaborates on the present invention by way of examples with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0022] Currently, one or more methods based on expert knowledge, intelligent optimization, artificial intelligence, and game theory are usually used to achieve autonomous decision-making for UAVs. However, all four have certain defects: (1) Autonomous decision-making based on expert knowledge, that is, the expert knowledge theory of using existing experience and knowledge for matching and optimization: For combat experience that cannot be described by a mathematical model, a decision-making method based on expert knowledge is formed by establishing military rules and combat rule bases. When making autonomous decisions based on expert knowledge, the inference system matches the perceived battlefield combat information with the combat rules in the knowledge base to find the corresponding combat rules for combat. However, this decision-making method has poor adaptability to the environment, requires a large amount of data sets to assist in decision-making, and the decision accuracy drops significantly when the relevant templates are lacking in the expert knowledge base.

[0023] (2) Autonomous decision-making based on intelligent optimization, that is, the intelligent optimization theory of using bionics characteristics to reduce the search complexity: Through the study of the collective behavior of organisms in nature, the cooperative foraging, hunting and other behaviors of biological clusters are mapped to the clusters of the studied carriers, thus deriving a large number of intelligent optimization algorithms. The intelligent optimization algorithm transforms the target solving problem into an optimization problem and solves it by drawing on bionics knowledge. The key to designing an intelligent optimization algorithm lies in the selection of the performance function and the optimization algorithm. However, when there are many optimization targets, the search efficiency decreases, and it is easy to fall into local optimal solutions and does not have global optimization capabilities.

[0024] (3) Autonomous decision-making based on artificial intelligence, that is, the artificial intelligence theory with autonomous learning ability: Reinforcement learning realizes the interaction between the agent and the environment through the behavior mechanism of reward and punishment. The UAV and the environment achieve intelligent autonomous decision-making through continuous interaction and learning. That is, reinforcement learning is a machine learning in which the system learns from the environment to maximize the reward. The reinforcement learning model learns through continuous trial and error and feedback, and is often used for sequential decision-making or control problems, such as game AI and UAVs. However, the design of the reward function is difficult, and learning from scratch takes a long time and requires a large number of samples for training.

[0025] (4) Game-based autonomous decision-making, namely the game theory applicable to highly adversarial problems: Game theory is a theory that studies how to make decisions when decision-making participants influence each other. In the autonomous decision-making of unmanned aerial vehicles (UAVs), there will be games between UAVs and friendly UAVs or enemy targets. Game-based autonomous decision-making can make decisions to improve the combat level for UAVs in complex environments. There will be cooperative games between UAVs and friendly UAVs in issues such as resource management, task allocation, and path planning. It is necessary to select the optimal game strategy to achieve the optimal of individuals or groups. However, it is relatively difficult to model based on game theory, and it cannot provide an analysis for the highly dynamic battlefield.

[0026] As can be seen from the above, the existing UAV autonomous decision-making technology needs to rely on a large amount of data for auxiliary decision-making, and the target matching efficiency is low and the real-time performance is low. For this reason, this application provides a target matching scheme for UAV clusters, which can train and converge to the optimal UAV formation with fewer training samples, and can improve the efficiency and real-time performance of target matching.

[0027] Please refer to Figure 1 , Figure 1 which is the flowchart of the target matching method for UAV clusters in the present invention. As Figure 1 shown, the target matching method for UAV clusters described in the embodiments of the present invention includes: Step S11, construct a UAV formation model, and train the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation.

[0028] In this embodiment, the optimal UAV formation is trained and inferred by constructing a UAV formation model. Specifically, a UAV formation model is constructed based on the Actor-Critic network, and then the UAV formation model is trained based on deep offline reinforcement learning to infer the optimal UAV formation. Among them, using deep offline reinforcement learning can obtain better training results with fewer samples in the case where it is difficult to obtain training samples. And deep reinforcement learning means using a deep learning model (such as a deep neural network) as a function approximator to solve the reinforcement learning problem, which enables RL to handle high-dimensional state spaces, such as images, texts, and action spaces. The core feature of offline reinforcement learning is that the agent only learns the optimal policy from a pre-collected, fixed and unchanging static dataset (generated by historical behavior policies) and does not interact with the environment during the learning process. Therefore, deep offline reinforcement learning refers to a reinforcement learning method that uses a deep neural network as a function approximator to learn the optimal policy from a pre-collected static dataset that does not involve environmental interaction.

[0029] In this embodiment, the UAV formation model is trained based on deep offline reinforcement learning to infer the optimal UAV formation. Specifically, it may include: in a preset simulation environment, using the initial UAV formation generated based on the expert policy to conduct multiple confrontations to obtain the corresponding confrontation results; determining the corresponding states, actions, and scores according to the confrontation results to obtain training samples for offline reinforcement learning; based on the preset twin-delayed deep deterministic policy gradient algorithm and behavior cloning algorithm, and using the training samples to train the UAV formation model to infer the optimal UAV formation. Among them, the training samples are datasets generated by human expert policies manually performing tasks in a machine learning task environment, usually having obvious behavioral strategies, and are generally used in supervised learning or imitation learning tasks. In addition to the expert policy, the initial UAV formation can also be generated based on the standard policy.

[0030] For example, as shown in Figure 2 , in a preset simulation environment, multiple confrontations are carried out using the initial UAV formation generated based on the expert policy, the confrontation results are collected, summarized into a series of states, actions, and scores, and these data are recorded in the experience pool. Then, the experience pool data in the experience pool is input into the training. Specifically, the BC+TD3 algorithm is used, that is, based on the preset twin-delayed deep deterministic policy gradient algorithm and behavior cloning algorithm, and using the experience pool data to train the UAV formation model constructed based on the Actor-Critic network to infer the optimal UAV formation. Among them, the inferred formation can also be input into the simulation again. The architecture of the BC+TD3 algorithm is shown in Figure 3 .

[0031] It should be noted that the TD3 algorithm and the BC algorithm are adopted during the training process. Among them, the TD3 (Twin Delayed Deep Deterministic policy gradient) algorithm is an online off-policy deep reinforcement learning algorithm for solving continuous control problems, which is improved on the basis of the DDPG (Deep Deterministic Policy Gradient) algorithm. Essentially, the TD3 algorithm integrates the idea of the Double Q-Learning algorithm into the DDPG algorithm, mainly to solve the overestimation problem of the DDPG algorithm. Compared with the DDPG algorithm, the improvements of the TD3 algorithm are as follows: First, two sets of Critic networks are adopted during the training process, and the smaller value of the two is taken when calculating the target value, so as to suppress the overestimation problem of the network; Second, during the process of calculating the target value, the algorithm adds a random perturbation value to the action in the next state, so as to make the value evaluation more accurate; Third, after the Critic network is updated multiple times, the Actor network is updated, rather than updating the Actor network every time the Critic network is updated, so as to ensure the more stable training of the Actor network.

[0032] Among them, the network structure of the Actor network is shown in Table 1: Table 1

[0033] The network structure of the Critic (Q1, Q2) network is shown in Table 2: Table 2

[0034] And the specific implementation of network update based on the TD3 algorithm is as follows: First, the update process of the Critic1 and Critic2 networks is as follows: First, the action under the state can be calculated by using the Target Actor network, that is: ; Among them, is the Target Actor network, is the Target Actor network internal learnable parameters such as weights and biases.

[0035] Then, based on the target policy smoothing regularization, noise is added to the target action , that is: ; Among them, , c is the boundary value of the clip truncation function, is the standard deviation of the noise, represents the normal distribution.

[0036] Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the parameters in the Critic1 and Critic2 networks, that is: ; Among them, represents the th loss function of the Critic network, represents the th Q value evaluated by the Critic network for the state-action , represents the th parameter of the Critic network, represents the target Q value.

[0037] Second, the update process of the Actor network is as follows: After the Critic1 and Critic2 networks are updated for d steps, the Actor network update is started. Using the Actor network, the action under the state s can be calculated, that is: ; Among them, represents the Actor network, represents the parameters of the Actor network.

[0038] At this time, no noise needs to be added after calculating the action, so that the Actor network can be updated in the direction of the maximum value. Adding noise has no meaning. Then, use the Critic1 network or the Critic2 network to calculate the evaluation value of the state-action pair , such as using the Critic1 network for evaluation, that is: ; Among them, represents the Q value evaluated by the Critic1 network for the state-action , represents the parameters of the Critic1 network.

[0039] Finally, the gradient ascent algorithm is used to maximize , thereby completing the update of the Actor network. Among them, any one of the Critic1 network and the Critic2 network can be used to calculate the Q value.

[0040] Third, the update process of the target network is: The target network is updated using a soft update method. A learning rate is introduced to perform a weighted average of the old target network parameters and the new corresponding network parameters, and then assigned to the target network, namely: ; ; Among them, the learning rate It can be 0.005.

[0041] Based on the original TD3 algorithm, the TD3 algorithm is combined with BC (Behavior Cloning), that is, the TD3+BC algorithm. By changing the Actor learning goal, compared with the previous method of increasing the amount of calculation and introducing redundant hyperparameters, it achieves better results at the same level of calculation as the TD3 algorithm.

[0042] Among them, the behavior cloning algorithm is a mainstream algorithm in imitation learning, and the other algorithm is adversarial imitation learning. Imitation learning was originally designed to enable intelligent agents to learn decisions from expert data sets, so that intelligent agents can perform some tasks like humans without complex conditional constraints. As a training method using expert data sets, similar to offline learning, it is possible that the initial strategy will accidentally go to a direction that does not exist in the data set. The behavior cloning algorithm can be seen as minimizing the difference between strategy actions. It is expected that the strategy can restore the expert's decision-making behavior well from the expert example, so that the decision maker's value function is relatively large and the strategy does not deviate from the expert strategy as much as possible.

[0043] The original TD3 learning objectives for policy network updates are as follows: ; in, Represents the data set State-Action Sampling Take the expected value, Represents state s and its actions in the Actor network The corresponding Q value.

[0044] The TD3+BC algorithm has made the following improvements on this basis: ; in, Indicates the impact weight of BC loss.

[0045] The first is to introduce a regularization term: The idea of the BC+TD3 algorithm is that if the current state s is in the data set, then the tuple consisting of the action a selected according to the strategy , where s must be in the dataset, then the training strategy and the behavior strategy As long as it is ensured that within a certain distance, the generated tuples also have a relatively high probability density in the dataset. Therefore, based on the idea of behavior cloning, the BC+TD3 algorithm uses the L2 loss to constrain the strategy The selected action from the action a in the dataset.

[0046] Second is to add weights: A new BC loss is introduced, and the influence degree of the BC loss needs to be constrained by weights. Among them, regarding the weights value, since the regularization term is the L2 loss constraint, after normalization, the actual loss size is very small, but the scale of the previous Q greatly affects the size of Q. Therefore, an absolute value averaging method based on the batchsize is introduced to ensure the size of Q, that is: ; Among them, represents the scaling coefficient, represents the batch size.

[0047] Step S12, when controlling the UAV cluster to form the optimal UAV formation, calculate the target superiority of the UAVs in the UAV cluster relative to the attack target and the target threat degree of the attack target to the UAVs.

[0048] It should be noted that the UAV cluster is an autonomous aerial intelligent system composed of multiple UAVs through intelligent algorithms and cooperative control strategies, which can realize information interaction, task cooperation and dynamic environment adaptation, so as to efficiently complete complex tasks.

[0049] In this embodiment, in the UAV cluster confrontation stage, control the UAV cluster to form the optimal UAV formation inferred by deep offline reinforcement learning, and calculate the target superiority of the UAVs in the UAV cluster relative to the attack target and the target threat degree of the attack target to the UAVs through situation assessment and analysis. It can be understood that before target allocation and maneuver decision-making, the situation assessment and analysis results can be used to judge whether the UAV is in an advantageous or disadvantageous state and other information affecting matching, so that the UAV can select a suitable maneuver decision and resource allocation. Among them, the UAV can be our own UAV, and the attack target can be the enemy UAV. Therefore, when controlling the UAV cluster to form the optimal UAV formation, calculate the target superiority of our own UAVs relative to the enemy UAVs and the target threat degree of the enemy UAVs to our own UAVs.

[0050] In this embodiment, calculating the target superiority degree of the unmanned aerial vehicles (UAVs) in the UAV cluster relative to the attack target specifically includes: calculating the angular superiority degree, distance superiority degree, and speed superiority degree of the UAVs relative to the attack target; and determining the target superiority degree of the UAVs relative to the attack target based on the angular superiority degree, distance superiority degree, and speed superiority degree.

[0051] For example, referring to Figure 4 as shown, the angular superiority degree, distance superiority degree, and speed superiority degree of the UAVs relative to the attack target are calculated using the angular superiority function, distance superiority function, and speed superiority function respectively; where F represents the UAV, T represents the attack target, and R is the target distance; is the UAV speed, is the attack target speed, the target line of sight FT, i.e., the line connecting the UAV and the attack target, is the target azimuth angle, is the target approach angle, and it is stipulated that the target azimuth angle and target approach angle are positive for right deviation and negative for left deviation, i.e., , .

[0052] Moreover, in order to achieve effective tracking of the target, it is required to maintain the target azimuth angle. At the same time, in order to avoid being attacked, the best target approach angle is 180°. Therefore, the angular superiority function can be: ; where, when and , , the angular superiority degree reaches the maximum value, forming a situation of attacking from the rear; while when and , , the angular superiority degree reaches the minimum value, and at this time the UAV is bitten by the attack target and is at a disadvantage in escaping.

[0053] It should be noted that assuming the minimum launch distance of the missile carried on the UAV is , and the maximum launch distance is , when the relative distance between the attack target and the UAV , the distance superiority is considered zero; as the relative distance decreases, the distance superiority gradually increases, and at , the distance superiority is considered to reach the maximum value; as the relative distance further decreases, the distance superiority gradually decreases again. Therefore, the distance superiority function can be: ; where, at , let , and we can get: .

[0054] Moreover, the higher the UAV speed, the greater the advantage. Therefore, the speed superiority function can be: ; In this embodiment, the angular advantage degree, distance advantage degree, and speed advantage degree can be specifically weighted and summed to calculate the target advantage degree of the UAV relative to the attack target, that is: ; Among them, , , are the weighting coefficients of the angular advantage, distance advantage, and speed advantage respectively, which can be determined through multiple simulations in the UAV swarm confrontation simulation platform.

[0055] In this embodiment, calculating the target threat degree of the attack target to the UAV specifically includes: calculating the angular threat degree, distance threat degree, and speed threat degree of the attack target to the UAV; determining the target threat degree of the attack target to the UAV based on the angular threat degree, distance threat degree, and speed threat degree.

[0056] For example, as shown in the above Figure 4 , the angular threat degree, distance threat degree, and speed threat degree of the UAV relative to the attack target are calculated by using the angular threat function, distance threat function, and speed threat function respectively.

[0057] Among them, the angular threat function is: ; The distance threat function is: ; The speed threat function is: ; Among them, is the angular threat degree, is the distance threat degree, is the speed threat degree, is the maximum launch distance of the missile carried on the UAV, is the attack distance of the attack target, is the maximum tracking distance of the UAV.

[0058] And the target threat degree of the attack target to the UAV can be calculated by weighted summing the angular threat degree, distance threat degree, and speed threat degree, that is: ; Among them, , , are the weighting coefficients of the angular threat index, distance threat index, and speed threat index respectively, which can be determined through multiple simulations in the UAV swarm confrontation simulation platform.

[0059] Step S13: Determine the pairing decision value between the UAV and the attack target according to the target advantage degree and the target threat degree.

[0060] In this embodiment, after calculating the target superiority of the UAV relative to the attack target and the target threat degree of the attack target to the UAV in the above step S12, the pairing decision value between the UAV and the attack target can be determined according to the target superiority and the target threat degree. Specifically, the decision function is used to calculate the target superiority and the target threat degree to determine the pairing decision value between the UAV and the attack target; The decision function is: ; where, is the pairing decision value, is the first preference weight, is the second preference weight, is the target superiority, is the target threat degree.

[0061] It should be noted that, and reflect the trade-off criterion of the combat department between attack and avoidance. When takes the maximum value, it means giving priority to matching the target with a large target superiority of the UAV relative to the attack target, that is, the target with a higher hit probability; When taking the maximum value, it means giving priority to matching the attack target that poses a greater threat to the UAV.

[0062] Step S14: Construct a weight network based on the pairing decision value, and use the network flow algorithm to determine the optimal matching result of the UAV matching the attack target under the weight network.

[0063] In this embodiment, the pairing decision value between the UAV and the attack target is calculated according to the target superiority of the UAV relative to the attack target and the target threat degree of the attack target to the UAV. Then, a weight network is constructed based on the pairing decision value, and finally, the network flow algorithm is used to determine the optimal matching result of the UAV matching the attack target under the weight network.

[0064] In this embodiment, constructing a weight network based on the pairing decision value may specifically include: using the pairing decision value between the UAV and the attack target as the weight of the edge, and establishing a corresponding weighted bipartite graph to obtain the weight network.

[0065] Moreover, the optimal matching result of the UAVs attacking the target in the weight network is determined by using the network flow algorithm, which specifically includes: constructing a network flow model based on the weight network; where the edge capacity is 1 and the cost is the negative of the pairing decision value; using the minimum cost maximum flow algorithm and solving the minimum cost maximum flow based on the network flow model to obtain the optimal matching result of the UAVs attacking the target. It can be understood that real-time performance is very important during the confrontation process. Using the network flow algorithm for target matching greatly improves the efficiency and real-time performance of target matching. Moreover, for a target enemy aircraft, the UAV with the largest pairing decision value can be used to match the attacking target. However, at the same time, it is necessary to consider the target matching principle of avoiding multiple UAVs matching the same attacking target and avoiding falling into the local optimal solution. The network flow algorithm can be used to solve the problems of repeated matching and local optimal solution in target allocation.

[0066] For example, taking four of our UAVs , and four enemy UAVs as an example, construct a cost function, that is: ; This cost function is used to describe the cost of our UAV attacking the enemy UAV . Specifically, since maximizing the decision function is equivalent to minimizing the cost, the cost function can be equivalent to , where is the decision function for our UAV to attack the enemy UAV

[0067] To ensure that when the numbers on both sides are equal, each UAV has a different attacking target as much as possible, construct an index function , which can only return 0 or 1. If it returns 0, it means the enemy UAV is not attacked. If it returns 1, it means that at least one of our UAVs has the attacking target . Therefore, to solve the optimal target matching, the following target equation system can be considered, that is: ; To solve the above target equation system, the maximum flow minimum cost flow model can be considered, and a network flow model is constructed. As shown in Figure 5 , in the network flow model, the edge capacity of 1 restricts the situation where multiple of our UAVs lock on to one enemy aircraft. Through the maximum flow algorithm, the in the target equation system can be satisfied. At the same time, for the cost per unit capacity, it depends on the cost between the two. Using the minimum cost algorithm can satisfy After obtaining the maximum flow from this equation, the matching result can be obtained by judging whether the capacity of the edge is 0.

[0068] It should be noted that the situation where the number of surviving enemy drones and the number of surviving our drones are not 4 is considered. If the numbers on both sides are equal, only the corresponding number of shot-down drones need to be subtracted from both sides in the network flow model; if the number of surviving our drones is less than the number of surviving enemy drones, similarly, the corresponding number of shot-down drones can be subtracted from both sides in the network flow model; if the number of surviving our drones is more than the number of surviving enemy drones, there will be a situation where multiple our drones attack one drone. In this case, only the edge capacity in the above network flow model needs to be modified to the number of surviving our drones minus the number of surviving enemy drones.

[0069] It can be seen that in the embodiments of the present invention, deep offline reinforcement learning and network flow technology are combined. That is, the formation model based on deep offline reinforcement learning optimizes the formation strategy in the process of UAV cooperative combat, and the target matching strategy based on the network flow algorithm realizes the global optimum of target allocation, achieving high-performance target matching. That is, for the target matching problem in UAV autonomous decision-making, first, the UAV formation pre-arrangement based on deep offline reinforcement learning is performed, which can use fewer training samples to train and converge to the optimal UAV formation. Then, when the UAV cluster forms the optimal UAV formation, the pairing decision value between the UAV and the attack target is calculated. Furthermore, a weight network is constructed based on the pairing decision value. Under the weight network, the network flow algorithm is used for target matching, which can improve the efficiency and real-time performance of target matching.

[0070] For example, the technical solution of the present application is applicable to the confrontation environment of UAV clusters, such as the UAV cluster confrontation simulation platform. And when UAVs are applied in the military, the technical solution of the present application is also applicable to the actual UAV battlefield confrontation, which can improve the autonomous combat ability of UAV clusters and is crucial for achieving air strategic advantages. Through the above technical solution of the present application, a way of firepower allocation can be provided for UAV clusters, and the global information of the environment where the UAV clusters are located is combined to allocate attack targets, that is, target matching objects, for each UAV in the UAV clusters, thereby effectively improving the combat ability of UAV clusters. Among them, compared with one-on-one air combat, the most significant difference in multi-aircraft air combat is that in the face of multiple enemy targets, it is necessary to allocate targets and firepower for each friendly aircraft according to our resources, and air combat situation assessment is a prerequisite for target allocation and maneuver decision-making. In air combat, through the situation assessment result, it can be judged whether our side is in an advantageous or disadvantageous state and other battlefield information at present, and then the maneuver decision suitable for the current battle situation and the allocation of resources can be selected.

[0071] See Figure 6As shown, in the training stage, first, pre - arrangement of combat formations based on deep offline reinforcement learning is carried out. That is, in a preset simulation environment, multiple confrontations are carried out using the initial UAV formation generated based on expert strategies to obtain corresponding confrontation results. According to the confrontation results, the corresponding states, actions, and scores are determined to obtain training samples for offline reinforcement learning. Based on the preset double - delay deep deterministic policy gradient algorithm and behavior cloning algorithm, and using the training samples to train the UAV formation model to infer the optimal UAV formation. This training method can converge to a better formation without relying on a large number of combat data samples for pre - training in a short time. Subsequently, in the UAV cluster confrontation stage, when controlling the UAV cluster to form the optimal initial formation inferred by deep offline reinforcement learning, various degrees of superiority and threat are calculated through situation analysis. According to the degrees of superiority and threat, the pairing decision values of the target UAV pairs to be paired are calculated. A weight network is constructed using the pairing decision values, and the optimal target matching result under the weight network is determined using the network flow algorithm, that is, to achieve target matching based on situation data. The pairing decision values calculated based on the current battle situation information can be used for target matching through the network flow algorithm, and at the same time, decisions on strike targets are made, without relying on a high - computing - power computing platform, and can also perform real - time target matching tasks with high performance. That is to say, through the formation tactics in multi - level cooperative air combat and the UAV formation model based on offline deep reinforcement learning in this application, the UAV cluster formation based on offline deep reinforcement learning is realized. Then, through the basic analysis of air combat situation, the analysis of attack superiority and target threat, and the target assignment for searching the enemy based on the network flow algorithm, the situation assessment and target assignment in air combat decision - making are realized.

[0072] In one embodiment, as Figure 7 shown, based on the above - mentioned target matching method for UAV clusters, the present invention also correspondingly provides a target matching device for UAV clusters, including: A formation model construction module 11, configured to construct a UAV formation model.

[0073] A formation model training module 12, configured to train the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation.

[0074] A matching index calculation module 13, configured to calculate the target superiority of the UAVs in the UAV cluster relative to the attack target and the target threat of the attack target to the UAVs when controlling the UAV cluster to form the optimal UAV formation.

[0075] A decision value determination module 14, configured to determine the pairing decision value between the UAV and the attack target according to the target superiority and the target threat.

[0076] A matching determination module 15, configured to construct a weight network based on the pairing decision value, and use a network flow algorithm to determine an optimal matching result of the UAV to match the attack target under the weight network.

[0077] In addition, it is worth noting that the working process of the target matching device for a UAV cluster provided in this embodiment is the same as that of the target matching method for the UAV cluster described above, and will not be elaborated here. Specifically, reference can be made to the working process of the target matching method for the UAV cluster described above.

[0078] Figure 8 The following is a schematic structural diagram of the terminal provided in the embodiment of the present application. The terminal may include: A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0079] When the processor 502 executes the program, it implements the target matching method for the UAV cluster provided in the above embodiment.

[0080] Furthermore, the terminal further includes: A communication interface 503, configured for communication between the memory 501 and the processor 502.

[0081] The memory 501 is used to store a computer program executable on the processor 502.

[0082] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0083] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0084] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.

[0085] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0086] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the target matching method for the drone cluster as described above is implemented.

[0087] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0088] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0089] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices.

[0090] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0091] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A target matching method for an unmanned aerial vehicle cluster, characterized in that, The method includes: Constructing a UAV formation model and training the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation; When controlling the UAV swarm to form the optimal UAV formation, calculating the target superiority degree of the UAVs in the UAV swarm relative to the attack target and the target threat degree of the attack target to the UAVs; Determining the pairing decision value between the UAVs and the attack target according to the target superiority degree and the target threat degree; Constructing a weight network based on the pairing decision value and using the network flow algorithm to determine the optimal matching result of the UAVs matching the attack target under the weight network.

2. The target matching method for the UAV cluster according to claim 1, wherein The constructing of the UAV formation model includes: Constructing a UAV formation model based on the Actor-Critic network; Wherein, the training of the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation includes: In a preset simulation environment, using the initial UAV formation generated based on the expert strategy to conduct multiple confrontations to obtain the corresponding confrontation results; Determining the corresponding states, actions and scores according to the confrontation results to obtain training samples for offline reinforcement learning; Based on the preset double-delay deep deterministic policy gradient algorithm and the behavior cloning algorithm, and using the training samples to train the UAV formation model to infer the optimal UAV formation.

3. The target matching method for the drone swarm according to claim 1, wherein The calculating of the target superiority degree of the UAVs in the UAV swarm relative to the attack target and the target threat degree of the attack target to the UAVs includes: Calculating the angular superiority degree, distance superiority degree and speed superiority degree of the UAVs relative to the attack target; Determining the target superiority degree of the UAVs relative to the attack target based on the angular superiority degree, the distance superiority degree and the speed superiority degree; Calculating the angular threat degree, distance threat degree and speed threat degree of the attack target to the UAVs; Determining the target threat degree of the attack target to the UAVs based on the angular threat degree, the distance threat degree and the speed threat degree.

4. The method for target matching of an unmanned aerial vehicle cluster according to claim 3, wherein The calculating of the angular superiority degree, distance superiority degree and speed superiority degree of the UAVs relative to the attack target includes: Calculating the angular superiority degree, distance superiority degree and speed superiority degree of the UAVs relative to the attack target by using the angular superiority function, distance superiority function and speed superiority function respectively; wherein, the angular advantage function is as follows: ; The distance advantage function is as follows: ; ; ; The speed advantage function is as follows: ; Among them, is the angular advantage degree, is the distance advantage degree, is the speed advantage degree, is the target azimuth angle, is the target entry angle, is the relative distance between the UAV and the attack target, is the maximum launch distance of the missile carried on the UAV, is the minimum launch distance of the missile carried on the UAV, is the UAV speed, is the attack target speed.

5. The method for target matching of the UAV cluster according to claim 3, wherein The calculating of the angular threat degree, distance threat degree and speed threat degree of the attack target to the UAVs includes: Calculating the angular threat degree, distance threat degree and speed threat degree of the UAVs relative to the attack target by using the angular threat function, distance threat function and speed threat function respectively; Among them, the angle threat function is as follows: ; The distance threat function is as follows: ; The speed threat function is as follows: ; Among them, is the angular threat level, is the distance threat level, is the speed threat level, is the target azimuth angle, is the target approach angle, is the relative distance between the UAV and the attack target, is the maximum launch distance of the missile carried on the UAV, is the attack distance of the attack target, is the maximum tracking distance of the UAV, is the UAV speed, is the attack target speed.

6. The target matching method for the UAV cluster according to claim 1, wherein The determining of the pairing decision value between the UAVs and the attack target according to the target superiority degree and the target threat degree includes: Using the decision function to calculate the target superiority degree and the target threat degree to determine the pairing decision value between the UAVs and the attack target; The decision function is as follows: ; Among them, is the pairing decision value, is the first preference weight, is the second preference weight, is the target superiority degree, is the target threat degree.

7. The target matching method for the UAV cluster according to any one of claims 1 to 6, characterized in that The constructing of the weight network based on the pairing decision value includes: Taking the pairing decision value between the UAVs and the attack target as the weight of the edge and establishing a corresponding weighted bipartite graph to obtain the weight network; Among them, determining the optimal matching result of the UAVs matching attack targets under the weight network by using the network flow algorithm includes: Constructing a network flow model based on the weight network; wherein, the edge capacity is 1, and the cost is the negative of the pairing decision value; Using the minimum cost maximum flow algorithm and solving the minimum cost maximum flow based on the network flow model to obtain the optimal matching result of the UAVs matching attack targets.

8. A target matching device for a drone swarm, characterized in that, The device includes: A formation model construction module, configured to construct a UAV formation model; A formation model training module, configured to train the UAV formation model based on deep offline reinforcement learning to infer the optimal UAV formation; A matching index calculation module, configured to calculate the target superiority of the UAVs in the UAV cluster relative to the attack targets and the target threat degree of the attack targets to the UAVs when controlling the UAV cluster to form the optimal UAV formation; A decision value determination module, configured to determine the pairing decision value between the UAVs and the attack targets according to the target superiority and the target threat degree; A matching determination module, configured to construct a weight network based on the pairing decision value and use the network flow algorithm to determine the optimal matching result of the UAVs matching attack targets under the weight network.

9. A terminal, characterized in that, Including: A memory, a processor, and a target matching program of the UAV cluster stored on the memory and executable on the processor. When the target matching program of the UAV cluster is executed by the processor, the steps of the target matching method of the UAV cluster according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the target matching method of the UAV cluster according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for controlling multi-agent formation

    CN110442129A

  • Task allocation method and related device

    CN111882152A

  • Unmanned aerial vehicle cluster multi-target search method and system based on deep reinforcement learning

    CN112947575A

  • Heterogeneous unmanned aerial vehicle formation autonomous decision-making implementation method based on cooperation of units and knowledge enhancement

    CN118838408A

  • Multi-unmanned aerial vehicle cooperative game autonomous decision-making method based on deep reinforcement learning

    CN120029053A