A method, system, device and medium for controlling flight of a drone swarm
By constructing a spatial association model between UAVs and detection mission targets and optimizing the DQN network model, the problem of poor mission performance of UAV swarms in dynamic environments was solved, and efficient and stable detection mission target allocation and flight attitude control were achieved.
Patent Information
- Application Number
- CN202510176214.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Traditional UAV swarm flight control methods struggle to respond quickly to dynamic environmental changes, resulting in poor mission performance and a lack of global coordination and intelligent allocation of detection mission targets.
By constructing a spatial association model between UAVs and detection mission targets, establishing a target swarm grayscale model, and optimizing the DQN network model, the system achieves integrated control of UAV swarm detection mission target allocation and flight attitude. It also quantifies the priority of detection mission targets by utilizing detection angle control error and historical detection counts, and constructs individual and swarm benefit functions for UAVs to provide a basis for real-time adjustments.
It improves the execution efficiency, global coordination and stability of UAV swarms in detection missions, enhances their adaptability to dynamic environments, and enables efficient target allocation and flight attitude control for UAV swarms in detection missions.
Smart Images

Figure CN120122719B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of unmanned aerial vehicle control, and particularly relates to a method, system, device and medium for controlling flight of a group of unmanned aerial vehicles. BACKGROUND
[0002] With the rapid development of intelligent technology, a group of unmanned aerial vehicles has gradually become an important tool for solving multi-detection task target detection tasks in a complex dynamic environment. Fixed-wing unmanned aerial vehicles have significant advantages in large-scale detection tasks due to their high speed, long endurance and low cost. However, due to the characteristics of weak hovering capability and insufficient flexibility of flight attitude adjustment, the intelligent requirements of fixed-wing unmanned aerial vehicles for detection task allocation and flight attitude control are higher.
[0003] Traditional control technology leads to a complex and lengthy overall task decision-making process, lacks global collaboration, and is difficult to quickly respond to dynamically changing task requirements, relies on predefined rules or detection task targets, and cannot flexibly adjust the detection task allocation scheme according to changes in the dynamic environment, thereby resulting in poor execution of the task by the group of unmanned aerial vehicles.
[0004] Therefore, how to provide a method for controlling flight of a group of unmanned aerial vehicles to ensure the effectiveness of the group of unmanned aerial vehicles in performing detection tasks has become a technical problem to be solved by those skilled in the art. SUMMARY
[0005] The present application provides a method, system, device and medium for controlling flight of a group of unmanned aerial vehicles to solve the technical problem of how to provide a method for controlling flight of a group of unmanned aerial vehicles to ensure the effectiveness of the group of unmanned aerial vehicles in performing detection tasks, improve the execution efficiency, global collaboration and stability of the group of unmanned aerial vehicles in detection tasks, and enhance the adaptability of the group of unmanned aerial vehicles to dynamic environments.
[0006] In a first aspect, the present application provides a method for controlling flight of a group of unmanned aerial vehicles, applied to a group of unmanned aerial vehicles composed of a plurality of unmanned aerial vehicles, the method comprising:
[0007] According to the established spatial correlation model between the unmanned aerial vehicle and the detection task target, an actual detection angle of the unmanned aerial vehicle in performing the detection task is obtained;
[0008] A target group gray degree model is constructed based on the historical number of times of being detected of the detection task target and a detection angle control error obtained from the actual detection angle, and the target group gray degree model is designed to quantitatively process the priority of the detection task target being detected;
[0009] constructing an individual expected benefit function of the UAV corresponding to the detection task target based on the target gray degree value output by the target group gray degree model, optimizing the individual expected benefit function according to preset information to obtain a UAV individual actual benefit function, and obtaining a UAV group benefit function corresponding to all detection task targets of the UAV group based on the UAV individual actual benefit function;
[0010] optimizing the constructed DQN network model according to each target gray degree value and the UAV group benefit function to obtain a target DQN network model;
[0011] In the actual UAV group flight control process, a flight control strategy output by the trained target DQN network model is executed, and the flight control strategy at least includes a detection task target allocation result of each UAV and detection angle data.
[0012] Preferably, the target group gray degree model is constructed based on the historical detection times of the detection task target and the detection angle control error obtained from the actual detection angle, and the target group gray degree model includes:
[0013] obtaining an expected detection angle of each UAV to the current detection task target according to a current flight control result of each UAV;
[0014] obtaining a detection angle control error according to the actual detection angle and the expected detection angle;
[0015] representing the historical detection times of the detection task target and the detection angle control error by using a nonlinear mapping relationship to obtain a detection task target group gray degree model;
[0016] The detection task target group gray degree model is represented as:
[0017] TG={Gr i |i=1,2,…,N}
[0018]
[0019] wherein g represents a nonlinear mapping relationship, Gr i represents a gray degree value of the ith detection task target, N represents a total number of detection task targets, C i represents a cumulative detection times of the ith detection task target, represents a detection angle control error of the ith detection task target detected by the jth UAV for the tth time.
[0020] Preferably, the target gray degree value output by the target group gray degree model is used to construct the individual expected benefit function of the UAV corresponding to the detection task target, and the individual expected benefit function includes:
[0021] calculating an individual maximum ideal benefit of the UAV according to the expected detection angle, the ideal detection angle of the UAV and the target gray scale value output by the target gray scale model;
[0022] adopting the detection angle control error and the individual maximum ideal benefit of the UAV to represent an individual expected benefit function of the UAV corresponding to the detection task target.
[0023] Preferably, the individual expected benefit function is optimized according to preset information, and a UAV group benefit function is obtained based on the optimization result of the individual expected benefit function of each UAV corresponding to each detection task target, comprising:
[0024] constructing a detection frequency decay function based on the influence of the cumulative detection frequency of the UAV on the actual detection effect;
[0025] calculating an importance coefficient of the detection task target according to the size of the detection task target;
[0026] adopting the detection frequency decay function and the importance coefficient to optimize the individual expected benefit function under the control error limitation of the UAV, to obtain an individual actual benefit function of the UAV.
[0027] Preferably, the UAV group benefit function is represented as:
[0028]
[0029] wherein M represents the total number of UAVs, A i represents the gray scale value of the i th detection task target, represents the expected detection angle of the i th detection task target when it is detected for the t th time by the j th UAV, represents the ideal detection angle of the i th detection task target, ε i represents the concealment coefficient of the i th detection task target, γ i represents the importance coefficient of the i th detection task target, μ represents a decay coefficient.
[0030] Preferably, the DQN network model is optimized according to each target gray scale value and the UAV group benefit function, to obtain a target DQN network model, comprising:
[0031] optimizing the state space of the DQN network model based on the target gray scale value, to obtain an initial target DQN network model;
[0032] The UAV group benefit function is taken as a reward function of the initial target DQN network model, and the initial target DQN network model is trained in a simulation environment for a UAV group multi-probe task target cooperative probe task, the parameters obtained by training are synchronized to a DQN network model carried by the UAV group, and a target DQN network model is obtained.
[0033] Preferably, the training of the initial target DQN network model in the simulation environment for the UAV group multi-probe task target cooperative probe task comprises:
[0034] The network training strategy is optimized by adopting a periodic repeated training strategy and a periodic parameter synchronization strategy, and the probe task target DQN network is trained in the simulation environment for the UAV group multi-probe task target cooperative probe task by using the optimized network training strategy.
[0035] In a second aspect, the application further provides a UAV group flight control applied to a UAV group composed of UAVs, and realizing the UAV group flight control method described above, the system comprises:
[0036] The system comprises a spatial correlation modeling unit, a target group gray degree model construction unit, a UAV group benefit function construction unit, a model optimization unit and a model application unit.
[0037] The spatial correlation modeling unit is configured to obtain an actual probe angle of the UAV when performing a probe task according to the established spatial correlation model between the UAV and the probe task target.
[0038] The target group gray degree model construction unit is configured to construct a target group gray degree model based on a historical number of times of being probed of the probe task target and a probe angle control error obtained from the actual probe angle, and the target group gray degree model is designed to quantitatively process a priority of the probe task target being probed.
[0039] The UAV group benefit function construction unit is configured to construct an individual expected benefit function of the UAV corresponding to the probe task target based on a target gray degree value output by the target group gray degree model, optimize the individual expected benefit function according to preset information to obtain a UAV individual actual benefit function, and obtain a UAV group benefit function of the UAV group corresponding to all probe task targets based on the UAV individual actual benefit function.
[0040] The model optimization unit is configured to optimize a constructed DQN network model according to each target gray degree value and the UAV group benefit function to obtain a target DQN network model.
[0041] The model application unit is configured to execute a flight control strategy output by the trained target DQN network model during actual flight control of the UAV group, the flight control strategy including at least a detection task target allocation result and detection angle data of each UAV.
[0042] In a third aspect, the present application provides a computer device, which comprises a memory, a processor and a transceiver connected through a bus, the memory is configured to store a set of computer program instructions and data, and transmit the stored data to the processor, the processor executes the program instructions stored in the memory to execute the UAV group flight control method described above.
[0043] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, when the computer program is executed, the UAV group flight control method described above is implemented.
[0044] The present application provides a UAV group flight control method, system, device and storage medium, compared with the prior art, the present application has the following beneficial effects:
[0045] The present application uses the historical detection times of the detection task target and the detection angle control error to represent the target gray value, quantifies the priority of the detection of the detection task target, uses the UAV group benefit function constructed by the target gray value, the detection angle control error, the target importance and the detection times decay function as the reward function of the DQN network model, provides real-time adjustment basis for the detection task target allocation and the UAV detection angle control, realizes the integrated control of the UAV group detection task target allocation and flight attitude, and improves the precision and efficiency of the UAV flight control. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 is a UAV group flight control method step schematic diagram provided by a preferred embodiment of the present application;
[0047] Figure 2 is a motion space correlation model of a UAV in the process of detecting a task target provided by a preferred embodiment of the present application;
[0048] Figure 3 is a training LOSS curve schematic diagram corresponding to the UAV group flight control method provided by a preferred embodiment of the present application;
[0049] Figure 4 is a training LOSS curve schematic diagram corresponding to a conventional DQN network model provided by a preferred embodiment of the present application;
[0050] Figure 5is a group reward curve diagram corresponding to the UAV group flight control method provided by a preferred embodiment of the present application;
[0051] Figure 6 is a group reward curve diagram corresponding to the conventional DQN network model provided by a preferred embodiment of the present application;
[0052] Figure 7 is a group reward comparison diagram of the UAV group flight control method provided by the present application and the optimal experience strategy and random strategy in the same scenario;
[0053] Figure 8 is a structure diagram of a UAV group flight control system provided by a preferred embodiment of the present application;
[0054] Figure 9 is a structure diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The embodiments of the present application will be described in detail below with reference to the drawings. The embodiments are presented only for the purpose of illustration and should not be understood as limiting the present application. The accompanying drawings are used for reference and illustration only and do not limit the scope of patent protection. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application. In the description of the present application, the terms "first", "second", "third" and the like are used only for the purpose of description and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second", "third" and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0056] In the description of the present application, it should be noted that unless otherwise explicitly defined and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication inside two elements. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are for illustrative purposes only, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. The term "and / or" used herein includes any and all combinations of one or more related listed items. The specific meanings of the above terms in the present application can be understood by the person skilled in the art.
[0057] In the description of the present application, it should be noted that unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as understood by those skilled in the art. The terms used in the specification of the present application are only for the purpose of describing the specific embodiments and are not intended to limit the present application. The specific meanings of the above terms in the present application can be understood by the person skilled in the art.
[0058] In military reconnaissance, the fixed-wing UAV group needs to efficiently allocate the task area, cooperatively detect multiple dynamic detection task target areas, and quickly obtain tactical information; in disaster relief, the fixed-wing UAV group needs to quickly cover the disaster area, identify and preferentially detect key areas; in resource exploration, the UAV group needs to optimize the detection path in complex terrain to achieve efficient coverage of multiple detection points.
[0059] Please refer to Figure 1 In an embodiment of the present application, a UAV group flight control method is provided, applied to a UAV group composed of individual UAVs, the method comprising:
[0060] S1, according to the established spatial correlation model between the UAV and the detection task target, the expected detection angle of the UAV in performing the detection task is obtained; the UAV movement involves dynamic state variables such as the position, speed, acceleration, camera line of sight direction and heading angle of the UAV, the dynamic state variables of the UAV movement and the detection task target position are jointly modeled to obtain the movement spatial correlation model of the UAV in the process of detecting the task target, such as Figure 2As shown, θ represents the current pitch angle of the aircraft, FOV (field of view) is the visual area that the observation camera can cover, LOS (line of sight) is the camera line of sight direction, generally the center line direction of the FOV, V U represents the speed vector of the unmanned aerial vehicle, a U represents the acceleration of the aircraft, used to adjust the flight attitude. Then the unmanned aerial vehicle position information in the determined coordinate system can be represented as:
[0061] p = [X, Y, Z] T
[0062] The speed vector of the unmanned aerial vehicle is represented as:
[0063] V U = [v x , v y , v z ] T
[0064] The acceleration vector of the unmanned aerial vehicle is represented as:
[0065] a U = [a x , a y , a z ] T
[0066] The position of the detection task target is represented as:
[0067] p = [x, y, z] T
[0068] Based on the Euclidean distance, the detection task target distance between the jth unmanned aerial vehicle and the detection task target is represented as:
[0069]
[0070] Generally, the camera line of sight direction is consistent with the nose direction in the horizontal plane, and in the longitudinal plane, the camera line of sight direction and the nose direction of the fixed-wing unmanned aerial vehicle have an angle β. According to the relationship between the camera line of sight direction and the heading angle, the rotation matrix representation R(β) between the camera line of sight direction and the heading angle is obtained, as shown below:
[0071]
[0072] The product of the rotation matrix representation and the speed vector is taken as the kinematic vector representation v LOS of the camera line of sight direction, as shown below:
[0073] v LOS = R(β)v U
[0074] The angle between the velocity vector of the camera line of sight direction and the position vector of the detection task target is a detection angle a, and a detection condition constraint equation is constructed as follows:
[0075]
[0076] When The detection task target can be detected.
[0077] The kinematic vector representation and the detection task target distance representation are substituted into the detection condition constraint equation to obtain a UAV detection angle equation, and the UAV detection angle equation is as follows:
[0078]
[0079] Wherein, a represents the detection angle, R(β) represents the rotation matrix, v U represents the velocity vector, and d represents the Euclidean distance.
[0080] According to the UAV detection angle equation, an actual detection angle representation equation of the UAV when performing the detection task is obtained.
[0081] S2, a target group gray degree model is constructed based on the historical detection times of the detection task target and the detection angle control error obtained from the actual detection angle, and the target group gray degree model is designed to quantitatively process the priority of the detection of the detection task target; in the preferred embodiment of the application, the historical detection times of the detection task target and the detection angle control error obtained from the actual detection angle are represented by a nonlinear mapping relationship to obtain the target group gray degree model, the historical detection times of each detection task target and the detection angle control error of each detection are input into the target group gray degree model, and the target gray degree value corresponding to each detection task target is output, the target gray degree value is used to quantitatively process the priority of the next detection of each detection task target in the detection task target group, specifically, when the gray degree value of the detection task target 1 decreases by 20, and the gray degree of the detection task target 2 decreases by 50, the importance of the detection task target 2 at the current time is higher than that of the detection task target 1. In a dynamic motion state of the UAV group, different previous detection results will lead to changes in subsequent decision-making, and the target gray degree value of the detection task target reflects the dynamic change of the detection task target.
[0082] In the preferred embodiment of the application, the expected detection angle of each UAV to the current detection task target is obtained according to the current flight control result of each UAV, and the expected detection angle is the UAV detection angle result obtained by the UAV group according to the UAV group flight control method provided by the application. Further, the absolute value of the difference between the actual detection angle and the expected detection angle is calculated and taken as the detection angle control error, and the detection angle control error is represented as follows:
[0083]
[0084] wherein, denotes the detection angle control error when the ith target is detected for the tth time by the jth UAV, denotes the expected detection angle when the ith detection task target is detected for the tth time by the jth UAV, denotes the actual detection angle when the ith detection task target is detected for the tth time by the jth UAV.
[0085] Finally, the historical detection times and the detection angle control error of the detection task target are represented by a nonlinear mapping relationship to obtain a target group grey degree model, and the target group grey degree model is specifically represented as:
[0086] TG={Gr i |i=1,2,…,N}
[0087]
[0088] wherein, g represents a nonlinear mapping relationship, Gr i denotes the grey degree value of the ith detection task target, N represents the total number of detection task targets, C i denotes the cumulative detection times of the ith detection task target.
[0089] Based on the target group grey degree model, the target grey degree values of each detection task target in the UAV group detection process can be obtained, which is fed back to the UAV group as a reward, and provides a real-time adjustment basis for the allocation and detection angle control of the detection task target, so as to ensure that the UAV group can more effectively identify the key detection task target in the current task, so as to prioritize the allocation of UAV resources for the detection of the key detection task target, and improve the detection efficiency of the detection task target.
[0090] S3, based on the target grey degree value output by the target group grey degree model, an individual expected return function corresponding to the UAV of the detection task target is constructed, the individual expected return function is optimized according to preset information, an individual actual return function of the UAV is obtained, and a UAV group return function corresponding to each detection task target of the UAV group is obtained based on the individual actual return function of the UAV; in the preferred embodiment of the present application, the UAV group return function is used as a reward function, the UAV group return is the sum of the individual actual return of the UAV, and the individual actual return of the UAV is related to the detection angle control error.
[0091] In the preferred embodiment of the present application, the individual expected return function is represented based on the detection angle control error, and is specifically:
[0092] The individual expected return of the UAV is represented as:
[0093]
[0094] wherein, represents the expected return of the jth UAV at the tth detection of the ith detection task target, R max represents the maximum ideal return of the UAV individual, k represents the error penalty coefficient.
[0095] For the maximum ideal return of the UAV individual R max , based on the gray value of the detection task target, it is represented as follows:
[0096]
[0097] wherein, A i represents the gray value of the ith detection task target, represents the ideal detection angle of the ith detection task target. Wherein, the ideal detection angle is the best return detection angle corresponding to the current detection task target type, and is a certain value, which is determined by priori.
[0098] Based on the detection angle control error and the maximum ideal return of the UAV individual, the expected return of the UAV individual is obtained, and the expected return of the UAV individual is represented as:
[0099]
[0100] Further, the concealment coefficient of the detection task target is introduced to the expected return of the UAV individual, the concealment coefficient of the detection task target represents the ability of the detection task target to avoid detection of the aircraft, which is determined according to the type of the detection task target. Since the total number of detection task targets in the region and the type of detection task target are known, the concealment coefficient of the detection task target is determined by priori.
[0101] Then, after introducing the concealment coefficient of the detection task target, the individual expected return function of the UAV is represented as:
[0102]
[0103] wherein, ε i represents the concealment coefficient of the ith detection task target.
[0104] Further, the individual expected benefit function of the UAV is introduced with a detection task target importance coefficient to adjust the expected benefit of each detection task target, the detection task target importance coefficient is related to the size of the detection task target, according to the hierarchical detection idea, the space-based device can provide preliminary intelligence information in the global range before the UAV enters the specified local area to perform detailed detection tasks, including the approximate size of each detection task target in the detection task target group and the boundary of the local area, therefore, the detection task target importance coefficient is defined as:
[0105]
[0106] Wherein, S i is the volume of the detection task target i.
[0107] Since multiple detections of the same detection task target will have a certain degree of information repetition, the influence of the cumulative detection times on the actual detection effect needs to be considered, that is, the ideal detection benefit of the detection task target i will gradually decrease with the increase of the cumulative detection times, therefore, a detection times decay function is established, which is expressed as:
[0108]
[0109] Wherein, μ represents the decay coefficient.
[0110] Under the control error limitation of the UAV, the individual expected benefit function is optimized by using the detection times decay function and the importance coefficient, to obtain the individual actual benefit function of the UAV, which is expressed as:
[0111]
[0112] Wherein, represents the actual benefit of the jth UAV to the detection task target i.
[0113] For a scene with N UAVs and M detection task targets, based on the individual actual benefit function of the UAV, the UAV group benefit function corresponding to each detection task target of the UAV group is obtained, which is expressed as:
[0114]
[0115] In the preferred embodiment of the present application, the UAV group benefit function fully considers the importance of the detection task target, the detection angle control error of the UAV, the increase of the detection times, and the influence of the target gray value of each detection task target on the detection benefit, which is fed back as a reward to the UAV, to provide real-time adjustment basis for the allocation of the detection task target and the detection angle control of the UAV, so as to improve the precision and efficiency of the flight control of the UAV.
[0116] S4, according to the target gray degree value and the UAV group benefit function, the constructed DQN network model is optimized to obtain a target DQN network model; in the preferred embodiment of the present application, the DQN network model is used as the model of the UAV flight control, the target gray degree value is introduced into the constructed state space to optimize the state space of the DQN network model to obtain an initial target DQN network model, and specifically:
[0117] The state space and the action space of the DQN network model are constructed, based on the control law of the longitudinal channel control coefficient of the UAV, the detection angle control error of the UAV at the observation point position is related to the distance from the UAV to the detection task target, the current pitch angle, the flight speed and the flight height, and the cumulative detection times affect the actual detection benefit of the current individual. In the preferred embodiment of the present application, based on the distance from the UAV to the detection task target, the current pitch angle, the flight speed, the flight height, the cumulative historical detection times of the detection task target and the target gray degree value, the state space of the DQN algorithm is constructed, as shown in Table 1 for the state space of the DQN network model:
[0118] Table 1
[0119] State variable name Symbol Dimension Number N 1 Distance D 1 Pitch angle θ 1 Speed v 1 Height h 1 Historical number of times of being detected by No.1 detection task target 1 Historical number of times of being detected by No.2 detection task target <k2> 1 Historical number of times of being detected by No.i detection task target k i ]]> 1 Gray value of detection task target Gr i ]] 1
[0120] Wherein, the number represents the position or stage of the current UAV in the task sequence.
[0121] For the action space of the target DQN network model, the UAV detection angle a is converted into the UAV tilt angle for decision-making because the camera installation angle is a fixed value. Generally, after trajectory planning, the detection task target is below the UAV, at this time, the detection task target can be classified as a low-height detection task target and a high-height detection task target according to the height characteristics of the detection task target, that is, the basic action corresponding to the detection task target selection result has two, each type of detection task target in the detection angle basic action has 26, and there are 52 basic actions, as shown in Table 2:
[0122] Table 2
[0123] Type of detection task target Optional tilt angle (°) High-altitude detection task target [20,45] Low-altitude detection task target [45,70]
[0124] Further, the UAV group benefit function is taken as the reward function of the initial target DQN network model, and the initial target DQN network is trained for a UAV group multi-probe task target cooperative probe task in a simulation environment, parameters obtained through training are synchronized to the DQN network model carried by the UAV group, and a target DQN network model is obtained; the UAV group benefit function fuses the change of the gray degree of the probe task target group, the probe frequency attenuation, the probe angle control error of the UAV, and the importance of the probe task target, takes them as the reward function of the initial target DQN network model, can respond to the change of the probe task target and the environment in real time, and the gray degree value of the probe task target group is used to balance the priority of the probe task target to be probed next time and the effectiveness of the probe angle adjustment, to adjust the priority and the allocation weight of the probe task target in real time, to ensure that the UAV group can quickly respond to the demand for the change of the number, position and characteristics of the probe task target in the dynamic environment, always probes the most important probe task target first, and makes the UAV group maximize the probe efficiency and the task benefit under limited resources.
[0125] The original training strategy of the DQN network model is that when the data storage quantity in the experience pool is greater than the minimum batch (the sample quantity input into the model in each iteration) quantity size required in the training process, a data set with a batch size (the sample quantity input into the model in each iteration) is extracted from the experience pool in a random manner to train the DQN network model, and the parameters obtained by training are synchronized to the DQN network model carried by the UAV group to update the parameters and obtain the target DQN network model. However, the above training strategy is low in efficiency, especially in a high-dimensional state space and action space, and is prone to unstable convergence or falling into a local optimum. In order to improve the convergence speed and stability of the DQN network model after convergence, the training strategy of the DQN network model is improved. The improvement direction includes two aspects, which are a periodic repeated training strategy and a periodic parameter synchronization strategy. The sample extraction method originally used by the DQN network model can cause the diversity of the randomly extracted sample data set to be poor. In addition, as the data storage quantity in the experience pool increases with the increase of the training round, frequent extraction causes the collected results to contain a large amount of repeated data, thereby reducing the effectiveness of training. In the preferred embodiment of the present application, the periodic repeated training strategy is adopted, that is, m times of training are performed every n rounds, so as to increase the proportion of new data in the training data, improve the utilization rate of the sample data, realize more effective exploration of the UAV, and improve the efficiency of model training. The parameter synchronization strategy originally used by the DQN network model is that the training parameter synchronization is performed once after each training. Frequent parameter synchronization causes the Q value of the target DQN network model used by the UAV-mounted DQN network model to fluctuate, affects the stability of convergence, and causes the training network to have no stable Q value reference in a certain period of time, thereby slowing down the convergence speed. In the preferred embodiment of the present application, the periodic parameter synchronization strategy is adopted, that is, the training parameter synchronization is performed once every fixed training time, thereby reducing the training fluctuation and accelerating the convergence speed of the network in the training process.
[0126] S5, in the actual UAV group flight control process, the flight control strategy output by the trained target DQN network model is executed, and the flight control strategy at least includes the detection target allocation result and the detection angle data of each UAV; in the actual application process, the state space data and the action space data of the UAV group are collected in real time, the real-time state space data and the real-time action space data are input into the target DQN network model carried by the UAV, and the detection task target allocation result and the UAV detection angle data of each UAV are obtained. The present application integrates the detection task target allocation and the flight attitude control into a joint decision problem, thereby improving the task execution efficiency and the cooperation.
[0127] In a preferred embodiment of the present application, the number of members in the fixed-wing UAV group in the scene is set to 10, the number of members in the detection task target group is set to 3, the flight performance of the UAVs is similar, the camera installation angles are different, the concealment of the members in the detection task target group is different, and the members belong to different categories. The UAV group related parameter setting rules are shown in Table 3.
[0128] Table 3
[0129] State variable name Value range Value mode Number [1,10] Sequential generation Longitude [142.9,143] Random Latitude [37.9,38] Random Pitch angle [-15,+15](°) Random Camera mounting angle [5,15](°) Random Speed [130, 200] (km / h) Random Height [3000,4000](m) Random
[0130] The detection task target group related parameter setting is shown in Table 4.
[0131]
[0132] The detection task target importance coefficient is (0.6, 0.3, 0.1), the detection task target concealment coefficient is (0.2, 0.2, 0.2), and the initial gray degree of the detection task target group is 500. The network hyperparameters used by the DQN algorithm are shown in Table 5.
[0133] Table 5
[0134] Parameter name Value Number of layers 3 batch_size 256 buffer_size 2500 max_epsilon 1 min_epsilon 0.01 epslon_decay 0.97 gamma 0.9 Training interval 5 Number of repeated training 10 Network parameter synchronization interval 20
[0135] The UAV group flight control method provided in the present application is compared with the conventional DQN network model, and the results are shown in Figure 3 and Figure 4 The training LOSS curves corresponding to the UAV group flight control method of the present application and the conventional DQN network model are shown in Figure 5 and 6, respectively. As shown in Figure 3 and Figure 4 Compared with the conventional DQN network model, the UAV group flight control method of the present application obviously accelerates the convergence speed of the LOSS curve, and the curve is more stable after convergence, proving that the UAV group flight control method of the present application maintains a stable Q value reference for a long time during the training process, accelerates the convergence speed, and the whole process is smoother. As shown in Figure 5 and Figure 6 The starting promotion point of the UAV group flight control method of the present application is 10 times earlier, the stable point is 57.1% earlier, and the reward mean value in the stable period is improved by 6.1%.
[0136] The UAV group flight control method provided in the present application is compared with the optimal experience strategy and the random strategy under the same scene, and the effectiveness of the algorithm is verified by setting the test scene. During the comparison process, the detection task target group parameters are unchanged, and the UAV group parameters are set as shown in Table 6.
[0137] Table 6
[0138] State variable name Value range Value mode Number [1,8] Sequential generation Longitude [137.8,138] Random Latitude [33.7,34] Random Pitch angle [-10,+10](°) Random Camera mounting angle [5,10](°) Random Speed [130, 210] (km / h) Random Height [2800,4200](m) Random
[0139] Based on the scene parameter setting, the test is repeated 200 times, and the UAV group flight control method of the application is compared with the optimal experience strategy and random strategy under the same scene, as shown in Figure 7 Fig. 6 shows the group reward comparison chart of the UAV group flight control method provided by the application and the optimal experience strategy and random strategy under the same scene.
[0140] In the 200 tests with different numbers of UAVs, the decision result of the UAV group flight control method provided by the application can make the gray degree drop value of the detection task target reach a high level, and has good stability. The experience strategy is affected by the UAV control law and the UAV flight performance, and cannot make the gray degree drop value of the detection task target reach a high level. The experience strategy has poor adaptability to different parameter situations in the scene, has large fluctuation range, and has poor stability. The statistical results are shown in Table 7.
[0141] Table 7
[0142] Mean reward Reward standard deviation Improved DQN method 209.66 5.31 Experience optimal strategy 131.38 34.29 Random strategy 34.74 158.13
[0143] The mean value of the UAV group flight control method provided by the application is increased by about 59.58% compared with the experience strategy method in the test of different numbers of UAVs, the standard deviation is decreased by about 7 times, and the efficiency is increased by about one order of magnitude compared with the random method, which proves the effectiveness and reliability of the UAV group flight control method of the application in practical application.
[0144] For the above scenarios, the application provides a UAV group flight control method, according to the established spatial correlation model between the UAV and the detection task target, the actual detection angle of the UAV when performing the detection task is obtained; based on the historical detection times of the detection task target and the detection angle control error obtained by the actual detection angle, a target group gray degree model is constructed, the target group gray degree model is designed to quantitatively process the priority of the detection of the detection task target; based on the target gray degree value output by the target group gray degree model, an individual expected income function of the UAV corresponding to the detection task target is constructed, the individual expected income function is optimized according to the preset information, and the individual actual income function of the UAV is obtained, and based on the individual actual income function of the UAV, the UAV group income function corresponding to all detection task targets of the UAV group is obtained; according to each target gray degree value and the UAV group income function, the constructed DQN network model is optimized to obtain a target DQN network model; in the actual UAV group flight control process, the flight control strategy output by the trained target DQN network model is executed, and the flight control strategy at least includes the detection task target allocation result of each UAV and the detection angle data. The UAV group flight control method provided by the application characterizes the target gray degree value with the historical detection times of the detection task target and the detection angle control error, quantifies the priority of the detection of the detection task target, constructs the UAV group income function with the target gray degree value, the detection angle control error, the target importance and the detection times decay function as the reward function of the DQN network model, provides real-time adjustment basis for the detection task target allocation and the UAV detection angle control, realizes the integrated control of the UAV group detection task target allocation and the flight attitude, and improves the precision and efficiency of the UAV flight control.
[0145] Correspondingly, as Figure 8 shown, based on a UAV group flight control method, the application embodiment further provides a UAV group flight control system applied to a carrying control system at least composed of a base body, a mechanical arm, a multi-point force sensor, an acceleration sensor and a micro vacuum degree sensor, realizes the UAV group flight control method disclosed by the application embodiment, and includes a spatial correlation modeling unit 1, a target group gray degree model construction unit 2, a UAV group income function construction unit 3, a model optimization unit 4 and a model application unit 5.
[0146] The spatial correlation modeling unit 1 is used for obtaining the actual detection angle of the UAV when performing the detection task according to the established spatial correlation model between the UAV and the detection task target;
[0147] The target group gray degree model construction unit 2 is configured to construct a target group gray degree model based on the historical detection times of the detection task targets and the detection angle control errors obtained from the actual detection angles, and the target group gray degree model is designed to quantitatively process the priorities of the detection task targets.
[0148] The UAV group benefit function construction unit 3 is configured to construct an individual expected benefit function of the UAV corresponding to the detection task target based on the target gray degree values output by the target group gray degree model, to optimize the individual expected benefit function according to preset information to obtain an individual actual benefit function of the UAV, and to obtain a UAV group benefit function of the UAV group corresponding to all detection task targets based on the individual actual benefit function of the UAV.
[0149] The model optimization unit 4 is configured to optimize the constructed DQN network model according to the target gray degree values and the UAV group benefit function to obtain a target DQN network model.
[0150] The model application unit 5 is configured to execute a flight control strategy output by the trained target DQN network model in an actual UAV group flight control process, and the flight control strategy at least includes a detection task target allocation result and detection angle data of each UAV.
[0151] The specific limitations of the UAV group flight control system can refer to the limitations of the UAV group flight control method described above, which will not be repeated here. Those skilled in the art can realize that the various modules and steps described in combination with the embodiments disclosed in the present application can be realized in hardware, software or both. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0152] As shown in Figure 9 The computer device provided by the embodiment of the present application includes a processor, a memory and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps in the UAV group flight control method embodiment described above are implemented, such as steps S1-S5 described in Figure 1 .
[0153] Those skilled in the art can understand that the schematic Figure 9The computer device is only an example and does not constitute a limitation on the computer device, which can include more or fewer components than shown, or combine some components, or have different components, for example, the computer device can also include an input / output device, a network access device, a bus, etc.
[0154] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor is the control center of the computer device, and connects all parts of the computer device through various interfaces and lines.
[0155] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required for a function (such as a sound playing function, an image playing function, etc.), etc.; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory device.
[0156] If the modules integrated in the computer device are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0157] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned various method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0158] Correspondingly, the embodiment of the present application provides a computer readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer readable storage medium is located to perform the steps in the UAV group flight control method of the above-mentioned embodiment, for example, the steps S1-S5 in the method described in the above-mentioned embodiment. Figure 1
[0159] The unmanned aerial vehicle group flight control method, system, device and medium provided in the embodiment are used to solve the technical problem of providing an unmanned aerial vehicle group flight control method and guaranteeing the effect of the unmanned aerial vehicle group when performing a detection task. The unmanned aerial vehicle group flight control method comprises the following steps: obtaining an actual detection angle of an unmanned aerial vehicle when performing a detection task according to a space correlation model established between the unmanned aerial vehicle and a detection task target; constructing a target group gray degree model based on a historical detection frequency of the detection task target and a detection angle control error obtained from the actual detection angle, the target group gray degree model being designed to quantitatively process a priority of the detection task target being detected; constructing an individual expected benefit function of the unmanned aerial vehicle corresponding to the detection task target based on a target gray degree value output by the target group gray degree model, optimizing the individual expected benefit function according to preset information, obtaining an individual actual benefit function of the unmanned aerial vehicle, and obtaining a group benefit function of the unmanned aerial vehicle group corresponding to all detection task targets based on the individual actual benefit function of the unmanned aerial vehicle; optimizing a constructed DQN network model according to each target gray degree value and the group benefit function of the unmanned aerial vehicle, and obtaining a target DQN network model; and in an actual unmanned aerial vehicle group flight control process, executing a flight control strategy output by the trained target DQN network model, the flight control strategy at least comprising a detection task target allocation result of each unmanned aerial vehicle and detection angle data. The unmanned aerial vehicle group flight control method provided in the application uses the historical detection frequency of the detection task target and the detection angle control error to represent the target gray degree value, quantifies the priority of the detection task target being detected, uses the group benefit function of the unmanned aerial vehicle group constructed based on the target gray degree value, the detection angle control error, the target importance and a detection frequency decay function as a reward function of the DQN network model, provides a real-time adjustment basis for the detection task target allocation and the unmanned aerial vehicle detection angle control, realizes integrated control of the unmanned aerial vehicle group detection task target allocation and flight attitude, and improves the precision and efficiency of the unmanned aerial vehicle flight control.
[0160] Each embodiment in the specification is described in a progressive manner, and the same or similar parts of each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the system embodiment is basically similar to the method embodiment, so the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment. It should be noted that each technical feature of the above embodiments can be combined arbitrarily, and in order to make the description simple, not all possible combinations of the technical features of the above embodiments are described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the specification.
[0161] The above described embodiments only express several preferred embodiments of the present application, which are described in more detail and in more specifically, but can not be understood as the limitation of the scope of the present application. It should be noted that for ordinary skilled in the art, several improvements and replacements can be made without departing from the technical principles of the present application, and these improvements and replacements should also be considered as the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for controlling the flight of an unmanned aerial vehicle (UAV) swarm, characterized in that, The method, applied to a swarm of drones consisting of individual drones, includes: Based on the established spatial correlation model between the UAV and the target of the detection mission, the actual detection angle of the UAV when performing the detection mission is obtained; A target group grayscale model is constructed based on the historical number of times the target of the detection mission is detected and the detection angle control error obtained from the actual detection angle. The target group grayscale model is designed to quantify the priority of the detection of the target of the detection mission. Based on the target gray value output by the target group gray model, an individual expected revenue function of the UAV corresponding to the detection mission target is constructed. The individual expected revenue function is optimized according to preset information to obtain the actual revenue function of the individual UAV. Based on the actual revenue function of the individual UAV, the UAV group revenue function corresponding to all detection mission targets is obtained. Based on the target gray values and the drone swarm benefit function, the constructed DQN network model is optimized to obtain the target DQN network model. In the actual flight control process of the UAV swarm, the flight control strategy output by the trained target DQN network model is executed. The flight control strategy includes at least the detection task target allocation result and detection angle data of each UAV.
2. The unmanned aerial vehicle swarm flight control method as described in claim 1, characterized in that, The construction of a target group greyscale model based on the historical number of times the target has been detected and the detection angle control error obtained from the actual detection angle includes: Based on the current flight control results of each UAV, the expected detection angle of each UAV for the current detection target is obtained; The detection angle control error is obtained based on the actual detection angle and the desired detection angle. A gray-scale model of the target group of the detection mission is obtained by using a nonlinear mapping relationship to represent the historical number of times the target of the detection mission has been detected and the detection angle control error. The gray-scale model of the target group of the detection mission is represented as follows: TG={Gr i |i=1,2,…,N} Where g represents a nonlinear mapping relationship, Gr i C represents the grayscale value of the i-th detection target, N represents the total number of detection targets, and C represents the grayscale value of the i-th detection target. i This represents the cumulative number of times the i-th detection target has been detected. This represents the detection angle control error when the i-th detection target is detected by the j-th UAV for the t-th time.
3. The unmanned aerial vehicle swarm flight control method as described in claim 2, characterized in that, The step of constructing an individual expected return function for the UAV corresponding to the detection mission target based on the target grayscale value output by the target group grayscale model includes: Based on the expected detection angle, the ideal detection angle of the UAV, and the target gray value output by the target group gray model, calculate the maximum ideal benefit of the individual UAV. The detection angle control error and the maximum ideal individual benefit of the UAV are used to represent the individual expected benefit function of the UAV corresponding to the detection mission target.
4. The unmanned aerial vehicle swarm flight control method as described in claim 1, characterized in that, The optimization of the individual expected return function based on preset information, and the obtaining of the UAV swarm return function based on the optimization results of the individual expected return function for each UAV corresponding to each detection mission target, includes: Based on the impact of the cumulative number of times a UAV is detected on the actual detection effect, a detection count decay function is constructed. Calculate the importance coefficient of the target based on the size of the target in the detection mission; Under the control rate error limit of the UAV, the individual expected revenue function is optimized by using the detection count decay function and the importance coefficient to obtain the actual revenue function of the UAV individual.
5. The unmanned aerial vehicle swarm flight control method as described in claim 2, characterized in that, The swarm payoff function for the drones is expressed as follows: Where M represents the total number of drones, A i This represents the grayscale value of the i-th detection target. This represents the expected detection angle when the i-th detection target is detected by the j-th UAV for the t-th time. Let ε represent the ideal detection angle for the i-th detection target. i γ represents the stealth coefficient of the i-th detection target. i Let μ represent the importance coefficient of the i-th detection target, μ represent the attenuation coefficient, k represent the error penalty coefficient, and σ represent the reward attenuation rate caused by the detection angle control error. It represents the actual detection angle when the i-th detection target is detected by the j-th UAV for the t-th time.
6. The unmanned aerial vehicle swarm flight control method as described in claim 1, characterized in that, The optimization of the constructed DQN network model based on each of the target grayscale values and the drone swarm benefit function to obtain the target DQN network model includes: Based on the target gray value, the state space of the DQN network model is optimized to obtain the initial target DQN network model; The reward function of the UAV swarm is used as the reward function of the initial target DQN network model. The initial target DQN network model is trained in a simulation environment using a multi-detection target collaborative detection task of the UAV swarm. The trained parameters are synchronized to the DQN network model carried by the UAV swarm to obtain the target DQN network model.
7. The unmanned aerial vehicle swarm flight control method as described in claim 6, characterized in that, The training of the initial target DQN network model in a simulation environment for a multi-target cooperative detection task involving a UAV swarm includes: The network training strategy is optimized by employing a periodic repetition training strategy and a periodic parameter synchronization strategy. The optimized network training strategy is then used in a simulation environment to train the DQN network for collaborative detection of multiple detection targets by UAV swarms.
8. A flight control system for a swarm of unmanned aerial vehicles (UAVs), characterized in that, The system, which is applied to a swarm of drones composed of individual drones, includes: a spatial association modeling unit, a target swarm grayscale model construction unit, a drone swarm benefit function construction unit, a model optimization unit, and a model application unit; The spatial association modeling unit is used to obtain the actual detection angle of the UAV when performing the detection mission based on the established spatial association model between the UAV and the detection mission target; The target group grayscale model construction unit is used to construct a target group grayscale model based on the historical number of times the target of the detection task has been detected and the detection angle control error obtained from the actual detection angle. The target group grayscale model is designed to quantify the priority of the detection of the target of the detection task. The drone swarm benefit function construction unit is used to construct the individual expected benefit function of the drone corresponding to the detection mission target based on the target gray value output by the target swarm gray value model, optimize the individual expected benefit function according to preset information to obtain the actual benefit function of the drone individual, and obtain the drone swarm benefit function corresponding to all detection mission targets based on the actual benefit function of the drone individual. The model optimization unit is used to optimize the constructed DQN network model according to each of the target gray values and the UAV swarm benefit function to obtain the target DQN network model. The model application unit is used to execute the flight control strategy output by the trained target DQN network model during the actual UAV swarm flight control process. The flight control strategy includes at least the detection task target allocation result and detection angle data for each UAV.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a transceiver, which are connected to each other via a bus; the memory is used to store a set of computer program instructions and data, and to transmit the stored data to the processor, and the processor executes the program instructions stored in the memory to perform the unmanned aerial vehicle swarm flight control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that, when executed, implements the unmanned aerial vehicle swarm flight control method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Unmanned aerial vehicle cluster formation safe flight control system based on block chain technology
CN117472076A
Multi-unmanned-system collaborative decision-making method, system and equipment based on Soar architecture and medium
CN118884829A