Heterogeneous unmanned aerial vehicle area coverage method, system and device

The heterogeneous drone coverage network model is trained through the ACTOR-CRITIC reinforcement learning method, which solves the problem that multi-UAV systems fail to effectively consider the performance differences and dynamic environment changes of heterogeneous drone during coverage control, and realizes the autonomous calculation of coverage paths and real-time control of the drone, improving coverage efficiency and accuracy.

CN120122694AActive Publication Date: 2025-06-10HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202510284348.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing multi-UAV system fails to effectively consider the performance differences and dynamic environmental changes of heterogeneous drones during coverage control, resulting in low coverage efficiency and accuracy, and the central control method is not time-sensitive and scalable.

Method used

ACTOR-CRITIC reinforcement learning method is adopted to train heterogeneous drone coverage network models through historical interactive data between multiple drones and target areas, use attention network mechanism to transmit messages, extract feature values ​​and update state value functions and policy network parameters, and realize the unmanned calculation of coverage paths independently.

Benefits of technology

The optimal coverage path planning and real-time control of heterogeneous drones in the target area are realized, avoiding the dependence of central control nodes, improving coverage efficiency and accuracy, and rapid deployment and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122694A_ABST
    Figure CN120122694A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous unmanned aerial vehicle area coverage method, system and device, and relates to the technical field of multi-agent system planning and control, and the method comprises the steps: collecting the interaction data of each unmanned aerial vehicle and the environment; a multi-agent reinforcement learning algorithm is trained based on interaction data of each unmanned aerial vehicle and the environment, and a heterogeneous unmanned aerial vehicle coverage network model is constructed; inputting the position information of each unmanned aerial vehicle into a heterogeneous unmanned aerial vehicle coverage network model, and outputting an expected attitude of each unmanned aerial vehicle; according to the method, the coverage path planning of the heterogeneous unmanned aerial vehicle is completed according to the expected attitude of each unmanned aerial vehicle, and in the method, each unmanned aerial vehicle can autonomously calculate the current optimal coverage path planning based on the current observation of the unmanned aerial vehicle to the environment and the information interaction with the neighbor unmanned aerial vehicle, so that the heterogeneous unmanned aerial vehicle can obtain the optimal control strategy; and real-time coverage planning and control are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-agent system planning and control, and particularly relates to a heterogeneous UAV area coverage method, system and device. Background Art

[0002] Full coverage and monitoring of a target area is one of the important tasks in the application of multi-UAV systems. When UAVs perform coverage tasks, it is specifically realized as multi-UAV joint vision coverage of a given area, and each UAV sends the collected information to a ground central controller. The ground controller can perform behaviors such as local map construction or path planning based on this information.

[0003] In a multi-UAV system, due to the communication burden between UAVs and the computing burden of the central processor, an increase in the number of UAVs does not necessarily improve the system performance. To ensure that the multi-UAV system can fully collect information about the target area while reducing the system burden, it is necessary to precisely online plan and control the movement and behavior of the UAV swarm to achieve optimal coverage of the target area.

[0004] Traditional UAV coverage control algorithms mainly use the Thiessen polygon method for area division. This method assumes that all UAVs have the same dynamic characteristics and task capabilities, thus simplifying UAVs into homogeneous objects for optimizing the coverage area. In practical applications, this assumption ignores the differences in performance and function among UAVs, such as different flight speeds, payload capacities, and endurance times, resulting in a significant reduction in coverage efficiency and accuracy. Especially when facing a dynamically changing environment or task requirements, the flexibility and adaptability of homogeneous algorithms are extremely low, and they cannot effectively adjust the specific roles and responsibility assignments of each UAV.

[0005] In recent years, coverage control strategies for multi-UAV systems have been proposed based on neural networks, achieving good results. However, such methods have very complex neural networks and require an excessively long training time. In addition, most coverage control methods only consider the planning level, that is, it is assumed that each UAV can accurately execute the desired actions, which does not conform to the actual situation. Therefore, the current coverage control methods for multi-UAV systems still have the following problems: (1) Most coverage methods are a centralized control method, which requires a central control node to perform real-time control on each UAV, lacking timeliness and scalability; (2) The heterogeneity of the multi-UAV system is not considered, that is, the structures and detection capabilities of each UAV in the multi-UAV system are different; this will reduce the coverage efficiency of the multi-UAV system; (3) The multi-UAV coverage control method based on neural networks has a complex network structure and a long training time, lacking the ability for rapid deployment; (4) The dynamic constraints of actual UAVs are not considered, that is, it cannot be guaranteed that UAVs will definitely execute the desired actions correctly. In addition, traditional coverage algorithms consider UAVs as homogeneous objects, while in actual situations, UAVs have different dynamic models and tasks in cluster missions, so it is difficult to transfer them to actual scenarios. Summary of the Invention

[0006] Aiming at the deficiency that the heterogeneity of the multi-UAV system is not considered when covering and controlling an area in the prior art, UAVs are regarded as homogeneous objects, and it is difficult to transfer them to actual scenarios. The present invention proposes a heterogeneous UAV area coverage method, system and device, which uses the method of reinforcement learning to complete the coverage path planning of heterogeneous UAVs, realizes real-time coverage planning and control, and thus solves the problems existing in the prior art.

[0007] A heterogeneous UAV area coverage method includes the following steps:

[0008] Collect historical interaction data of multiple UAVs and the target area; the interaction data includes the current observation data of each UAV itself and the observation data of neighboring UAVs;

[0009] Train the ACTOR-CRITIC reinforcement learning method using the historical interaction data of multiple drones with the target area to construct a heterogeneous drone coverage network model. Specifically, it includes: based on the current observation data of each drone itself and the observation data of neighboring drones, obtain the reward feedback of the target area, the current observation value of the drone, and the new observation value obtained after executing an action, and form the new observation value into a state tuple and save it to the experience replay pool of the state value function network CRITIC. Select any state tuple, and use the message passing mechanism based on the attention network to extract the eigenvalue of each drone when performing multiple rounds of message passing with neighboring drones. Generate the state value function network CRITIC for each drone. Use the stochastic gradient descent method to update the parameters of the state value function network CRITIC and the parameters of the policy network ACTOR to construct a heterogeneous drone coverage network model;

[0010] Input the real-time interaction data of multiple drones with the target area into the heterogeneous drone coverage network model, and output the expected attitude of each drone. Each drone realizes the coverage of the target area by tracking its own expected attitude.

[0011] Further, the steps of obtaining the reward feedback of the target area, the current observation value of the drone, and the new observation value obtained after executing an action, and forming the new observation value into a state tuple and saving it to the experience replay pool of the state value function network CRITIC are specifically as follows:

[0012] Define the set composed of the m vertex coordinates of the target area as [e 1 ,..., e m , the coordinate of drone i is p i , and the coordinate set of the N i drones closest to drone i is Each drone i obtains the observation value t of the current environment s as:

[0013]

[0014] According to the current policy network π θ execute an action where θ is the parameter of the policy network;

[0015] Each drone i executes an action according to the current state and observation Update the state of the drone after interacting with the environment to At the same time, obtain the reward feedback of the environment and the new observation The environment represents the target area;

[0016] The reward feedback of the environment and the execution actions of the drone The observed values form a state tuple and are saved in the experience replay pool maintained by the value function network.

[0017] Furthermore, to select any state tuple and use the message passing mechanism based on the attention network to extract the eigenvalue of each drone when performing multiple rounds of message passing with neighboring drones, the following steps are included:

[0018] Select d state tuple data from the current experience replay pool to form

[0019] According to the state tuple of the i-th drone in D Obtain the state information p of the i-th drone i , and according to the group encoding g of this drone i Obtain the position vector p' after encoding the group information of the drone i = concat[p i , g i , and traverse N drones to get [p' 1 , p' 2 ,..., p' N ;

[0020] Define the message sequence received by drone i from neighboring drones as k refers to the message in the k-th round of transmission during one message passing. For each drone i: For the message received by drone i at the initial stage is

[0021] The drone performs k rounds of message passing with neighboring drones, and each round of message passing is based on the message processed in the previous round; and according to the attention network mechanism, the neighboring heterogeneous information is encoded through the attention unit to obtain Q, K, V, which is expressed as:

[0022]

[0023] where W Q , W K , W V are different linear transformation matrices of the attention unit. Then the message transmitted by the drone in the (k + 1)-th round is:

[0024]

[0025] where d k is the dimension of the query vector Q, and f aRepresents the attention network;

[0026] The eigenvalue is obtained after k rounds of transmission Among them, the message obtained by the i-th drone after the k-th round of transmission Is expressed as:

[0027]

[0028] Furthermore, the state value function CRITIC of each drone is calculated according to the eigenvalue. The state value function network calculation formula when the drone i is in the state s i Is:

[0029]

[0030] Among them, f w Represents a fully connected network, and all drones share a state value function network V φ .

[0031] Furthermore, it also includes calculating the reward of each drone i, the global reward r glob , the local reward r loc And the final reward r used by the drone i to train the state value function network parameters i train , specifically expressed as:

[0032]

[0033] Among them, k 1 , k loc , k glob Are the reward weights corresponding to the reward, the local reward r loc , the global reward r glob respectively, Is the flight energy consumption coefficient of the i-th drone.

[0034] Furthermore, the parameters of the state value function network CRITIC and the policy network ACTOR are updated by using the stochastic gradient descent method. Specifically, it includes:

[0035] The parameters φ of the state value function network are updated by using the stochastic gradient descent method, expressed as

[0036]

[0037] Among them, Is the sample of the k-th iteration, and the objective value of the function Is the cumulative discounted return, which is estimated by the Monte Carlo cumulative return estimation method and is expressed as:

[0038]

[0039] Among them, γ is the discount rate;

[0040] Updating the policy network parameter θ using the stochastic gradient descent method is expressed as:

[0041]

[0042] Among them, π θ (a t |s t ) is the policy function, is the state-action value function, V φ (s t ) is the value function, and ε ∈ [0, 1] is a custom hyperparameter.

[0043] The present invention also proposes a control system for heterogeneous UAV area coverage, including:

[0044] An acquisition module for acquiring historical interaction data of multiple UAVs and a target area; the interaction data includes the current observation data of each UAV itself and the observation data of neighboring UAVs;

[0045] A model construction module for training the ACTO R-CRITIC reinforcement learning method using the historical interaction data of multiple UAVs and a target area to construct a heterogeneous UAV coverage network model; specifically including: based on the current observation data of each UAV itself and the observation data of neighboring UAVs, obtaining the reward feedback of the target area, the current observation value of the UAV, and the new observation value obtained after executing an action, and forming the new observation value into a state tuple and storing it in the experience replay pool of the state value function network CRITIC, selecting any state tuple, using the message passing mechanism based on the attention network to extract the eigenvalue of each UAV when performing multiple rounds of message passing with neighboring UAVs; generating the state value function network CRITIC of each UAV according to the eigenvalue; using the stochastic gradient descent method to update the parameters of the state value function network CRITIC and the policy network ACTOR parameters to construct a heterogeneous UAV coverage network model;

[0046] A coverage module for inputting the real-time interaction data of multiple UAVs and a target area into the heterogeneous UAV coverage network model and outputting the expected attitude of each UAV; each UAV realizes the coverage of the target area by tracking its own expected attitude.

[0047] The present invention also proposes a control computer device for heterogeneous UAV area coverage, including: a memory, a processor, and a computer program stored in the memory, and the processor implements the steps of the control of heterogeneous UAV area coverage when executing the computer program.

[0048] The present invention also provides a readable storage medium storing a computer program, where the computer program includes program instructions, and when the program instructions are executed by a processor, they are used to execute the steps of the heterogeneous UAV area coverage method described above.

[0049] The present invention provides a heterogeneous UAV area coverage method, system and device, which have the following beneficial effects:

[0050] According to the present invention, by using the observation data of each UAV itself and the observation data of neighboring UAVs, a coverage network model of heterogeneous UAVs is constructed by using a multi-agent reinforcement learning algorithm. When covering a target area, no central control node or manual control is required; each UAV can autonomously calculate the current optimal coverage path plan based on its current observation of the environment and information interaction with neighboring UAVs, so as to complete the coverage path plan of heterogeneous UAVs, enabling heterogeneous UAVs to obtain an optimal control strategy and realizing real-time coverage planning and control. Description of the Drawings

[0051] Figure 1 It is a flowchart for training a heterogeneous UAV coverage network model in an embodiment of the present invention;

[0052] Figure 2 It is a schematic diagram of the structure of a state value function network using an attention mechanism in an embodiment of the present invention;

[0053] Figure 3 It is a schematic diagram of the structure of a heterogeneous UAV policy network in an embodiment of the present invention;

[0054] Figure 4 It is a block diagram of a heterogeneous UAV coverage system based on multi-agent reinforcement learning in an embodiment of the present invention;

[0055] Figure 5 It is a hardware architecture diagram of a heterogeneous UAV coverage system based on multi-agent reinforcement learning in an embodiment of the present invention;

[0056] Figure 6 It is a flowchart of a method for heterogeneous UAV coverage in an embodiment of the present invention. Detailed Embodiments

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0058] The present invention proposes a method for heterogeneous UAV coverage, as Figure 6As shown in the figure; the method of reinforcement learning is used to complete the coverage path planning of heterogeneous drones, and an attitude controller is designed to track the attitude output by the neural network. Specifically, it includes the offline training stage of the deep reinforcement learning network and the online usage stage; in the offline training, the present invention adopts the ACTOR-CRITIC method to train the network, as Figure 1 shown. Among them, the ACTOR network is used to generate the expected policy of each drone at the next moment, and the CRITIC network is used to evaluate the reward of the drone after executing the current action. In addition, the specific structure of the CRITIC network is composed of an attention unit and an MLP fully connected layer, and the structure is as Figure 2 shown. The purpose of such design is to enable the network to extract effective information from a large amount of information, so as to reduce the instability caused by the continuous change of the learning target. The structure of the ACTOR network is as Figure 3 shown, which consists of two layers of fully connected neuron networks and one layer of non-linear activation unit RELU; specifically includes the following steps:

[0059] Steps in the offline training stage:

[0060] S1. According to the heterogeneous characteristics of the drones, group and encode the drones. The group encoding of drone i is g i For agents in the same group, they have the same encoding parameters; initialize the ACTOR and CRITIC networks, and initialize the network parameters as θ, θ', φ, φ' respectively. Initialize the message passing mechanism based on the attention network for each drone, and its role is to establish a communication mechanism between drones, and its purpose is to enable drones to obtain more comprehensive information.

[0061] S2. Collect data on the interaction between the drones and the environment. Define the vertex coordinates of the area to be covered as e i , the coordinates of drone i as p i , and each drone i obtains the observation value of the current state s t as:

[0062]

[0063] Each drone i obtains the state information p of the i-th drone from the observation i at the current moment, and obtains the new state information p’ i according to the group encoding g i of this drone = concat[p i , g i . Traverse N agents to get [p 1 ', p' 2 ,..., p' N ​As the input of the ACTOR network, and then obtain the network output As the expected action of the drone at the next moment. After the drone completes the action Update to the new state s t →s t+1 The CRITIC network evaluates the effectiveness of the action and gives a reward feedback And obtain the observation at the next moment The reward feedback And the executed action of the drone The observation values form a state tuple Save it to the experience replay pool maintained by the value function network

[0064] S3. Train the ACTOR and CRITIC networks. After step S2 is executed a finite number of times, d state tuple data are taken from the current experience replay pool To form To obtain more comprehensive global information and avoid the influence of local observation problems on the algorithm convergence, the drone performs multiple rounds of message passing with neighbor drones, uses a message passing mechanism based on an attention network to extract features, and obtains feature values (f 1 , f 2 ,... f N ), where each round of message passing is based on the message processed in the previous round. The specific steps are as follows

[0065] S3.1. Obtain the state information p of the i-th drone from the state tuple of the i-th drone in D And obtain the new state information p i According to the group encoding g of this drone i ' = concat[p i , g i , g i , and traverse N agents to get [p’ 1 , p' 2 ,..., p' N .

[0066] S3.2. Define the message sequence received by agent i from neighbor agents as Where k refers to the message passed in the k-th round in one message passing. For each drone i, initialize The message received by drone i at the initial stage is

[0067] The agent and neighbor agents perform k rounds of message passing. Each round of message passing is based on the message processed in the previous round. According to the basic principle of the attention unit, the drone encodes the neighbor information through the attention unit to obtain Q, K, V

[0068]

[0069] Among them, h(*) represents the non-linear encoding of the neural network, and W Q , W K , W V are different linear transformation matrices of the attention unit. Then the message transmitted by the UAV in the (k + 1)-th round is:

[0070]

[0071] where d k is the dimension of the query vector Q.

[0072] After k rounds of propagation, the eigenvalue According to the above formula, the message obtained by the i-th agent after k rounds of propagation can be expanded as follows:

[0073]

[0074] S4. Calculate the state value function v 1 , f 2 ,... f N ) of the data unit from the eigenvalues (f φ (f 1 ), v φ (f 2 ),... v φ (N). All agents share a CRITIC network V φ . For agent i, the calculation method of the value function when it is in state s i is:

[0075]

[0076] where f w represents a fully connected network, f a represents an attention network, and all UAVs share a value function network V φ .

[0077] S5. Update the parameters φ of the ACTOR network.

[0078] S5.1. Design the reward function based on the reward reconstruction mechanism.

[0079] Given a target area Q ∈ R 2 , define the points in the target area as q, and the set of points in the target area is represented by s obj . The coverage range of UAV i is defined as:

[0080]

[0081] The scalar function φ(q): R 2 →R + is a mapping from the coordinates of each point in the region to the interval (0, 1), which is used to represent the importance of the point. The position of the controller is p μ , and the rate of change of the weight is φ(q) is defined as:

[0082]

[0083] where in the heterogeneous UAV swarm, different UAVs have different energy loss coefficients.

[0084] In the design of the reward function for heterogeneous multi-UAV coverage control, due to the differences among agents, the reward function is slightly different from that of the homogeneous multi-agent algorithm. Agents have different speeds and coverage ranges. In this case, using the global reward to train the model will lead to the emergence of lazy agents. Therefore, it is considered to introduce a reward based on local observations for each agent on the basis of the global reward to solve the problem of reputation allocation of agents regarding the global reward. First, calculate the reward of each agent i by the following formula. The global reward r glob and the local reward r loc and the final reward r i train used by agent i to train the parameters of the value function network, where: and φ(q) are defined as above.

[0085]

[0086] r i train = k glob ·r glob + k loc ·r loc

[0087] S5.2. Update the parameters φ of the ACTOR network using the stochastic gradient descent method.

[0088]

[0089] Define as the sample at the k-th iteration, γ as the discount rate, and the target value of the function is the cumulative discounted return. The Monte Carlo cumulative return estimation method is adopted as follows:

[0090]

[0091] S6. Update the parameters θ of the policy network using the stochastic gradient descent method through the following formula:

[0092]

[0093] S7. Keep the target network parameters the same as the original network parameters, i.e., θ′ = θ and φ′ = φ.

[0094] S8. Repeat S3 - S7 until the network training is completed.

[0095] The steps in the online usage phase include:

[0096] S1. Each drone i obtains the observation at the current moment

[0097] S2. Based on the information interaction mechanism, each drone exchanges the observation values with its neighboring drones. Assume the state information p i of the i-th drone, and according to the group encoding g i of this drone, obtain the new state information p’ i = concat[p i , g i . Traverse N agents to get [p’ 1 , p' 2 ,..., p' N .

[0098] S3. Drone i takes [p’ 1 , p' 2 ,..., p' N as the input of the ACTOR network, and then obtains the network output as its expected action at the next moment.

[0099] S4. Inside each drone, perform the tracking of its respective expected trajectory. Figure 4 is a schematic diagram of the tracking control algorithm architecture. For the speed information output by the policy network, the expected pose of the drone at the next time step can be obtained. The attitude controller of the drone obtains the control torque based on the current pose of the drone and the expected pose output by the neural network, and outputs it as an instruction to the drone model. The drone outputs the motor speed according to the relationship between the motor and the control torque. The framework is as Figure 3 shown. The output of the policy network is the speed of the drone. First, simplify the drone particle and the policy output of the network to a first-order model as follows:

[0100]

[0101] where s irepresents the positions of the current UAV \(i\) and its neighboring UAVs relative to the global coordinate system. \(x\), \(v\), \(m\), \(R\), \(\Omega\), \(J\), \(M\) respectively represent the position, velocity, mass, attitude matrix, rotational angular velocity, moment of inertia, and torque input of the UAV. For the following flight dynamics model of the UAV:

[0102]

[0103] The differential equation of the attitude error can be obtained as follows, where the subscripts \(e\) and \(d\) represent the error value and the expected value of each physical quantity respectively.

[0104]

[0105] Then the expected torque input is:

[0106]

[0107] Among them, the attitude error matrix can be obtained by mutual conversion according to the quaternion error vector \((q e0 , q e1 , q e2 , q e3 ) and the error matrix, where \(q ei represents the \(i\)-th dimension error. The conversion relationship is as shown in the formula:

[0108]

[0109] Each UAV uses an external positioning system to obtain its own position, such as GPS, which provides the coordinates of each UAV in real time. The systems communicate through ROS, as shown in Figure 5 . The host subscribes to the coordinate topic published by VICON, so as to obtain the positions of each UAV in the UAV swarm in real time. The model is deployed on the upper computer. After the host gets the data, it converts the data into the standard input of the neural network, and then transmits the data to the neural network to get the action output. The next waypoint of the UAV is obtained through the obtained action and the current position of the UAV, and then the waypoint is published to each UAV through the PA wireless sensor matching with CRAZYFLIE. The attitude tracking controller of each UAV tries to track the target attitude to complete the coverage of the target area.

[0110] Based on the same inventive concept, the present invention also proposes a control system for heterogeneous UAV area coverage, including:

[0111] An acquisition module, configured to acquire historical interaction data of multiple UAVs and a target area; the interaction data includes the current own observation data of each UAV and the observation data of neighboring UAVs.

[0112] A model construction module is used to train the ACTOR-CRITIC reinforcement learning method by using the historical interaction data between multiple drones and the target area, and construct a heterogeneous drone coverage network model. Specifically, it includes: based on the current observation data of each drone itself and the observation data of neighboring drones, obtaining the reward feedback of the target area, the current observation value of the drone, and the new observation value obtained after performing an action, and forming the new observation value into a state tuple and saving it to the experience replay pool of the state value function network CRITIC. Selecting any state tuple, using the message passing mechanism based on the attention network to extract the eigenvalue of each drone when performing multiple rounds of message passing with neighboring drones; generating the state value function network CRITIC of each drone; using the stochastic gradient descent method to update the parameters of the state value function network CRITIC and the parameters of the policy network ACTOR, and constructing a heterogeneous drone coverage network model.

[0113] A coverage module is used to input the real-time interaction data between multiple drones and the target area into the heterogeneous drone coverage network model and output the expected attitudes of each drone; each drone realizes the coverage of the target area by tracking its own expected attitude.

[0114] The present invention also proposes a control computer device for heterogeneous drone area coverage, including: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of controlling heterogeneous drone area coverage are realized.

[0115] The present invention also proposes a readable storage medium. The readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, they are used to execute the steps of the heterogeneous drone area coverage method.

[0116] As described above, only the preferred specific implementation manners of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A heterogeneous UAV area coverage method, characterized in that: The following steps are involved: Collect historical interaction data between multiple drones and the target area; the interaction data includes each drone's current observation data and neighboring drones' observation data; The ACTOR-CRITIC reinforcement learning method is trained using historical interaction data between multiple drones and the target area to construct a heterogeneous drone coverage network model. Specifically, based on each drone's current observation data and the observation data of its neighboring drones, the reward feedback of the target area, the drone's current observation value, and the new observation value obtained after the action is executed are obtained, and the new observation value is combined into a state tuple and saved in the experience replay pool of the state value function network CRITIC. Any state tuple is selected, and the message passing mechanism based on the attention network is used to extract the feature value of each drone through multiple rounds of message passing with neighboring drones. Generate the state value function network CRITIC of each drone based on the eigenvalues; use the stochastic gradient descent method to update the state value function network CRITIC parameters and the policy network ACTOR parameters to build a heterogeneous drone coverage network model; The real-time interaction data between multiple drones and the target area is input into the heterogeneous drone coverage network model, and the expected posture of each drone is output; each drone achieves coverage of the target area by tracking its own expected posture.

2. A heterogeneous UAV area coverage method according to claim 1, characterized in that: The step of obtaining the reward feedback of the target area, the current observation value of the drone, and the new observation value obtained after executing the action, and storing the new observation value into a state tuple in the experience playback pool of the state value function network CRITIC, specifically includes the following steps: The set of m vertex coordinates defining the target area is [e1,...,e m ], the coordinates of drone i are p i , the N closest to UAV i i The coordinate set of the UAVs is Each drone i obtains the current environment s t Observed value for: According to the current strategy network π θ Execute an action Where θ is the policy network parameter; Each drone i is based on the current state With observation Execute an action After interacting with the environment, the drone's state is updated to At the same time, get reward feedback from the environment and new observations The environment represents a target area; Feedback the environment’s rewards And the execution actions of the drone Observations form state tuples Save to the experience replay pool maintained by the value function network.

3. A heterogeneous UAV area coverage method according to claim 2, characterized in that: The method of selecting any state tuple and using the attention network-based message passing mechanism to extract the feature value of each drone after multiple rounds of message passing with neighboring drones includes the following steps: Select d state tuple data from the current experience replay pool constitute According to the state tuple of the i-th drone in D Get the status information p of the i-th drone i , and according to the group code g of the drone i The position vector p after obtaining the encoded drone group information i '=concat[p i ,g i ], traverse N drones to obtain [p1',p'2,...,p' N ]; Define the message sequence that drone i receives from neighboring drones as k refers to the message delivered in the kth round in a message delivery. For each drone i: For the initial stage, the message received by drone i is The drone and its neighboring drones perform k rounds of message transmission, and each round of message transmission is based on the message processed in the previous round; and according to the attention network mechanism, the neighbor heterogeneous information is encoded through the attention unit to obtain Q, K, V, which is expressed as: Where W Q ,W K ,W V is the linear transformation matrix of different attention units, then the message transmitted by the drone in the k+1th round is: where d k is the dimension of the query vector Q, f a represents the attention network; After k rounds of transmission, the eigenvalues ​​are obtained Among them, the message obtained by the i-th drone after the k-th round of transmission is It is expressed as:

4. A heterogeneous UAV area coverage method according to claim 3, characterized in that: The state value function CRITIC of each drone is calculated according to the characteristic value, and the drone i is in state s i The state value function network calculation formula at this time is: Among them, f w Represents a fully connected network, and all drones share a state-value function network V φ .

5. A heterogeneous UAV area coverage method according to claim 4, characterized in that: It also includes using the reward reconstruction mechanism to calculate the reward for each drone i and the global reward r glob , local reward r loc and the final reward r of drone i for training the state-value function network parameters i train , specifically expressed as: Among them, k1,k loc ,k glob Reward, local reward r loc 、Global Reward r glob The corresponding reward weight, k2 i is the flight energy consumption coefficient of the i-th UAV.

6. A heterogeneous UAV area coverage method according to claim 4, characterized in that: The stochastic gradient descent method is used to update the state value function network CRITIC parameters and the policy network ACTO R parameters, specifically including: The stochastic gradient descent method is used to update the state value function network parameter φ, which is expressed as in, is the sample of k iterations, the target value of the function is the cumulative discounted return, using the Monte Carlo cumulative return estimation method, which is expressed as: Among them, γ is the discount rate; The stochastic gradient descent method is used to update the policy network parameters θ as follows: Among them, π θ (a t |s t ) is the policy function, is the state-action value function, V φ (s t ) is the value function, and ε∈[0,1] is a custom hyperparameter.

7. A heterogeneous UAV regional coverage system, characterized in that: include: A collection module is used to collect historical interaction data between multiple drones and a target area; the interaction data includes each drone's current observation data and neighboring drones' observation data; The model building module is used to train the ACTO R-CRITIC reinforcement learning method using the historical interaction data between multiple drones and the target area, and construct a heterogeneous drone coverage network model; specifically, it includes: based on the current observation data of each drone and the observation data of neighboring drones, obtaining the reward feedback of the target area, the current observation value of the drone, and the new observation value obtained after the action is executed, and the new observation value is composed of a state tuple and saved in the experience playback pool of the state value function network CRITIC, selecting any state tuple, and using the message passing mechanism based on the attention network to extract the characteristic value of each drone when it passes multiple rounds of messages with neighboring drones; generating the state value function network CRITIC of each drone according to the characteristic value; using the stochastic gradient descent method to update the state value function network CRITIC parameters and the policy network ACTOR parameters, and constructing a heterogeneous drone coverage network model; The coverage module is used to input the real-time interaction data between multiple drones and the target area into the heterogeneous drone coverage network model and output the expected posture of each drone; each drone achieves coverage of the target area by tracking its own expected posture.

8. A heterogeneous UAV area coverage computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of the heterogeneous UAV area coverage method described in any one of claims 1 to 6 are implemented.

9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the heterogeneous drone area coverage method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle autonomous coverage method based on pheromone inspiration

    CN116643587A

  • Multi-unmanned aerial vehicle path planning method and device based on deep reinforcement learning

    CN118295452A

  • Flight decision generation method and apparatus, computer device, and storage medium

    WO2023142316A1

Cited By

  • River and lake inspection image real-time transmission method based on unmanned aerial vehicle group

    CN121864953A