Resource allocation and trajectory planning method for UAV data acquisition system based on non-orthogonal multiple access assistance

By using a drone data acquisition system based on non-orthogonal multiple access, combined with deep reinforcement learning and mixed integer programming, the channel allocation and trajectory planning of drones are optimized, solving the low latency and high efficiency problems in sensor data acquisition and achieving more efficient data acquisition.

CN119767266BActive Publication Date: 2025-09-26ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411711671.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-09-26
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Sensor data collection faces the difficulties of achieving low latency, high energy efficiency and high reliability requirements. Traditional orthogonal multiple access cannot meet the needs of a large number of sensor accesses. The complexity of hybrid non-orthogonal multiple access leads to increased receiver delay. In addition, the drone data collection optimization problem is a non-convex problem and is difficult to solve using traditional optimization methods.

Method used

A UAV data acquisition system assisted by non-orthogonal multiple access is adopted. Through channel allocation and trajectory planning methods, deep reinforcement learning and mixed integer programming are used to decompose the optimization problem into discrete and continuous variable sub-problems. Combined with the trajectory planning neural network and the channel allocation neural network, the proximal gradient optimization and distributed soft actor-critic algorithm are used to optimize the UAV's flight trajectory and channel allocation.

Benefits of technology

It effectively reduces the time of drone data collection tasks, improves spectrum utilization and data collection efficiency, and achieves faster data collection completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119767266B_ABST
    Figure CN119767266B_ABST
Patent Text Reader

Abstract

The present application discloses a resource allocation and trajectory planning method for a UAV data acquisition system assisted by non-orthogonal multiple access. This application designs an optimization problem for minimizing the UAV acquisition task time while ensuring complete data acquisition and taking into account the UAV's speed, three-dimensional trajectory, channel allocation of sensor nodes, and access time slot length. For the H-NOMA channel allocation problem, a DRL algorithm based on a combination of rolling baseline and proximal gradient optimization is designed. For the same DSAC trajectory optimization algorithm, the RPPO channel allocation algorithm proposed in this application has a shorter completion time; this application designs a distributed DSAC algorithm, and the DSAC trajectory optimization algorithm proposed in this application has a shorter completion time. The RPPO algorithm and DSAC algorithm adopted in this application can optimize the acquisition task time in a UAV-NOMA data acquisition system with dynamic channel access, and the task time is lower than the benchmark method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless communication technology, and in particular to a resource allocation and trajectory planning method for an unmanned aerial vehicle (UAV) data acquisition system assisted by non-orthogonal multiple access. Background Art

[0002] With the development of the Internet of Things (IoT) technology, sensors are increasingly being used in applications such as hydrological monitoring, traffic control, smart cities, and smart agriculture. It is predicted that by 2030, the number of sensor nodes (SNs) worldwide will exceed one trillion. This massive number of sensors poses challenges to sensor data collection, such as the difficulty in achieving low latency, high energy efficiency, and high reliability. Unmanned aerial vehicles (UAVs) offer excellent coverage and channel characteristics. UAVs use sensors to establish line-of-sight channels, significantly improving communication performance. Furthermore, the high maneuverability of UAVs facilitates sensor data collection, especially in complex or remote environments. Using UAVs for data collection offers advantages such as wide coverage, high collection efficiency, and low energy consumption. Therefore, UAV-assisted data collection is a promising technology. In recent years, research on UAV data collection has included two-dimensional and three-dimensional trajectory planning for UAVs, sensor transmit power, and channel allocation.

[0003] As the number of sensors increases, limited communication resources are one of the factors that constrain the development of data collection. Traditional orthogonal multiple access (OMA) struggles to accommodate a large number of sensors. Fortunately, non-orthogonal multiple access (NOMA) allows multiple users to access the same resource block (RB) simultaneously and uses successive interference cancellation (SIC) at the receiver to recover user data. NOMA can effectively improve spectrum efficiency and consistently outperforms OMA in any communication system, even when both employ optimal resource allocation. Therefore, NOMA is being applied in feasible drone data collection networks. While NOMA can achieve higher spectrum efficiency through SIC, the complexity of SIC increases linearly with the number of users, significantly increasing latency at the receiver. Hybrid NOMA (H-NOMA) is a compromise between latency and spectrum efficiency. Users are grouped into clusters. Users within a cluster use NOMA to access the same RB, while users between clusters use OMA to access different RBs. The performance of H-NOMA networks depends on the clustering method. Experiments have shown that H-NOMA can effectively improve the efficiency of drone data collection.

[0004] Optimization problems in drone data collection are often non-convex, making them difficult to solve using traditional optimization methods. Deep reinforcement learning (DRL) leverages the trial-and-error learning of reinforcement learning and the powerful data processing capabilities of deep neural networks (DNNs) to solve these problems. In short, DRL uses the DNN to output feasible solutions to the problem. The agent executes these feasible solutions to obtain corresponding rewards. The agent then optimizes the DNN's weight parameters to maximize the cumulative reward. After iterative learning, the DNN requires only a simple forward propagation to obtain an approximate solution to the problem. In recent years, DRL has been used to solve drone control problems. Summary of the Invention

[0005] This application provides a resource allocation and trajectory planning method for a drone data acquisition system based on non-orthogonal multiple access assistance. The purpose of this application is to minimize the time required to complete the drone data acquisition task by optimizing the drone's channel allocation algorithm, trajectory planning, and time slot allocation algorithm.

[0006] This application provides a resource allocation and trajectory planning method for a UAV data acquisition system based on non-orthogonal multiple access assistance, the method comprising:

[0007] Step 1: Model the H-NOMA-assisted UAV data collection system and determine the minimum UAV data collection optimization problem;

[0008] Step 2: Decompose the optimization problem P0 into two sub-problems: the first sub-problem P1 and the second sub-problem P2. The optimization variables of the first sub-problem P1 are all discrete variables, while the optimization variables of the second sub-problem P2 are all continuous variables.

[0009] Step 3: The drone uses a trajectory planning neural network to decide the flight direction, speed, and duration of the drone's next time slot.

[0010] Step 4: Obtain the channel gain of the current UAV position sensor, select the top M sensors with the highest channel gain for data collection, and use the channel allocation neural network to allocate channels to the M sensors.

[0011] Step 5: Update the channel assignment neural network weights using the trajectory reference mechanism and proximal gradient optimization algorithm;

[0012] Step 6: Calculate and save the amount of data collected by the drone in the current time slot, the state of the drone in the previous time slot, and the current state; loop through steps 3 to 5 until all sensor data collection is completed;

[0013] Step 7: The drone returns to the starting point to unload the experience and uses the SAC algorithm to update the neural network weights;

[0014] Step 8: Loop through steps 3 to 6 until the channel allocation algorithm, trajectory planning, and time slot allocation algorithm converge.

[0015] Step 1: Model the H-NOMA-assisted UAV data collection system and determine the minimum UAV data collection optimization problem.

[0016] Consider the data acquisition problem of a single UAV with N sensors. The sensor set is represented as The drone starts from the starting position and uses H-NOMA to collect data from each sensor. After all sensors have collected data, the drone returns to the starting position.

[0017] Assuming that the sensor position is known and unchanged within the time of the acquisition task completion, the UAV locates the sensor position based on the received signal strength RSS;

[0018] Without loss of generality, the G2A channel is considered to be a fading channel; the coordinates of the sensor and the UAV at the tth time slot are denoted as SP n =[x n ,y n ,0] and Q u (t) = [x u (t),y u (t),z u (t)], the distance from the UAV to the nth sensor in the tth time slot is:

[0019]

[0020] The n-th sensor channel gain in the t-th time slot is expressed as:

[0021]

[0022] In formula (2), α represents the path loss index, β0 represents the path loss value at the reference distance; g n (t) is the channel fading, which obeys the Rice distribution:

[0023]

[0024] In formula (3), g is the additional fading of the LoS component, is the fading of the NLoS component that obeys the Rayleigh distribution and is assumed to be constant within a time slot; K n (t) is the Rice fading factor. When there is no LoS ​​component between the UAV and the nth sensor in the tth time slot, the Rayleigh fading channel K n (t)=0, when Kn When (t) = ∞, it is a pure LoS channel;

[0025] Previous studies have shown that K n (t) is related to the elevation angle from the UAV to the sensor. As the elevation angle increases, the LoS component increases. K n (t) is modeled using the following exponential function:

[0026] K n (t) = A1exp(A2θ n (t)),(4)

[0027] θ n =arcsin(z n (t) / d n (t)), (5)

[0028] K in formulas (4) and (5) n (t)∈[K min ,K max ],A1=K min , Assume that the elevation angle of the UAV is constant within a time slot, that is, K n (t) is also constant within a time slot;

[0029] Use H-NOMA to collect sensor data; let L and M represent the number of channels and the maximum number of sensors that the drone can access in a time slot, respectively. denote the lth channel and the mth access user respectively; the channel gain from the sensor to the UAV in a time slot is expressed as the set H = {h1,…,h n ,…,h N}; s = argsort(H) indicates that sensors are sorted in descending order of channel gain; assuming that M sensors with the largest channel gain are selected for data collection in each time slot;

[0030] make represents the channel set of sensors accessing the drone in the current time slot; let U l represents the number of sensors allocated in each channel and satisfies All M sensors need to be assigned to channels; define ω j,l ∈M represents the jth sensor j∈{1,…,U l},vector represents the set of sensors assigned to the lth channel, and the matrix Ω = [Ω1; ...; Ω L ] represents the allocation of M sensors on L channels;

[0031] The j-th signal-to-noise ratio SINR of the l-th channel is expressed as:

[0032]

[0033] In formula (6) represents the channel gain of the sensor assigned to the lth channel; P and σ 2 are the sensor's transmission power and additive white Gaussian noise power respectively; the transmission rate of the jth sensor to the lth channel is expressed as:

[0034]

[0035] In formula (7), ε g is the demodulation threshold of H-NOMA, B is the bandwidth of each channel;

[0036] Assume that the UAV needs to collect fixed C max The amount of bit data collected by the jth sensor of the lth channel in the tth time slot is expressed as:

[0037]

[0038] In formula (8), δ(t) is the duration of the time slot, is the amount of data remaining in the sensor;

[0039]

[0040] The optimization problem is used to minimize the time of the data collection task: by jointly optimizing the UAV flight speed v(t), horizontal direction angle θ(t) and vertical direction angle φ(t) in each time slot, the channel allocation method Ω(t) and the time slot duration δ(t), the data collection completion time is minimized. The optimization problem is expressed as:

[0041]

[0042] In formula (10), T is the total number of time slots; constraint C1 indicates that data from all sensors must be collected; C2-C5 represent the ranges of δ(t), v(t), θ(t), and φ(t), respectively; and condition C6 represents the flight range of the UAV.

[0043] Step 2: Since the optimization problem (P0) is a mixed integer programming problem, the optimization problem P0 is decomposed into two sub-problems: the first sub-problem P1 and the second sub-problem P2. The optimization variables of the first sub-problem P1 are all discrete variables, and the optimization variables of the second sub-problem P2 are all continuous variables.

[0044] The flight trajectory and time slot length of the drone are continuous variables, while the channel allocation is a discrete variable. Therefore, (P0) is a mixed integer programming problem. To facilitate problem solving, this application decomposes (P0) into two sub-problems (P1) and (P2).

[0045] The first sub-problem optimization problem P1 is expressed as:

[0046]

[0047] Formula (11) is a channel allocation problem within t time slots. The goal is to maximize the amount of data collected by the UAV in a single time slot. Different channel allocation methods Ω(t) will affect the data acquisition speed, thereby affecting the length of a single time slot δ(t) and the total number of time slots T. After solving problem (P1), the second sub-problem P2 is expressed as:

[0048]

[0049] In the second sub-problem P2, all optimization variables are continuous variables, which makes it easier to solve the problem.

[0050] In step 3, the drone uses a trajectory planning neural network to decide the flight direction, speed, and duration of the drone's next time slot.

[0051] Define the four-dimensional motion of the drone a t =[v(t),θ(t),φ(t),δ(t)] to represent the optimization variables of the second sub-problem P2, and at needs to satisfy the conditions C2-C5 in the second sub-problem P2; when executing action a t The position of the drone is:

[0052]

[0053] The action selection method adopts the random strategy method, that is, the neural network does not directly output a specific action, but outputs a parameter of a probability distribution related to the action, and generates a specific probability distribution based on the parameter for sampling, and then obtains a specific action, which makes the neural network more exploratory than the deterministic strategy; previous studies mostly used Gaussian distribution, but the sampling range of Gaussian distribution is infinite. In order to meet the needs of C2-C5, the Gaussian distribution is truncated, which reduces the learning efficiency. Therefore, this application uses β distribution as the random strategy distribution; unlike the Gaussian distribution, the sampling range of β distribution is [0,1], which helps to expand the action space to different ranges. The probability density function of β distribution is:

[0054]

[0055] In formula (14), a d >0 and βd >0 are the two determining parameters of the β distribution, Γ(·) is the Γ function; the neural network is used to output four independent sets of a d , β d , then according to a d , β d Generate four independent beta distributions and sample the distributions to obtain the drone action a t .

[0056] Step 4: Obtain the channel gain of the current UAV position sensor, select the top M sensors with the highest channel gain for data collection, and use the channel allocation neural network to allocate channels to the M sensors.

[0057] It is necessary to allocate channels to M sensors, and the state space of M steps State={state1,...,state m ,...,state M}, each step allocates a channel to a sensor, and the corresponding channel gain is set up represents the channel allocation status of the mth step, Indicates that the sensor is assigned to the mth channel of the lth channel; when all channels are assigned, SE m Expressed as:

[0058]

[0059] In formula (15), SE M The first U in the first row l The non-zero elements are Ω l (t), so according to SE M Get the corresponding Ω(t);

[0060] The goal of the first sub-problem P1 is to maximize the amount of data collected in the current time slot, so the current time slot length δ(t) and the remaining data collected by the sensor are Also added to state m In the state space, it is represented as:

[0061]

[0062] In the channel allocation algorithm, the m-th step action represents the channel gain The sensor is assigned to the action m channel, where

[0063] Use neural network to obtain channel allocation action; transform the state space state m Divided into three groups: [H s ,Rd], and SE m ; Among them, [H s ,Rd] represents the static information of the sensor, which remains unchanged during the channel allocation process; Indicates the information of the sensor that performs channel allocation in step m; at the same time, SE m Indicates the overall change of channels during the channel allocation process;

[0064] First, CNN is used to extract the s ,Rd] and Extract features and get ref1 and q1 respectively; use the attention mechanism to match the two features to get Logit, and then use CNN to match SE m The rows and columns are feature extracted to obtain ref2 and ref3; ref2 is multiplied by Logits to obtain q2, and then the attention mechanism is used to match ref3 with q2 to obtain the probability p of channel assignment. m ; sampling p m Produce a specific action m ;

[0065] The attention mechanism is specifically manifested as follows:

[0066]

[0067] In formula (17) W qk is the fully connected layer parameter, V k is a learnable parameter and d is the number of hidden units of CNN.

[0068] Step 5: Update the channel assignment neural network weights using the trajectory reference mechanism and proximal gradient optimization algorithm.

[0069] The reward of the channel assignment algorithm is set to the data collected by the UAV from each sensor, expressed as:

[0070]

[0071] Because all sensors need to be assigned channels to obtain rewards, similar to combinatorial optimization problems, the Rollout Baseline and Proximal Policy Optimization (PPO) algorithms are used to train the neural network. PPO is an on-policy algorithm, but compared to other on-policy algorithms, PPO uses importance sampling to train the same data multiple times, resulting in higher training efficiency. PPO is based on an actor-critic framework, where the actor is responsible for taking an action, while the value function estimation module (critic) is responsible for evaluating the quality of that action under the current state. This requires synchronous training of both networks. Rollout Baseline is a self-evaluation method, meaning that the actor and critic have the same network structure. During training, the critic's parameters are based on the actor's highest historical reward.

[0072] According to the PPO algorithm, the loss function of Actor is expressed as:

[0073]

[0074] In formula (19), represents the similarity between the new and old channel allocation strategies, entropy is the entropy of the current strategy, and ξ is the entropy coefficient. According to the Rollout Baseline, the advantage function adv is expressed as:

[0075]

[0076] In formula (20), reward m A and reward M C Represents the rewards obtained by Actor and Critic respectively; when the mean of adv is greater than 0 and the p-value of the one-sided t-test of adv is greater than 0.005, it indicates that the Actor has been significantly improved compared to the Critic. At this time, the Actor parameters are assigned to the Critic.

[0077] Step 6: Calculate and save the amount of data collected by the drone in the current time slot, the state of the drone in the previous time slot, and the current state; loop through steps 3 to 5 until all sensor data collection is completed.

[0078] Step 7: The drone returns to the starting point to unload the experience and uses the SAC algorithm to update the neural network weights.

[0079] This application uses a distributed soft actor-critic (SAC) algorithm based on DRL to solve the trajectory planning and time slot allocation problems.

[0080] In the trajectory planning algorithm introduced in this application, the state space includes the current position Q of the drone u (t), the position of the sensor {SP1,...,SP N} and the remaining data volume Rd of the sensor, the current state is expressed as:

[0081] s t =[Q u (t),SP1,SP2,...,SP N ,Rd]. (21)

[0082] The reward function in trajectory planning is designed as follows:

[0083]

[0084] In formula (22), μ is the execution t The penalty for failing to satisfy the C6 constraint prevents the drone from flying out of the boundary. t is the reward after the collection task is completed, F t Expressed as:

[0085]

[0086] Formula (23)F t It represents the time given to G minus the time consumed to complete the task collection, and ρ is the weight coefficient of time;

[0087] The SAC algorithm has three modules, namely the Actor module, the Critic module, and the Critic-Target module. The multi-layer perceptron MLP is used to construct the modules of the SAC algorithm. The Actor module is used to output the actions of the drone, while the Critic module and the Critic-Target module are used to evaluate the quality of the Actor module's actions. The Critic module and the Critic-Target module use clipped double-Q to avoid the problem of overestimation, that is, the Critic module and the Critic-Target module have two evaluation networks. Let the neural network parameters of the Critic module and the Critic-Target module be θ respectively. i and Where i∈{0,1} represents the i-th evaluation network; the loss function of the Critic module is:

[0088]

[0089] In formula (24)

[0090]

[0091] In formula (25), Qtar represents the learning objective of the action-value function; is in t Execute a in the state t The value of the action is given by the parameter θ i Neural network evaluation, γ is the discount factor, logπ φ (a t+1 |s t+1 ) is the entropy of the current policy, α H is the entropy weight coefficient; the parameters updated by Critic-Target using softupdate are as follows:

[0092]

[0093] In formula (26), τ∈[0,1] is the weight of softupdate; the loss function of the Actor module is expressed as:

[0094]

[0095] α H ∈[0,1] is set as a hyperparameter and adjusted manually, or it can be adjusted using self-learning methods. However, the self-learning method needs to determine the size of the target entropy. According to existing research, it is better to set the target entropy equal to the negative of the action space dimension. Then α H The loss function is expressed as:

[0096]

[0097] The distributed Apex architecture is used to implement the SAC algorithm. The architecture divides the SAC algorithm into four parts, namely the experience collection module rolloutworker, the experience cache module replaybuffer parameter server module parameter server and the SAC training module SAC trainer;

[0098] Specifically, when the drone starts from the starting point, the rollout worker perceives the state of the environment and obtains the action a t , and save the experience t ,a t ,r t ,s t+1}; After data collection is completed, the drone will return to the starting point and unload the experience into the experience cache module; then, the SAC training module performs neural network training by sampling experience from the experience cache module and uploads the trained neural network parameters to the parameter server module; when the drone performs the next collection mission, it downloads the latest parameters from the parameter server module.

[0099] Step 8: Loop through steps 3 to 6 until the channel allocation algorithm, trajectory planning, and time slot allocation algorithm converge.

[0100] This application proposes a channel allocation and trajectory planning method for drone data acquisition assisted by hybrid non-orthogonal multiple access (H-NOMA) that minimizes acquisition time. This method allows a larger number of users to complete data acquisition faster with the help of the RPPO and DSAC algorithms. This application decomposes the optimization problem into channel allocation, trajectory planning, and time slot allocation, and theoretically analyzes the solutions to these two optimization problems. This application is superior to other drone data acquisition time optimization methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] Figure 1 This is the model diagram of the H-NOMA UAV data acquisition system;

[0102] Figure 2 This is a diagram of the channel allocation neural network structure used in this application;

[0103] Figure 3 It is a schematic diagram showing the actions of the drone in this application;

[0104] Figure 4 This is the block diagram of the DSAC algorithm in this application;

[0105] Figure 5 It is the trajectory planning learning curve of DSAC and DDPG in Example 1 of this application;

[0106] Figure 6 is the 3D trajectory of the drone optimized by DSAC for different data amounts in Example 1 of this application, where (a) C max =10Mbit; (b) C max =100Mbit; (c)C max =200Mbit;

[0107] Figure 7 is the speed of the drone and the size of the collected data in each time slot in Example 1 of this application;

[0108] Figure 8 is the completion time of the drone data collection task under different data amounts in Example 1 of this application;

[0109] Figure 9 is the amount of data collected by the drone in each time slot in Example 1 of this application. DETAILED DESCRIPTION

[0110] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0111] This application provides a channel allocation and trajectory planning method for H-NOMA-assisted drone data acquisition that minimizes the completion time of the system drone data acquisition task. Figure 1 Model shown.

[0112] Step 1: Model the H-NOMA-assisted UAV data acquisition system and determine the optimization problem.

[0113] Consider the data acquisition problem of a single UAV with N sensors. The sensor set is represented as The drone starts from a starting position and uses H-NOMA to collect data from each sensor. After all sensors have collected data, the drone returns to the starting position. This application assumes that the sensor positions are known and unchanged during the acquisition task completion time. The drone can locate the sensor positions based on the received signal strength (RSS).

[0114] Without loss of generality, the G2A channel is considered to be a fading channel. The coordinates of the sensor and the UAV at the tth time slot are denoted as SP n =[x n ,y n ,0] and Q u (t) = [x u (t),y u (t),z u (t)], the distance from the UAV to the nth sensor in the tth time slot is:

[0115]

[0116] The n-th sensor channel gain in the t-th time slot is expressed as:

[0117]

[0118] In formula (30), α represents the path loss index, and β0 represents the path loss value at the reference distance. n (t) is the channel fading, which obeys the Rice distribution,

[0119]

[0120] In formula (31), g is the additional fading of the LoS component, The fading of the NLoS component follows the Rayleigh distribution, which is assumed to be constant within a time slot in this application. n (t) is the Rice fading factor. When there is no LoS ​​component between the UAV and the nth sensor in the tth time slot, the Rayleigh fading channel K n (t)=0, when Kn When (t) = ∞, it is a pure LoS channel. Previous studies have shown that K n (t) is related to the elevation angle from the UAV to the sensor. As the elevation angle increases, the LoS component increases. K n (t) can be modeled by the following exponential function

[0121] K n (t) = A1exp(A2θ n (t)),(32)

[0122] θ n =arcsin(z n (t) / d n (t)), (33)

[0123] K in formulas (32) and (33) n (t)∈[K min ,K max ],A1=K min , This application assumes that the elevation angle of the UAV is constant within a time slot, that is, K n (t) is also constant within a time slot.

[0124] H-NOMA is used to collect sensor data. Let L and M represent the number of channels and the maximum number of sensors that the drone can access in a time slot, respectively. Denote the lth channel and the mth access user respectively. The channel gain from the sensor to the UAV in a time slot is represented by the set H = {h1,...,h n ,...,h N}. s = argsort(H) indicates that sensors are sorted in descending order of channel gain. Assume that M sensors with the largest channel gain are selected in each time slot for data collection. Further let Denotes the channel set of sensors accessing the drone in the current time slot. Let U l represents the number of sensors allocated in each channel and satisfies This formula indicates that all M sensors need to be assigned to channels. Define ω j,l ∈M represents the jth (j∈{1,...,U l}) sensors, vector represents the set of sensors assigned to the lth channel, and the matrix Ω = [Ω1; ...; Ω L ] represents the allocation of M sensors on L channels.

[0125] The jth signal-to-noise ratio (SINR) of the lth channel is expressed as

[0126]

[0127] In formula (34) represents the channel gain of the sensor assigned to the lth channel. P and σ 2 are the sensor's transmission power and additive white Gaussian noise power respectively. The transmission rate of the jth sensor to the lth channel is expressed as:

[0128]

[0129] In formula (35), ε g is the demodulation threshold of H-NOMA, and B is the bandwidth of each channel.

[0130] Assume that the UAV needs to collect fixed C max The amount of data collected by the jth sensor of the lth channel in the tth time slot is expressed as:

[0131]

[0132] In formula (36), δ(t) is the duration of the time slot, is the amount of data remaining in the sensor.

[0133]

[0134] In step 1, an optimization problem is designed to minimize the time required for the data collection task. By jointly optimizing the UAV flight speed v(t), horizontal angle θ(t), and vertical angle φ(t) for each time slot, the channel allocation method Ω(t), and the time slot duration δ(t), the data collection completion time is minimized. The optimization problem is expressed as:

[0135]

[0136] In Equation (38), T is the total number of time slots. Constraint C1 indicates that data from all sensors must be collected. C2-C5 represent the ranges of δ(t), v(t), θ(t), and φ(t), respectively. Condition C6 represents the flight range of the drone.

[0137] Step 2: Since (P0) is a mixed integer programming problem, the optimization problem (P0) is decomposed into two sub-problems (P1) and (P2), where the optimization variables of (P1) are all discrete variables and the optimization variables of (P2) are all continuous variables.

[0138] The flight trajectory and time slot length of the drone are continuous variables, while the channel allocation is a discrete variable. Therefore, (P0) is a mixed integer programming problem. To facilitate problem solving, this application decomposes (P0) into two sub-problems (P1) and (P2).

[0139] Therefore, the optimization problem (P1) is expressed as:

[0140]

[0141] Equation (39) shows that this is a channel allocation problem within time slots t. The goal is to maximize the amount of data collected by the drone within a single time slot. Different channel allocation methods Ω(t) will affect the data acquisition speed, thereby affecting the length of a single time slot δ(t) and the total number of time slots T. After solving problem (P1), problem (P2) can be expressed as:

[0142]

[0143] In equation (40), all optimization variables in the (P2) problem are continuous variables, which facilitates the solution of the problem.

[0144] In step 3, the drone uses a trajectory planning neural network to decide the flight direction, speed, and duration of the drone's next time slot.

[0145] like Figure 2 As shown in the diagram of drone motion representation, this application defines the four-dimensional motion of the drone a t =[v(t),θ(t),φ(t),δ(t)] to represent the optimization variables of problem (P2), and a t The conditions C2-C5 in (P2) need to be met. t The position of the drone is:

[0146]

[0147] The action selection method adopts the random strategy method, that is, the neural network does not directly output a specific action, but outputs a parameter of a probability distribution related to the action, and generates a specific probability distribution based on the parameter for sampling, and then obtains a specific action, which makes the neural network more exploratory than deterministic strategies. Previous studies mostly used Gaussian distribution, but the sampling range of Gaussian distribution is infinite. In order to meet the needs of C2-C5, the Gaussian distribution is truncated, which reduces the learning efficiency. Therefore, this application uses β distribution as the random strategy distribution. Unlike the Gaussian distribution, the sampling range of β distribution is [0,1], which helps to expand the action space to different ranges. The probability density function of β distribution is:

[0148]

[0149] In formula (42), a d >0 and β d>0 are the two determining parameters of the β distribution, and Γ(·) is the Γ function. This application uses a neural network to output four independent sets of a d , β d , then according to a d , β d Generate four independent beta distributions and sample these distributions to obtain the drone action a t .

[0150] Step 4: Obtain the channel gain of the current UAV position sensor, select the top M sensors with the highest channel gain for data collection, and use the channel allocation neural network to allocate channels to the M sensors.

[0151] It is necessary to allocate channels to M sensors, so this application designs an M-step state space State={state1,...,state m ,...,state M}, each step allocates a channel to a sensor, and the corresponding channel gain is set up represents the channel allocation status of the mth step, Indicates that the sensor is assigned to the mth channel of the lth channel. When all channels are assigned, SE m Expressed as

[0152]

[0153] In formula (43), SE M The first U in the first row l The non-zero elements are Ω l (t), so according to SE M The corresponding Ω(t) can be obtained.

[0154] The goal of the optimization problem (P1) is to maximize the amount of data collected in the current time slot, so the current time slot length δ(t) and the remaining data collected by the sensor are Also added to state m The state space is represented as:

[0155]

[0156] In the channel allocation algorithm, the m-th step action represents the channel gain The sensor is assigned to the action m channel, where

[0157] Subsequently, the present application designed a neural network to obtain the channel allocation action. The structure diagram of the channel allocation neural network is shown in FIG. Figure 3As shown. This application will state space state m Divided into three groups: [H s ,Rd], and SE m Among them, [H s ,Rd] represents the static information of the sensor, which remains unchanged during the channel allocation process; Indicates the information of the sensor that performs channel allocation in step m; at the same time, SE m represents the overall change of the channel during the channel allocation process. First, CNN is used to extract the s ,Rd] and Extract features and get ref1 and q1 respectively. Use the attention mechanism to match the two features to get Logit, and then use CNN to SE m The rows and columns are feature extracted to obtain ref2 and ref3. Ref2 is multiplied by Logits to obtain q2, and then the attention mechanism is used to match ref3 with q2 to obtain the probability p of channel assignment. m . Sampling p m Produce a specific action m .

[0158] The attention mechanism is specifically manifested as follows

[0159]

[0160] In formula (45) W qk is the fully connected layer parameter, V k is a learnable parameter and d is the number of hidden units of CNN.

[0161] In step 5, the trajectory reference mechanism and the proximal gradient optimization algorithm update the channel allocation neural network weights.

[0162] The reward of the channel allocation algorithm is set to the data collected by the UAV from each sensor, expressed as

[0163]

[0164] Since channels need to be allocated to all sensors to obtain rewards, which is similar to a combinatorial optimization problem, this application refers to the Rollout Baseline and the existing popular proximal policy optimization (PPO) algorithm for neural network training. The PPO algorithm is an on-policy algorithm, but compared with other on-policy algorithms, PPO trains the same batch of data multiple times based on importance sampling, which has higher training efficiency. PPO is based on the actor-critic framework. The Actor is responsible for giving actions, and the value function estimation module (Critic) is responsible for evaluating the quality of the action in the current state, which requires the two networks to be trained synchronously. Rollout Baseline is a self-evaluation method, that is, the Actor and Critic have the same network structure, and during the training process, the parameters of the Critic are the parameters with the highest historical rewards of the Actor.

[0165] According to the PPO algorithm, the loss function of Actor is expressed as:

[0166]

[0167] In formula (47), Indicates the similarity between the new and old channel allocation strategies, entropy is the entropy of the current strategy, and ξ is the entropy coefficient. According to the Rollout Baseline, the advantage function adv is expressed as:

[0168]

[0169] In formula (48), reward m A and reward M C Represents the rewards received by the Actor and Critic, respectively. When the mean of adv is greater than 0 and the p-value of the one-sided t-test of adv is greater than 0.005, it indicates that the Actor has significantly improved compared to the Critic. At this time, the Actor parameters are assigned to the Critic.

[0170] Step 6: Calculate and save the amount of data collected by the drone in the current time slot, the drone's status in the previous time slot, and the current status. Repeat steps 3 to 5 until all sensor data collection is completed.

[0171] Step 7: The drone returns to the starting point to unload the experience and uses the SAC algorithm to update the neural network weights.

[0172] This application uses a distributed soft actor-critic (SAC) algorithm based on DRL to solve the trajectory planning and time slot allocation problems.

[0173] In the trajectory planning algorithm introduced in this application, the state space includes the current position Q of the drone u (t), the position of the sensor {SP1,...,SP N} and the remaining data amount Rd of the sensor, the current state can be expressed as

[0174] s t =[Q u (t),SP1,SP2,...,SP N ,Rd]. (49)

[0175] The reward function in trajectory planning is designed as follows

[0176]

[0177] In formula (50), μ is the execution t The penalty for failing to satisfy the C6 constraint prevents the drone from flying out of the boundary. t is the reward after the collection task is completed, F t Expressed as

[0178]

[0179] Formula (51)F t It represents the time given to G minus the time consumed to complete the task collection, and ρ is the weight coefficient of time.

[0180] The SAC algorithm has three modules, namely the Actor module, the Critic module, and the Critic-Target module. This application uses a simple multi-layer perceptron (MLP) to construct these modules. The Actor module is used to output the actions of the drone, while the Critic module and the Critic-Target module are used to evaluate the quality of the Actor module's actions. The Critic module and the Critic-Target module use clipped double-Q to avoid the problem of overestimation of value, that is, the Critic module and the Critic-Target module have two evaluation networks. Let the neural network parameters of the Critic module and the Critic-Target module be θ i and Where i∈{0,1} represents the i-th evaluation network. The loss function of the Critic module is

[0181]

[0182] In formula (52)

[0183]

[0184] In formula (53), Q tar represents the learning objective of the action-value function. It is in t Execute a in the state t The value of the action is given by the parameter θ i Neural network evaluation, γ is the discount factor, logπ φ (a t+1 |s t+1 ) is the entropy of the current policy, α H Is the entropy weight coefficient. The parameters updated by Critic-Target using softupdate are as follows

[0185]

[0186] In formula (54), τ∈[0,1] is the weight of softupdate. The loss function of the Actor module is expressed as

[0187]

[0188] α H ∈[0,1] is set as a hyperparameter and can be adjusted manually or by self-learning. However, the self-learning method requires determining the size of the target entropy. According to existing research, setting the target entropy equal to the negative of the action space latitude is more effective. Then α H The loss function is expressed as

[0189]

[0190] like Figure 4 As shown in the DSAC algorithm block diagram, this application uses a distributed Apex architecture to implement the SAC algorithm, which divides the SAC algorithm into four parts, namely the experience collection module (rollout worker), the experience cache module (replay buffer), the parameter server module (parameter server) and the SAC training module (SAC trainer). Specifically, when the drone starts from the starting point, the rollout worker perceives the environment state and obtains the action a. t , and save the experience t ,a t ,r t ,s t+1After data collection is complete, the drone returns to its starting point and offloads the experience to the experience cache module. The SAC training module then trains the neural network by sampling experience from the experience cache module and uploads the trained neural network parameters to the parameter server module. When the drone performs its next collection mission, it downloads the latest parameters from the parameter server module.

[0191] Step 8: Loop steps 3 to 6 until the channel allocation algorithm, trajectory planning, and time slot allocation algorithm converge.

[0192] The specific implementation examples of the present invention are as follows:

[0193] Example 1: Assume that the flight range of the drone is 2000m×2000m and the maximum flight altitude is 200m, that is, Q min =[-1000,1000,0],Q max = [-1000, 1000, 200]. Assume that N sensors are randomly distributed in the environment. The drone starts data collection at a random location and returns to its starting position after completion. Furthermore, assume that the sensor data size is 10-200 Mbit. The remaining detailed parameter settings are shown in Table 1.

[0194] Table 1 Simulation parameter settings

[0195]

[0196] The amount of data from each sensor is set to C max =200Mbit, DSAC and DDPG algorithms use the same number of pre-training steps, i.e. Ep=100.

[0197] The optimized results were obtained according to the specific method provided in this application.

[0198] Figure 5 The learning curves of the proposed DSAC algorithm and the existing popular trajectory planning algorithm DDPG are compared, where the solid line is the curve after sliding average and the shadow is the actual reward. Figure 5 It can be seen that the performance of DSAC algorithm after convergence is significantly better than that of DDPG algorithm. In addition, thanks to the maximum entropy learning, DSAC algorithm has better stability and smaller fluctuation after convergence than DDPG algorithm. max When the bandwidth is 100 Mbit, DDPG is difficult to converge, so the performance of the DDPG algorithm is not compared in the following text.

[0199] Figure 6 Shows different C max When the amount of sensor data is small, that is, C max= 10Mbit, the drone will choose a shorter flight path and a faster flight speed to obtain the fastest acquisition completion time. max When C increases, the drone will slow down appropriately in locations with a large number of sensors to ensure that data from each sensor is collected. For sensors at edge locations, the drone will approach the sensors to obtain faster data collection speed, although this will increase the flight distance to a certain extent. max = 200Mbit, in areas with dense sensors, the drone's speed will further decrease or even stop, while in areas with fewer sensors, the drone will choose to accelerate. This shows that the DSAC algorithm proposed in this application can well optimize the drone trajectory and minimize the acquisition completion time. Moreover, when C max When changes, the DSAC algorithm can also converge to a better trajectory.

[0200] Figure 7 The length of each time slot, the speed of the drone in each time slot, and the amount of data collected are given. It can be seen that the length of each time slot is relatively uniform, indicating that the time slot length has little impact on the optimization results. When the drone needs a longer time slot, it will choose to reduce its speed and choose the same action in the next time slot instead of changing the time slot length. Figure 7 It can also be seen that the drone will reduce its speed to ensure that more data can be captured in this time slot.

[0201] Figure 8 The performance of the proposed solution in terms of completion time is demonstrated. We compare the sensor access method between the OMA method and H-NOMA using H2T pairing. This application compares the Kmeans algorithm and the particle swarm optimization (PSO) algorithm for pre-training. Figure 8 The results show that when paired with the same channel allocation algorithm H2T, the proposed DSAC trajectory optimization algorithm has a shorter completion time compared to other benchmark trajectory optimization algorithms. For the same DSAC trajectory optimization algorithm, the proposed RPPO channel allocation algorithm has a shorter completion time compared to H2T and OMA.

[0202] Figure 9 By comparing the amount of data captured in each time slot under different methods, it can be found that the RPPO channel allocation algorithm can obtain the maximum amount of captured data in the same time slot.

[0203] This application proposes a channel allocation and trajectory planning method for drone data acquisition assisted by hybrid non-orthogonal multiple access (H-NOMA) that minimizes acquisition time. This method allows a larger number of users to complete data acquisition faster with the help of the RPPO and DSAC algorithms. This application decomposes the optimization problem into channel allocation, trajectory planning, and time slot allocation, and theoretically analyzes the solutions to these two optimization problems. This application is superior to other drone data acquisition time optimization methods.

[0204] The above-described embodiments of the present application do not constitute a limitation on the scope of protection of the present application.

Claims

1. A resource allocation and trajectory planning method for a UAV data acquisition system based on non-orthogonal multiple access assistance, characterized in that: The method comprises: Step 1: Model the H-NOMA-assisted UAV data collection system and determine the minimum UAV data collection optimization problem; Step 2: Decompose the optimization problem P0 into two sub-problems: the first sub-problem P1 and the second sub-problem P2. The optimization variables of the first sub-problem P1 are all discrete variables, while the optimization variables of the second sub-problem P2 are all continuous variables. Step 3: The drone uses a trajectory planning neural network to decide the flight direction, speed, and duration of the drone's next time slot. Step 4: Obtain the channel gain of the current UAV position sensor, select the top M sensors with the highest channel gain for data collection, and use the channel allocation neural network to allocate channels to the M sensors. Step 5: Update the channel assignment neural network weights using the trajectory reference mechanism and proximal gradient optimization algorithm; Step 6: Calculate and save the amount of data collected by the drone in the current time slot, the state of the drone in the previous time slot, and the current state; loop through steps 3 to 5 until all sensor data collection is completed; Step 7: The drone returns to the starting point to unload the experience and uses the SAC algorithm to update the neural network weights; Step 8: Loop through steps 3 to 6 until the channel allocation algorithm, trajectory planning, and time slot allocation algorithm converge.

2. The method according to claim 1, characterized in that Step 1: Model the H-NOMA-assisted UAV data collection system and determine the optimization problem of minimizing the UAV data collection, including: Consider the data acquisition problem of a single UAV with N sensors. The sensor set is represented as The drone starts from the starting position and uses H-NOMA to collect data from each sensor. After all sensors have collected data, the drone returns to the starting position. Assuming that the sensor position is known and unchanged within the time of the acquisition task completion, the UAV locates the sensor position based on the received signal strength RSS; The G2A channel is considered to be a fading channel; the coordinates of the sensor and the UAV at the tth time slot are represented as SP n =[x n ,y n ,0] and Q u (t) = [x u (t),y u (t),z u (t)], the distance from the UAV to the nth sensor in the tth time slot is: The n-th sensor channel gain in the t-th time slot is expressed as: In formula (2), α represents the path loss index, β0 represents the path loss value at the reference distance; g n (t) is the channel fading, which obeys the Rice distribution: In formula (3), g is the additional fading of the LoS component, is the fading of the NLoS component that obeys the Rayleigh distribution and is assumed to be constant within a time slot; K n (t) is the Rice fading factor. When there is no LoS ​​component between the UAV and the nth sensor in the tth time slot, the Rayleigh fading channel K n (t)=0, when K n When (t) = ∞, it is a pure LoS channel; K n (t) is related to the elevation angle from the UAV to the sensor. As the elevation angle increases, the LoS component increases. K n (t) is modeled using the following exponential function: K n (t)=A1exp(A2θ n (t)), (4) θ n =arcsin(z n (t) / d n (t)), (5) K in formulas (4) and (5) n (t)∈[K min ,K max ],A1=K min , Assume that the elevation angle of the UAV is constant within a time slot, that is, K n (t) is also constant within a time slot; Use H-NOMA to collect sensor data; let L and M represent the number of channels and the maximum number of sensors that the drone can access in a time slot, respectively. denote the lth channel and the mth access user respectively; the channel gain from the sensor to the UAV in a time slot is expressed as the set H = {h1,...,h n ,...,h N }; s = argsort(H) indicates that sensors are sorted in descending order of channel gain; assuming that M sensors with the largest channel gain are selected for data collection in each time slot; make represents the channel set of sensors accessing the drone in the current time slot; let U l represents the number of sensors allocated in each channel and satisfies All M sensors need to be assigned to channels; define ω j,l ∈M represents the jth sensor j∈{1,...,U l },vector represents the set of sensors assigned to the lth channel, and the matrix Ω = [Ω1; ...; Ω L ] represents the allocation of M sensors on L channels; The j-th signal-to-noise ratio SINR of the l-th channel is expressed as: In formula (6) represents the channel gain of the sensor assigned to the lth channel; P and σ 2 are the sensor's transmission power and additive white Gaussian noise power respectively; the transmission rate of the jth sensor to the lth channel is expressed as: In formula (7), ε g is the demodulation threshold of H-NOMA, B is the bandwidth of each channel; Assume that the UAV needs to collect fixed C max The amount of bit data collected by the jth sensor of the lth channel in the tth time slot is expressed as: In formula (8), δ(t) is the duration of the time slot, is the amount of data remaining in the sensor; The optimization problem is used to minimize the time of the data collection task: by jointly optimizing the UAV flight speed v(t), horizontal direction angle θ(t) and vertical direction angle φ(t) in each time slot, the channel allocation method Ω(t) and the time slot duration δ(t), the data collection completion time is minimized. The optimization problem is expressed as: In formula (10), T is the total number of time slots; constraint C1 indicates that data from all sensors must be collected; C2-C5 represent the ranges of δ(t), v(t), θ(t), and φ(t), respectively; and condition C6 represents the flight range of the UAV.

3. The method according to claim 1, characterized in that Step 2: Decompose the optimization problem P0 into two sub-problems, including: The first sub-problem optimization problem P1 is expressed as: Formula (11) is a channel allocation problem within t time slots. The goal is to maximize the amount of data collected by the UAV in a single time slot. Different channel allocation methods Ω(t) will affect the data acquisition speed, thereby affecting the length of a single time slot δ(t) and the total number of time slots T. The second sub-problem P2 is expressed as: All optimization variables in the second sub-problem P2 are continuous variables.

4. The method according to claim 1, wherein Step 3: The drone uses a trajectory planning neural network to determine the flight direction, speed, and duration of the drone's next time slot, including: Define the four-dimensional motion of the drone a t =[v(t),θ(t),φ(t),δ(t)] to represent the optimization variables of the second subproblem P2, and a t The conditions C2-C5 in the second sub-problem P2 need to be met; when performing action a t The position of the drone is: The action selection method adopts the random strategy method, that is, the neural network does not directly output a specific action, but outputs a parameter of the probability distribution related to the action, and generates a specific probability distribution based on the parameter for sampling, thereby obtaining a specific action; the β distribution is used as the random strategy distribution; the sampling range of the β distribution is [0,1], and the probability density function of the β distribution is: In formula (14), a d >0 and β d >0 are the two determining parameters of the β distribution, Γ(·) is the Γ function; the neural network is used to output four independent sets of a d , β d , then according to a d , β d Generate four independent beta distributions and sample the distributions to obtain the drone action a t .

5. The method according to claim 1, wherein Step 4: Obtain the channel gain of the current UAV position sensor and select the top M sensors for data collection. Use the channel allocation neural network to allocate channels to the M sensors, including: It is necessary to allocate channels to M sensors, and the state space of M steps State={state1,...,state m ,...,state M }, each step allocates a channel to a sensor, and the corresponding channel gain is set up represents the channel allocation status of the mth step, Indicates that the sensor is assigned to the mth channel of the lth channel; when all channels are assigned, SE M Expressed as: In formula (15), SE M The first U in the first row l The non-zero elements are Ω l (t), so according to SE M Get the corresponding Ω(t); The goal of the first sub-problem P1 is to maximize the amount of data collected in the current time slot, so the current time slot length δ(t) and the remaining data collected by the sensor are Also added to state m In the state space, it is represented as: In the channel allocation algorithm, the m-th step action represents the channel gain The sensor is assigned to the action m channel, where Use neural network to obtain channel allocation action; transform the state space state m Divided into three groups: [H s ,Rd], and SE m ; Among them, [H s ,Rd] represents the static information of the sensor, which remains unchanged during the channel allocation process; Indicates the information of the sensor that performs channel allocation in step m; at the same time, SE m Indicates the overall change of channels during the channel allocation process; First, CNN is used to extract the s ,Rd] and Extract features and get ref1 and q1 respectively; use the attention mechanism to match the two features to get Logit, and then use CNN to match SE m The rows and columns are feature extracted to obtain ref2 and ref3; ref2 is multiplied by Logits to obtain q2, and then the attention mechanism is used to match ref3 with q2 to obtain the probability p of channel assignment. m ; sampling p m Produce a specific action m ; The attention mechanism is specifically manifested as follows: In formula (17) k∈{1,2}, W qk is the fully connected layer parameter, V k is a learnable parameter and d is the number of hidden units of CNN.

6. The method according to claim 1, characterized in that Step 5: Update the channel assignment neural network weights using the trajectory reference mechanism and proximal gradient optimization algorithm, including: The reward of the channel assignment algorithm is set to the data collected by the UAV from each sensor, expressed as: Use Rollout Baseline and Proximal Policy Optimization (PPO) algorithm to train neural networks; According to the PPO algorithm, the loss function of Actor is expressed as: In formula (19), represents the similarity between the new and old channel allocation strategies, entropy is the entropy of the current strategy, and ξ is the entropy coefficient. According to the Rollout Baseline, the advantage function adv is expressed as: In formula (20), reward m A and reward M C Represents the rewards obtained by Actor and Critic respectively; when the mean of adv is greater than 0 and the p-value of the one-sided t-test of adv is greater than 0.005, the Actor parameter is assigned to Critic.

7. The method according to claim 1, characterized in that Step 7: The drone returns to the starting point to unload experience and uses the SAC algorithm to update the neural network weights, including: The state space includes the current position Q of the drone u (t), the position of the sensor {SP1,...,SP N } and the remaining data volume Rd of the sensor, the current state is expressed as: s t =[Q u (t),SP1,SP2,...,SP N ,Rd]. (21) The reward function in trajectory planning is designed as follows: In formula (22), μ is the execution t The penalty for failing to satisfy the C6 constraint prevents the drone from flying out of the boundary. t is the reward after the collection task is completed, F t Expressed as: Formula (23)F t It represents the time given to G minus the time consumed to complete the task collection, and ρ is the weight coefficient of time; The SAC algorithm has three modules, namely the Actor module, the Critic module, and the Critic-Target module. The multi-layer perceptron (MLP) is used to construct the modules of the SAC algorithm. The Actor module is used to output the actions of the drone, while the Critic module and the Critic-Target module are used to evaluate the quality of the Actor module’s actions. The Critic module and the Critic-Target module use clipped double-Q to avoid the problem of overestimation, that is, the Critic module and the Critic-Target module each have two evaluation networks. Let the neural network parameters of the Critic module and the Critic-Target module be θ respectively. i and Where i∈{0,1} represents the i-th evaluation network; the loss function of the Critic module is: In formula (24) In formula (25), Q tar represents the learning objective of the action-value function; It is in t Execute a in the state t The value of the action is given by the parameter θ i Neural network evaluation, γ is the discount factor, logπ φ (a t+1 |s t+1 ) is the entropy of the current policy, α H is the entropy weight coefficient; the parameters updated by Critic-Target using softupdate are as follows: In formula (26), τ∈[0,1] is the weight of softupdate; the loss function of the Actor module is expressed as: α H ∈[0,1]; then α H The loss function is expressed as: The distributed Apex architecture is used to implement the SAC algorithm. The architecture divides the SAC algorithm into four parts, namely the experience collection module rolloutworker, the experience cache module replaybuffer parameter server module parameter server and the SAC training module SAC trainer; When the drone starts from the starting point, the rollout worker perceives the state of the environment and takes action a t , and save the experience t ,a t ,r t ,s t+1 }; After data collection is completed, the drone will return to the starting point and unload the experience into the experience cache module; then, the SAC training module performs neural network training by sampling experience from the experience cache module and uploads the trained neural network parameters to the parameter server module; when the drone performs the next collection mission, it downloads the latest parameters from the parameter server module.