Unmanned aerial vehicle trajectory planning method based on noise dual-depth Q network

By optimizing UAV trajectory planning through a noisy dual-depth Q-network, the problems of insufficient adaptability and exploration in existing technologies are solved, achieving a higher trajectory planning success rate and data collection efficiency.

CN120685088APending Publication Date: 2025-09-23CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510787624.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing technologies, UAV trajectory planning based on deep Q-network has limited adaptability in high-dimensional continuous space, and the ε-greedy strategy is insufficiently explored in complex or dynamic environments, resulting in unstable trajectory planning.

Method used

A noisy dual-depth Q-network is adopted, combined with a noisy neural network and periodic delayed synchronous update. Noise bias terms and non-uniform sampling are introduced. A composite loss function and trajectory smoothness constraints are designed to optimize UAV path planning.

Benefits of technology

It improves the success rate and stability of drone trajectory planning, enhances the exploration capability in different environments, and ensures path continuity and data collection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120685088A_ABST
    Figure CN120685088A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent unmanned aerial vehicle autonomous control and decision making, in particular to an unmanned aerial vehicle trajectory planning method based on a noise dual-depth Q network, which comprises the following steps: acquiring unmanned aerial vehicle sensor data; establishing a forest fire model according to unmanned aerial vehicle sensor data; establishing an unmanned aerial vehicle constraint condition and an optimization function according to the forest fire model, and converting the unmanned aerial vehicle constraint condition and the optimization function into a partially observable Markov decision process; solving the partial observable Markov decision process by adopting a noise double-depth Q network to obtain an optimal unmanned aerial vehicle path planning strategy; according to the invention, the node areas are taken as decision objects, and value grading is carried out on different node areas, so that the decision is more in line with the reality; a noise bias item is introduced into a neural network weight to balance learning stability and exploration performance, so that the network can dynamically adapt to exploration behaviors in different environments; and a trajectory smoothness constraint is designed to ensure that the flight path is continuous, and the success rate of unmanned aerial vehicle trajectory planning is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous control and decision-making of intelligent unmanned aerial vehicles (UAVs), and in particular to a UAV trajectory planning method based on a noisy dual-depth Q network. Background Art

[0002] In forest fires, drones are often used for data collection tasks, and drones need to perform trajectory planning when performing data collection tasks.

[0003] Existing technologies often combine deep Q networks for drone path planning. For example, Chinese invention patent publication number CN120010547A discloses a method for collecting drone data in forest fire scenarios, which uses dual deep Q networks for path planning.

[0004] The dual-depth Q network in this paper involves a Dueling structure. When processing high-dimensional states or continuous action spaces, this structure is limited by the generalization ability of the advantage function estimation, and its adaptability in high-dimensional continuous spaces will be restricted. In addition, the dual-depth Q network in this paper adopts an ε-greedy strategy for action selection. This strategy dynamically adjusts the exploration probability ε through a decay curve. As the number of iterations increases, ε decays exponentially or linearly, gradually focusing on the known optimal strategy. The exploration ability is affected by the design of the ε decay curve, and insufficient exploration is prone to occur in complex or dynamic environments, leading to unstable drone trajectory planning and cruising. Summary of the Invention

[0005] In view of this, the present invention discloses a UAV trajectory planning method based on a noisy dual-depth Q network to solve the above problems; the method comprises:

[0006] S1, obtain drone sensor data;

[0007] S2, establish a forest fire model based on drone sensor data;

[0008] S3. Establish UAV constraints and optimization functions based on the forest fire model, and convert the UAV constraints and optimization functions into a partially observable Markov decision process;

[0009] S4, using a noisy dual-depth Q-network to solve a partially observable Markov decision process and obtain the optimal UAV path planning strategy;

[0010] Furthermore, the drone sensor data includes: weather information, drone transmission and reception frequency information, drone motion information, and environmental information;

[0011] Furthermore, the forest fire model includes:

[0012] A forest fire spread model, which uses cellular automata to delineate fire spread areas based on meteorological information;

[0013] Fire environment model, used to describe wind speed changes, smoke diffusion, and terrain undulations based on environmental information;

[0014] The UAV motion model is used to describe the UAV state based on the UAV motion information;

[0015] The communication scheduling model is used to describe the UAV communication rate and communication scheduling strategy based on the UAV's transmission and reception frequency information;

[0016] Regional data value model, used to quantify the data value of the node area;

[0017] Furthermore, the noisy dual deep Q-network is based on a parameterized noise mechanism. Compared with the deep Q-network, a noise neural network is introduced as the target network to learn the noise distribution.

[0018] Furthermore, the noisy dual-depth Q network adopts periodic delayed synchronous update;

[0019] Furthermore, when the noisy dual-depth Q network performs experience replay, non-uniform sampling is adopted and regularization constraints are combined to calculate the priority sampling weights.

[0020] The beneficial effects of the present invention include:

[0021] By adopting node areas as the basic collection objects, the decision-making is more in line with the reality;

[0022] By classifying node areas according to data value, quantifying the data value of forest fire areas, and introducing corresponding priority sampling weights in subsequent processing, the UAV path planning decision-making is more reasonable and has a higher success rate than other path planning methods under the same conditions.

[0023] By introducing a noise bias term into the neural network weights, we further balance the stability of learning and exploration performance, enabling the network to dynamically adapt to exploration behavior in different environments.

[0024] A composite loss function consisting of the Huber loss function, TD error priority sampling weights, and the Frobenius norm regularization term is designed. The sample weights are adjusted based on the TD error value of each experience pool sample and the priority of the corresponding experience. The Frobenius norm regularization term suppresses overfitting and achieves a balanced contribution of different samples to learning.

[0025] By designing trajectory smoothness constraints, the continuity of the flight path is ensured, further improving the success rate of UAV trajectory planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematic diagram of the UAV trajectory planning method based on the noisy dual-depth Q network in the present invention;

[0027] Figure 2 is a grayscale image of fire spread at the 20th time step of the cellular automaton in an embodiment of the present invention;

[0028] Figure 3 is a grayscale image of fire spread at the 100th time step of the cellular automaton in an embodiment of the present invention;

[0029] Figure 4 is a grayscale image of fire spread at the 200th time step of the cellular automaton in an embodiment of the present invention;

[0030] Figure 5 is a grayscale image of fire spread at the 300th time step of the cellular automaton in an embodiment of the present invention;

[0031] Figure 6 Schematic diagram of data transmission between a firefighter area and a drone in an embodiment of the present invention;

[0032] Figure 7 Schematic diagram of data transmission between a facility area and a drone in an embodiment of the present invention;

[0033] Figure 8 Schematic diagram of data transmission between other areas and drones in an embodiment of the present invention;

[0034] Figure 9 Schematic diagram comparing the total reward values ​​of the noisy dual-depth Q network and the comparison algorithm in an embodiment of the present invention;

[0035] Figure 10 Schematic diagram comparing the total data collection rate of the noisy dual-depth Q network and the comparison algorithm in an embodiment of the present invention;

[0036] Figure 11 Schematic diagram comparing the high-value data collection rate of the noisy dual-depth Q network and the comparison algorithm in an embodiment of the present invention;

[0037] Figure 12 Schematic diagram comparing the trajectory planning success rate of the noisy dual-depth Q network and the comparison algorithm in an embodiment of the present invention;

[0038] Figure 13 This is a comparison chart of the collection rates of the noisy dual-depth Q network for data of different value levels in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the objectives, technical solutions, features and advantages of the present invention more clearly understood, the present invention is further described below with reference to the accompanying drawings and embodiments.

[0040] This embodiment includes a UAV trajectory planning method based on a noisy dual-depth Q network, such as Figure 1 As shown, including:

[0041] S1. Obtain drone sensor data.

[0042] Specifically, drone sensor data includes: weather information, drone transmission and reception frequency information, drone motion information, and environmental information.

[0043] S2. Establish a forest fire model based on drone sensor data.

[0044] The forest fire model includes: forest fire spread model, fire environment model, drone movement model, communication scheduling model, and regional data value model. Specifically:

[0045] The forest fire spread model is used to use cellular automata to delineate the fire spread area based on meteorological information, and obtain the failure state of the node area according to the fire spread area. The cellular automaton divides the forest area into cells of different states according to the burning situation, and represents the cells where the forest fire has spread as a set Determine whether the node's coordinates belong to If yes, the node area is considered as a failed node area; if no, the node area is considered as a valid node area. When calculating failed nodes, the cellular automaton uses the modified forest fire spread rate of Wang Zhengfei's model. The calculation of the modified forest fire spread rate includes:

[0046] Calculate the initial rate of forest fire spread using the formula:

[0047] R F0 =a*T F +b*V F +c*(100-H F )+d

[0048] Wherein, a, b, c, and d represent empirical parameters, which are preferably 0.03, 0.05, 0.01, and -0.3, respectively, in this embodiment. F Indicates temperature, V F Indicates wind force level, H F Indicates relative humidity of the air. Temperature, wind force level and relative humidity of the air are included in meteorological information.

[0049] The initial spread rate is corrected according to the fuel type, terrain slope and wind vector to obtain the corrected forest fire spread rate.

[0050] R=R F0 *K f *K s *K w

[0051] Among them, R represents the modified forest fire spread rate, K f Indicates the correction coefficient of the combustible type, which is determined by the combustible type, K s Indicates the influence factor of slope on forest fire spread rate, K w Indicates the factors that influence the spread rate due to wind force level and wind direction conditions.

[0052] Furthermore, if Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 , which is a schematic diagram of using cellular automata to delineate fire spread areas at different time steps in this embodiment.

[0053] The fire environment model is used to describe wind speed changes, smoke diffusion, and terrain fluctuations based on environmental information. Specifically:

[0054] The description of wind speed change includes: setting the wind speed uncertainty factor, calculating the uncertainty of wind speed based on the wind speed uncertainty factor, using the uncertainty of wind speed to dynamically correct the flight speed of the UAV, and calculating the ambient wind speed v based on the wind speed data and the uncertainty of wind speed. w [n]:

[0055]

[0056] Where n represents the time slot, v w [n] represents the ambient wind speed, Represents wind speed data, Δv w [n] represents the uncertainty of wind speed, and α represents the wind speed uncertainty factor, which is preferably 0.5 to 0.7 in this embodiment.

[0057] The description of smoke diffusion includes: setting the smoke diffusion speed correction coefficient, and calculating the smoke diffusion speed based on the smoke diffusion speed correction coefficient and the ambient wind speed. The formula used is:

[0058] R s [n] = μv w [n]

[0059] Among them, R s [n] represents the smoke diffusion speed, and μ represents the smoke diffusion speed correction coefficient, which is preferably 0.8 in this embodiment.

[0060] Furthermore, the smoke diffuses linearly according to the wind speed with the fire spreading area as the origin, and the starting time of the node area covered by the smoke is recorded.

[0061] The description of terrain undulation includes: constructing a mountain and terrain undulation model through coordinate mapping and exponential decay. Specifically, the altitude of the drone is calculated based on the coordinates, terrain correction parameters, the coordinates of the highest point of the mountain, and the attenuation of the mountain height. The formula used is:

[0062]

[0063] Among them, z(x,y) represents the altitude of the drone, (x,y) represents the coordinates of the drone, (x i ,y i ) represents the coordinates of the highest point of the i-th mountain, h i Represents the terrain correction parameter, which is used to control the overall height of the i-th mountain, x si represents the attenuation of the height of the i-th mountain along the x-axis, y si represents the attenuation of the height of the i-th mountain along the y-axis, x si with y si Used to describe the slope of a mountain.

[0064] The UAV motion model is used to describe the UAV state based on its motion information. This includes setting a discount factor for wind speed effects and calculating the UAV's ground speed based on the discount factor, ambient wind speed, and the UAV's ground speed in a windless environment. The formula used is:

[0065] v[n]=v u [n]+λv w [n]

[0066] Among them, v[n] represents the UAV’s ground speed, v u [n] represents the ground speed of the UAV in a windless environment, and λ represents the discount factor for the influence of wind speed, which is preferably 0.8 in this embodiment.

[0067] Furthermore, based on the description of the terrain undulation, the obstacle area in the UAV flight altitude plane is represented as a set Obstacle information is the terrain information of the mountain, which is described by the fire environment model. The formula is:

[0068]

[0069] According to the position coordinates q of the UAV at time slot n u [n] and the j-th obstacle area D j The relationship between z(x,y)≥h u , to determine whether the drone is in obstacle avoidance state. If satisfied, then Determine whether the drone is in obstacle avoidance state; if not, then q u [n]∈D j, it is judged that the drone is in non-obstacle avoidance state.

[0070] The communication scheduling model is used to describe the UAV communication rate and communication scheduling strategy based on the UAV's transmission and reception frequency information, including:

[0071] The dynamic path loss between the UAV and the ground wireless sensor node is calculated based on the carrier frequency. The formula used is:

[0072] PL k [n] = 20log 10 f c +10n1log 10 θ k [n]+10n2log 10 d k [n]+X σ

[0073]

[0074] Among them, PL k [n] represents the dynamic path loss, k represents the node area, d k [n] represents the distance between the UAV and the node area k, f c represents the carrier frequency, θ k [n] represents the elevation angle of the center point of the ground node area relative to the drone, X σ represents shadow fading, which conforms to Gaussian distribution, n1 represents elevation angle influence factor, n2 represents path loss index, d k [n] represents the distance between the UAV and the boundary of the node area k, (x[n], y[n], h u ) represents the coordinates of the drone, (x k ,y k ,h k ) represents the coordinates of the node area, r k Indicates the radius of the node area; in this embodiment, the node is selected as the center, r k The unit of radius is used as the node area, and the unit area can also be other shapes.

[0075] The communication rate of the UAV is calculated based on the dynamic path loss, and the formula used is:

[0076]

[0077] γ k [n] = p0 - PL k [n]-p n

[0078] Among them, R k [n] represents the UAV communication rate, B represents the channel bandwidth, γk [n] represents the signal-to-noise ratio of the drone receiving end, p0 represents the node transmission power, and p n represents the noise power at the UAV receiver.

[0079] Set the signal-to-noise ratio threshold γ0, and judge the connection status between the drone and the node based on the signal-to-noise ratio threshold and the signal-to-noise ratio of the drone receiving end.

[0080] Specifically, γ k [n]<γ0, then the node region k is a disconnected node region; γ k [n] ≥ γ0, then the node area k is the connection area. Furthermore, the UAV establishes a communication connection with at most one node in each time slot. In order to improve the efficiency of data collection, the UAV prioritizes data transmission with the node area with good communication channel conditions.

[0081] Step 4: Calculate the communication scheduling strategy based on the node area connection status and the signal-to-noise ratio of the drone receiving end. The formula used is:

[0082]

[0083] in, Indicates the communication scheduling strategy, otherwise includes disconnected areas and failed node areas.

[0084] The regional data value model is used to quantify the data value of the node area. The present invention distinguishes and divides the node area according to the actual situation of the importance of data collection. In this embodiment, the divided node areas include: firefighter area, facility area, and other areas. The firefighter area is as follows: Figure 6 As shown, the facility area is as follows Figure 7 Other areas such as Figure 8 As shown. Calculate the total data value of node areas with different data importance respectively; the formula used to calculate the total data value is:

[0085] v k [n] = (1-α)v k,a +αv k,b [n]

[0086] Among them, v k [n] represents the total data value, v k,a Indicates the data distance value, v k,b [n] represents the data security value, α represents the data security value weight of the pre-set node area, and the formula used to calculate the data security value and data distance value is:

[0087]

[0088] Among them, G represents the number of fire areas, dg [k] represents the shortest distance from the edge of node area k to the edge of fire area g, η represents the preset smoke hazard coefficient, t n =n-n0 represents the duration of the node area being covered by smoke until time n, n0 represents the start time of the node area covered by smoke, (x g ,y g , h g ) represents the coordinates of the fire area, (x k ,y k , h k ) represents the node area coordinates, r k Represents the node area radius, r g Indicates the radius of the fire area.

[0089] The regional data value model classifies the importance according to the data distance value, taking into account the boundary distance between the target area and the fire area. The regional data value model with priority distinction is more in line with the actual situation and has practical application value.

[0090] S3. Establish UAV constraints and optimization functions based on the forest fire model, and convert the UAV constraints and optimization functions into a partially observable Markov decision process (POMDP).

[0091] Specifically, the UAV constraints include: UAV speed performance constraints, UAV steering performance constraints, environmental wind constraints, UAV obstacle avoidance constraints, communication scheduling constraints, single-node area data transmission constraints, and mission time constraints.

[0092] In this embodiment, the drone constraint condition adopts the formula:

[0093]

[0094] C6:Q k ≤Q max

[0095] C7:T≤T max

[0096] Among them, C1 is the UAV speed performance constraint, which limits the UAV flight speed to no more than the maximum speed v allowed by the performance. max ; C2 is the steering performance constraint of the UAV, which limits the heading angle change to no more than the maximum heading angle change C3 is the environmental wind constraint, which limits the wind speed change vector to no more than the maximum allowable amplitude of wind speed change. C4 is the obstacle avoidance constraint of the UAV, which limits the three-dimensional coordinates q of the UAV u[n] The extended area D that is not at the jth obstacle j C5 is the communication scheduling constraint, which limits the communication scheduling strategy in the optimization objective. C6 is the single-node area data transmission constraint, which limits the amount of data Q that the drone can send to a single node area k. k The upper limit Q max ; C7 is the task time constraint, which limits the total task time T to less than the maximum time T max .

[0097] Among them, the drone collects a total amount of data Q from the node area k k Calculated based on the communication rate between the UAV and the node area k, the communication scheduling strategy, and the indicator function δ[n]; the formula used is:

[0098]

[0099] The indicator function δ[n] takes the value of 0 or 1, where a value of 1 indicates that data has been collected and a value of 0 indicates that no data has been collected. The communication scheduling model comprehensively considers the quality of the communication link, dynamic resource scheduling, and task integrity constraints to ensure an efficient and reliable data transmission mechanism.

[0100] Establish an optimization function. In forest fires, drones need to plan safe trajectories in real time according to the state of the environment and node areas to collect the most valuable data. This paper transforms the trajectory planning problem of data collection drones in forest fire scenarios into maximizing the total value of data collected from all wireless sensor node areas. The strategy decoding is performed with the maximization of the total value of node area data as the objective function. According to the data value v of node area k in time slot n, the optimal strategy is established. k [n], the data rate R of the UAV from node k in time slot n k [n], indicator function δ[n] and communication scheduling variables Calculate the optimization function P1, indicating that the function is also equivalent to the duration of time slot n. If there is a time when the value of δ[n] is 1, it means that data is still being collected.

[0101]

[0102] Furthermore, a partially observable Markov decision process typically consists of a state space, an action space, a reward function, and a discount factor. The observation space S represents the set of all possible states, the action space A represents the set of all possible actions, the reward function R(s,a) represents the reward obtained by taking action a in state s, and the discount factor γ∈[0,1) is used to balance the importance of current rewards with future rewards.

[0103] Define the observation space, and concatenate the drone's state information vector, the drone's detection environment information vector, the global node area information vector, and the time consumed in executing the task to obtain the environment information vector S[n].

[0104] S[n]=(S u [n],S e [n],S s [n],S t )

[0105] S s [n]=(S k [n])for k∈{1,2,…,K}

[0106]

[0107] Among them, S u [n] represents the state information vector of the UAV, including the coordinate data, speed data and heading angle data of the UAV; S e [n] represents the information vector of the drone’s detection environment. The drone sensor obtains the distance d of the nearest obstacle in each horizontal direction. m [n], in this embodiment, the drone sensor can obtain distance information in 8 directions, S e [n]=(d1[n],d2[n],…,d m [n],…,d8[n]), the sensor is set with a maximum effective detection distance limit. In the absence of obstacles, the sensor reports the maximum effective detection distance value. Obtaining distance information can be achieved by detecting the surrounding environment through airborne sensors such as laser radar, binocular camera or infrared camera; S s [n] represents the global node region information vector, and the global node is composed of the information vector of a single node region; S k [n] represents the information vector of a single node area, S t Indicates the time consumed in executing the task.

[0108] Define the action space and calculate the executable actions according to the change in the speed and heading of the drone. Specifically, the actions of the drone are limited by the mechanical properties of the drone, and only executable actions can be sampled from legal actions. The change in the speed of the drone v c and heading change composition.

[0109] Define a reward function. Specifically, actions performed by the drone will change the state of the system environment, and the environment will provide feedback to the drone. Actions that contribute to achieving the goal are rewarded, while actions that do not contribute to achieving the goal are penalized. The drone needs to optimize its trajectory planning strategy based on environmental feedback. Therefore, the design of the reward function is based on the desired goal. Designing a reasonable reward function based on the task model and optimization problem can help the drone quickly learn the optimal strategy guided by the task goal. The reward function includes:

[0110] Drone Data Collection Reward R dc , data collection rewards are based on the data value v k [n] and the amount of new data in time slot n calculate.

[0111] Drone Time Saving Bonus R tk , the time saving reward is scaled by the remaining time of the task, and the remaining time is Calculate the drone time saving reward R tk .

[0112] Drones approaching unknown obstacles will be penalized R uo The obstacle approach penalty is based on the minimum distance d between the drone and the obstacle in all directions and the longest effective detection distance D of the drone's distance sensor. s calculate.

[0113] Drone collision penalty R so , the collision penalty is triggered when a path conflict occurs or a space constraint is violated, according to the position coordinate q of the UAV at time slot n u [n] Whether it belongs to the jth obstacle area D j calculate.

[0114] Drone invalid action penalty R ia ,Invalid action penalty is set as a fixed negative value for action repetition, failure or violation, and is calculated based on the invalid action penalty value.

[0115] Furthermore, the above reward functions are linearly weighted and combined to obtain a global reward function, which serves as the total reward of the drone at each decision step.

[0116] In this embodiment, the formula of the reward function is:

[0117] R=R bc +R tk +R bo +R so +R ia

[0118]

[0119] R ia =-β5

[0120] Among them, β is the scaling factor used to measure the weight, β1 represents the data collection amount scaling factor, β2 represents, β3 represents the obstacle avoidance penalty coefficient, β4 represents the collision penalty coefficient, and β5 represents the invalid action penalty coefficient.

[0121] S4. A noisy double-depth Q-network is used to solve the partially observable Markov decision process to obtain the optimal UAV path planning strategy; the noisy double-depth Q-network is based on a parameterized noise mechanism and introduces a noise neural network as the target network, which is used to learn the noise distribution.

[0122] Specifically, the standard deep Q network uses the same network for action selection and Q value estimation. Each time the max operation is estimated, the largest Q value in the current estimate is selected. Due to errors in network prediction, the errors sometimes cause the Q value of non-optimal actions to be overestimated, resulting in over-estimation bias. In addition, the traditional ∈-greedy strategy controls the degree of exploration through a pre-set exploration value. Since the selected actions are randomly generated, the direction of strategy learning will fluctuate greatly. Especially in some critical states, random exploration may cause the agent to deviate from the learning goal, while the controllable noise exploration of the noise network can better balance the stability of learning and the need for exploration. The noise network introduces a noise bias term into the neural network weights, so that the network dynamically adapts to the exploration behavior in different environments. The present invention adopts a noisy dual-depth Q network (NoisyNet-DDQN). The noisy dual-depth Q network uses the standard deep Q network as the estimation network and introduces a noise neural network as the target network. The noise neural network is used to learn the noise distribution. The estimation network and the target network together constitute the noisy dual-depth Q network. The formula of the noisy dual-depth Q network is:

[0123] Q(s,a;θ t )=Q(s,a;w,B)

[0124] Among them, Q(s,a;θ t ) represents the noisy dual-depth Q network, s represents the current state, a represents the action, W represents the weight parameter, w represents the noise weight, B represents the noise mean, w and B together constitute the learnable noise bias term θ of the noisy dual-depth Q network t .

[0125] w=W+Σ⊙ε W

[0126] B=b+σ⊙ε b

[0127] Among them, W represents the mean part of weight w, Σ∈R d×drepresents the noise covariance matrix, ⊙ represents element-wise multiplication, ε W represents the noise sample applied to the weight w, b represents the mean part of the bias B, and σ∈R d Represents the standard deviation of the bias B, that is, the bias noise vector, b represents the bias parameter, ε b represents the noise sample applied to the bias b, and the Gaussian random variable ε W ,ε b The noise parameters (Σ,σ) are sampled independently during each forward propagation and participate in the back-propagation training.

[0128] Specifically, the noise neural network uses the noise covariance matrix, bias noise vector and Gaussian random variables to perform periodic delayed synchronous updates on the deep Q network.

[0129] Periodic delay synchronization is used to enhance the stability of the training process. Specifically, a hard synchronization operation is performed after each fixed step K. The main weight parameters of the target network are calculated based on the main weight parameters of the estimated network. The noise covariance matrix of the target network is calculated based on the noise covariance matrix of the estimated network. The bias noise vector of the target network is calculated based on the bias noise vector of the estimated network. The update process is expressed as:

[0130]

[0131] Among them, θ t Represents the main weight parameters of the estimated network, represents the main weight parameters of the target network, Σ t represents the noise covariance matrix of the estimation network, represents the noise covariance matrix of the target network, σ t represents the bias noise vector of the estimation network, Represents the bias noise vector of the target network.

[0132] When the noise dual-depth Q network designed in the present invention performs experience playback, the priority sampling is calculated in combination with the regularization constraint; specifically, the formula used to calculate the priority sampling weight is:

[0133] w j =(N·P(j)) -β ,P(j)∝|y j -Q(s j ,a j )|

[0134] Among them, N represents the experience pool capacity, β represents the decay coefficient, P(j) represents the priority of experience j, and y j Indicates the target Q value, Q(s j ,a j ) means in state s jNext, perform action a j The Q value predicted by the estimated network.

[0135] Furthermore, since the priority experience replay uses non-uniform sampling, experience pool samples with large TD errors will be sampled more frequently, resulting in the excessive amplification of the contribution of these experience pool samples to the Q value update. In order to correct the sampling bias and ensure the performance of the unbiased estimation strategy, the present invention adopts a composite loss function composed of the Huber loss function, the TD error priority sampling weight and the Frobenius norm regularization term. The sample weight is adjusted according to the TD error value of each experience pool sample and the priority of the corresponding experience. The Frobenius norm regularization term suppresses overfitting and achieves a balance in the contribution of different samples to learning. The formula of the composite loss function designed by the present invention is:

[0136]

[0137] Among them, B represents the total amount of data of experience j, w j represents the data priority sampling weight of experience j, L δ represents the Huber loss function, Represents the Frobenius norm, and α represents the regularization coefficient of the Frobenius norm. The formula of the Huber loss function is:

[0138]

[0139] Here, δ represents a hyperparameter.

[0140] The noise dual-depth Q network designed in this invention also has a trajectory smoothness constraint to ensure the continuity of the flight path. The updated network parameters are calculated based on the original network parameters, the learning rate, the gradient of the loss function, the trajectory smoothing coefficient, the square of the second-order difference term, and the gradient of the trajectory smoothing term, which are accumulated over time steps. The formula used is:

[0141]

[0142] Among them, Q new represents the updated network parameters, Q represents the original network parameters, η represents the learning rate, Represents the gradient of the loss function, μ represents the trajectory smoothing coefficient, which is dynamically adjusted according to the maximum angular velocity of the drone. Indicates time step accumulation, P k represents the coordinates of the UAV trajectory points, ‖P k -2P k-1 +P k-2 ‖ 2 Indicates P k The second-order difference term of .

[0143] During trajectory planning, the drone perceives the grid environment in real time and generates an optimal path planning strategy that satisfies constraints based on a deep reinforcement learning framework. This strategy determines the next movement direction and calculation method. During the action execution phase, the drone moves according to the decision and completes the corresponding task, while simultaneously updating state information, including location, in real time. This feedback is used as input to the noisy dual-depth Q-network. During network training, the policy network parameters are continuously optimized through priority experience sampling and loss function calculation, gradually improving the quality of decisions.

[0144] Furthermore, the noise neural network processes data specifically including:

[0145] Step 1: Initialize the parameters of the noisy dual-depth Q network.

[0146] Step 2: Calculate the legal action space and the range of actions that can be executed.

[0147] Step 3: Set the round step index k to 1 and the time step index t to 1.

[0148] Step 4: Initialize the environment and adaptive noise training mechanism.

[0149] Step 5: Get the environment information vector.

[0150] Step 6: Based on the parameterized noise injection mechanism and priority sampling weights, select action a in the legal action space t [n].

[0151] Step 7: According to action a t [n] Perform state transfer, obtain the reward function value, update the environment information vector according to the reward function value, and store the generated experience pool samples into the experience pool.

[0152] Step 8: Determine whether the number of experience pool samples is less than the preset experience pool sample number threshold. If so, t=t+1, and return to step 6. If not, proceed to the next step.

[0153] Step 9: Calculate the Q value of all legal actions and the TD error of each experience pool sample.

[0154] Step 10: Based on the TD error, the weight correction parameter β, and the capacity N of the experience replay pool, calculate the priority sampling weight of the experience pool samples; according to the composite loss function Correct the TD error of each experience pool sample.

[0155] Step 11: Update the parameters θ of the target network every K steps - ←τθ+(1-τ)θ - .

[0156] Step 12: Based on the corrected TD error of each experience pool sample, the noise bias term of the Q network is updated according to the trajectory smoothness constraint.

[0157] Step 13: Determine whether k is less than the preset maximum round step index. If so, k = k + 1 and return to step 4. If not, the optimal UAV path planning strategy that meets the constraints is obtained.

[0158] In this embodiment, the settings of nodes and drones are shown in Table 1, and the network parameter settings of the noisy double-depth Q network are shown in Table 2.

[0159] Table 1: Parameter reference range

[0160]

[0161] Table 2: Network parameter settings for the noisy dual-depth Q network

[0162]

[0163]

[0164] Figure 9 、 Figure 10 、 Figure 11 、 Figure 12 The total data collection characteristics of the proposed method are compared with those of other methods, using the DDQN, Dueling QAN, and PER DDQN models. The proposed algorithm is the Noisy Dual Deep Q Network (DDQN) in the figure. As shown in the figure, with increasing training rounds, the total reward value obtained by the DDQN gradually approaches a maximum, stabilizing at around 920; the total data collection rate gradually approaches a maximum, stabilizing at nearly 100%; the high-value data collection rate gradually approaches a maximum, stabilizing at 90%; and the trajectory planning success rate gradually approaches a maximum, stabilizing at nearly 100%. After 2000 rounds, the total reward value, total data collection rate, high-value data collection rate, and trajectory planning success rate of the DDQN are more stable than those of the other algorithms. Furthermore, algorithm stability is crucial for trajectory planning in forest fire scenarios. As shown in the figure, after a limited number of training rounds, the DDQN achieves convergence, effectively improving the decision-making robustness of the drone trajectory planning strategy.

[0165] Figure 13 This figure compares the data collection rate characteristics of the noisy dual-DQ network for different value levels. The first, second, and third value levels correspond to the other area, the facility area, and the firefighter area, respectively. The third value level has the highest data collection rate, stabilizing at around 90%, demonstrating the importance of priority sampling weights for data collection rate.

[0166] Finally, it should be noted that the above only describes some embodiments of the present invention. For those skilled in the art, it is conceivable that various changes, modifications, substitutions and deformations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents, and the above-mentioned actions should be covered within the scope of protection of the present invention.

Claims

1. A UAV trajectory planning method based on noisy dual-depth Q network, applied to forest fire scenarios, characterized by: include: S1, obtain drone sensor data; S2, establish a forest fire model based on drone sensor data; S3. Establish UAV constraints and optimization functions based on the forest fire model, and convert the UAV constraints and optimization functions into a partially observable Markov decision process; S4. A noisy double-depth Q-network is used to solve the partially observable Markov decision process and obtain the optimal UAV path planning strategy; the noisy double-depth Q-network is based on a parameterized noise mechanism.

2. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 1 is characterized in that: The bushfire model includes: A forest fire spread model is used to delineate the fire spread area using cellular automata; Fire environment model, used to describe wind speed changes, smoke spread, and terrain undulations; UAV motion model, used to describe the UAV state; Communication scheduling model, used to describe the UAV communication rate and communication scheduling strategy; Regional data value model, used to quantify the data value of the node area.

3. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 2 is characterized in that: The description of smoke diffusion includes: setting the smoke diffusion speed correction coefficient, and calculating the smoke diffusion speed based on the smoke diffusion speed correction coefficient and the ambient wind speed. The formula used is: R s [n]=μv w [n] Among them, R s [n] represents the smoke diffusion speed, μ represents the smoke diffusion speed correction coefficient, v w [n] represents the ambient wind speed, which is obtained from the drone sensor data. The smoke diffuses linearly according to the wind speed, with the fire spread area as the origin. The start time of the node area covered by the smoke is recorded. The start time of the node area covered by the smoke is used in the regional data value model to quantify the data value of the node area.

4. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 2 is characterized in that: The regional data value model divides the node areas into different categories and calculates the total data value of different node areas separately. The formula used to calculate the total data value is: v k [n]=(1-α)v k,a +av k,b [n] Among them, v k [n] represents the total data value, v k,a Indicates the data distance value, v k,b [n] represents the data security value, α represents the data security value weight of the pre-set node area, G represents the number of fire areas, d g [k] represents the shortest distance from the edge of node area k to the edge of fire area g, η represents the preset smoke hazard coefficient, t n =n-n0 represents the duration of the node area being covered by smoke until time n, n0 represents the starting time of the node area covered by smoke, which is obtained from the description of smoke diffusion. (x g ,y g , h g ) represents the coordinates of the fire area, (x k ,y k , h k ) represents the node area coordinates, r k Represents the node area radius, r g Indicates the radius of the fire area.

5. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 1 is characterized in that: The parameterized noise mechanism includes: the noisy dual deep Q network uses the standard deep Q network as the estimation network and introduces a noise neural network as the target network. The noise neural network is used to learn the noise distribution. The estimation network and the target network together constitute the noisy dual deep Q network.

6. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 5 is characterized in that: The formula for the noisy dual deep Q network is: Q(s,a;θ t )=Q(s,a;w,B) Among them, Q(s,a;θ t ) represents the noisy dual-depth Q network, s represents the current state, a represents the action, W represents the weight parameter, w represents the noise weight, B represents the noise bias, w and B together constitute the learnable noise bias term θ of the noisy dual-depth Q network t .

7. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 6 is characterized in that: The calculation formulas for noise weight and noise bias are: w=W+Σ⊙ε W B=b+σ⊙ε b Where W represents the mean part of the noise weight w, ⊙ represents element-wise multiplication, Σ∈R d×d represents the noise covariance matrix, ε W represents the noise sample applied to the noise weight w, b represents the mean part of the noise bias B, σ∈R d Represents the standard deviation of the noise bias B, that is, the bias noise vector, b represents the bias parameter, ε b represents the noise sample applied to the bias b, and the Gaussian random variable ε W ,ε b The noise parameters (Σ,σ) are sampled independently during each forward propagation and participate in the back-propagation training.

8. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 5 is characterized in that: The noisy dual-depth Q network adopts periodic delayed synchronization update, including: performing a hard synchronization operation after each fixed step; calculating the main weight parameters of the target network based on the main weight parameters of the estimated network; calculating the noise covariance matrix of the target network based on the noise covariance matrix of the estimated network; and calculating the bias noise vector of the target network based on the bias noise vector of the estimated network.

9. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 5 is characterized in that: The noisy dual-depth Q network uses a composite loss function, and the formula of the composite loss function is: w j =(N·P(j)) -β ,P(j)∝|y j -Q(s j ,a j )| in, represents the composite loss function, B represents the total amount of data for experience j, and w j represents the data priority sampling weight of experience j, L δ represents the Huber loss function, represents the Frobenius norm, α represents the regularization coefficient of the Frobenius norm, N represents the experience pool capacity, β represents the decay coefficient, P(j) represents the priority of experience j, y j Indicates the target Q value, Q(s j ,a j ) means in state s j Next, perform action a j The Q value predicted by the estimated network.

10. The UAV trajectory planning method based on noisy dual-depth Q network according to claim 5 is characterized in that: The noisy dual-depth Q network uses trajectory smoothness constraints to update the Q parameters. The formula for the trajectory smoothness constraint is: Among them, Q new represents the updated network parameters, Q represents the original network parameters, η represents the learning rate, Represents the gradient of the loss function, μ represents the trajectory smoothing coefficient, which is dynamically adjusted according to the maximum angular velocity of the drone. Indicates time step accumulation, P k represents the coordinates of the UAV trajectory points, ‖P k -2P k-1 +P k-2 ‖ 2 Indicates P k The second-order difference term of .

Citation Information

Patent Citations

  • Unmanned aerial vehicle data collection method in forest fire scene

    CN120010547A

Cited By

  • Unmanned aerial vehicle slope inspection autonomous navigation path planning system fused with neural network

    CN121829569A

  • Unmanned aerial vehicle autonomous navigation path planning system for slope inspection based on fusion neural network

    CN121829569B

  • Unmanned aerial vehicle adaptive motion planning method and device for different environments

    CN122083954A