Mobile node-oriented obstacle avoidance unmanned aerial vehicle data collection method in forest fire
Through the deep reinforcement learning model, the problem of data transmission instability in traditional data collection methods when mobile nodes are frequently moved is solved, efficient and complete data collection is achieved, and the problem of dynamic obstacle avoidance of drones is solved.
Patent Information
- Application Number
- CN202510166088.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-14
AI Technical Summary
At forest fire scenes, traditional data collection methods are difficult to ensure stable data transmission, especially when mobile nodes move frequently, resulting in data loss or delay. At the same time, the problem of dynamic obstacle avoidance of drones increases challenges.
The deep reinforcement learning model is used to build optimization problems. By obtaining data from ground mobile nodes and drone nodes, building a system model, planning obstacle avoidance paths, deploying the intelligent body to the drone, giving action instructions based on real-time environmental data, and performing data collection tasks.
It realizes more efficient data collection and covers a larger range of data sources, greatly improving the integrity of the data, providing a solid and reliable data foundation for subsequent in-depth analysis and decision-making, and solving the dynamic obstacle avoidance problem between drones.
Smart Images

Figure CN120010516A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles and the Internet of Things, and in particular to an obstacle avoidance unmanned aerial vehicle data collection method for mobile nodes in forest fires. Background Art
[0002] At the scene of a forest fire, the safety of firefighters and the real-time changes in the spread of the fire are very important. Traditional data collection methods, such as multi-hop transmission through wired network connections, wireless network nodes, or reliance on fixed base station transmission, have many limitations. Wired connection methods are expensive to lay and cannot adapt to the dynamic changes of mobile nodes; multi-hop transmission of wireless network nodes consumes a lot of energy, and mobile nodes also cause frequent changes in network topology; fixed base stations are limited by coverage and are difficult to achieve full coverage in some remote areas or areas with complex terrain. In addition, when mobile nodes are in a state of frequent movement, traditional methods are difficult to ensure stable data transmission, resulting in data loss or delay.
[0003] The rapid development of crowd sensing and the Internet of Things has made it possible to deploy a large number of mobile sensors at the fire scene. These data collection nodes benefit from a certain degree of mobility and can survive for a long time in the fire scene, continuously collect continuously changing complete data, and provide a large amount of environmental information at the forest fire scene and the vital signs of firefighters, thus playing an irreplaceable and important role in ensuring the safety of personnel at the forest fire scene. The use of drones to assist in data collection tasks further solves the problem of data collection and transmission, avoiding the energy consumption and random degradation of transmission performance caused by traditional multi-hop transmission.
[0004] However, the mobility of nodes makes the relative positions of different nodes and the fire scene change during the data collection task, and the value of their data also changes accordingly. At the same time, the moving aerial nodes and other non-cooperative drones increase dynamic obstacles in the flight environment.
[0005] The different values of data from different nodes and the problem of dynamic obstacle avoidance of UAVs bring more challenges to the UAV-assisted data collection task. Summary of the invention
[0006] In view of this, the present invention discloses a method for collecting data of a mobile node-avoiding drone in a forest fire to solve the above-mentioned problem; the method comprises the following steps:
[0007] S1. Acquire the data of the ground mobile nodes and the data of the UAV nodes, and establish a system model according to the data of the ground mobile nodes and the data of the UAV nodes;
[0008] S2, construct optimization problems based on the system model;
[0009] S3. Solve the optimization problem using a deep reinforcement learning model; obtain an intelligent agent for planning an obstacle avoidance path based on the solution result;
[0010] S4. Deploy the intelligent agent to the drone. The drone obtains real-time environmental data. The intelligent agent gives the drone action instructions for the next moment based on the real-time environmental data. The drone performs the data collection task according to the action instructions.
[0011] The beneficial effects of the present invention include: by establishing a mobile node model, quantitatively analyzing the unit data value of the mobile node according to the spread of forest fires, achieving more efficient data collection, covering a wider range of data sources in a short time, greatly improving the integrity of the data, and providing a solid and reliable data foundation for subsequent in-depth analysis and decision-making; by establishing safe distance constraints between drones based on the state vector relationship between the task drone and the node drone and the non-cooperative drone, the dynamic obstacle avoidance problem between drones is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 Schematic diagram of the process of the obstacle avoidance drone data collection method of the present invention;
[0013] Figure 2 Schematic diagram of the system architecture of the obstacle avoidance drone data collection method of the present invention;
[0014] Figure 3 This is a comparison diagram of the reward value convergence of the agent during the training process under different algorithms in the embodiment of the present invention;
[0015] Figure 4 This is a comparison chart of the total data collection rate of the mission drone under different algorithms in the embodiment of the present invention;
[0016] Figure 5 This is a comparison chart of data collection rates of aerial nodes by task drones under different algorithms in an embodiment of the present invention;
[0017] Figure 6 This is a comparison chart of data collection rates of mission drones to ground mobile nodes under different algorithms in an embodiment of the present invention;
[0018] Figure 7 This is a comparison chart of data collection rates of high-value nodes by mission drones under different algorithms in an embodiment of the present invention;
[0019] Figure 8 This is a comparison chart of the success rate of safe trajectory planning for a mission drone under different algorithms in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution, characteristics and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0021] This embodiment includes a method for collecting data of obstacle-avoiding drones for mobile nodes in forest fires. The flowchart of the method is as follows: Figure 1 As shown, the steps include:
[0022] S1. Acquire the data of the ground mobile nodes and the data of the UAV nodes, and establish a system model based on the data of the ground mobile nodes and the data of the UAV nodes.
[0023] Furthermore, the system model includes a system architecture, a fire area model, a mobile node model, and a data collection task model.
[0024] Specifically, the system architecture is described as follows: Figure 2 As shown in the figure, the mission drone maintains a safe altitude flight, and the spatial position of the mission drone is expressed as q[n] = (x[n], y[n], h u ), where h u is the flight altitude of the drone. The projection coordinates of the drone on the flight plane are q u [n]=(x[n],y[n]); the ground speed of the drone is v[n]=(v x [n],v y [n]); the heading of the drone is Furthermore, the state vector of the mission drone is
[0025] Furthermore, there are K wireless sensor nodes in the system, which are divided into fast-moving nodes in the air and slow-moving nodes on the ground. The fast-moving nodes in the air include K nodes that perform data collection tasks at the fire scene. 1 UAV nodes, slow-moving ground nodes including K 2 A driverless car and K 3 firefighters. K = K 1 +K 2 +K 3 The number of drones performing other tasks is recorded as L. There is no direct dispatch communication between drones performing other tasks and the task drones in this system, so they will also bring obstacles to the flight of the task drones. When these drones approach the task drone, the task drone obtains the basic information of these non-cooperative drones through the onboard sensors. The state vectors of other non-cooperative drones at the fire scene are q l [n]=(x l ,y l ) represents the position of the non-cooperative UAV, v l[n] represents the speed of the non-cooperative drone, Indicates the heading of the non-cooperative UAV.
[0026] Specifically, the fire area model is described as follows: At the forest fire site, wind force will affect the spread of the forest fire and also affect the flight speed of the drone. w [n] is represented by in represents the data of the UAV wind speed sensor, n represents the nth time slot obtained by evenly dividing the maximum mission time, Δv w [n] represents the uncertainty of wind speed, which is determined by the difference between meteorological data and on-site wind speed. The larger the difference, the greater the uncertainty, the faster the wind changes, and the maximum uncertainty of wind speed. The larger the value, the greater the flying speed of the drone. u [n]+λv w [n], where v u [n] is the ground speed of the UAV in a windless environment, λ is the discount factor for the wind speed, and v u [n] is the ground speed of the UAV in a windless environment, λ is the discount factor for the wind speed, the preferred range is 0.005 to 0.01, and the preferred value in this embodiment is 0.01. Due to the performance of the UAV itself, there is a constraint ||v u [n]|| 2 ≤v max , and
[0027] Furthermore, the influence of mountains and vegetation on terrain is expressed as Where (x, y) represents the xy plane coordinates, z is the altitude of the location, h is the altitude of the location, i represents the overall height of the i-th mountain, x i With y i Control the slope of the mountain along the x-axis and y-axis respectively.
[0028] Furthermore, the UAV needs to avoid known terrain obstacles and unknown environmental obstacles during flight, so the area of various obstacles in the UAV flight altitude plane is represented as a set D = {D 1 ,D 2 ,…D j}, where D j represents the jth obstacle, D j That is, z(x,y)>h u is the set of points whose altitude is higher than the flight altitude of the UAV; the obstacle avoidance constraint of the UAV is The drone detects environmental information through the airborne distance sensor, and the surrounding environmental information vector obtained is S o [n]=(d 1 [n],d 2 [n],…,d 8 [n]), where d i [n],i∈{1,2,…,8} represents the distance between the UAV and the obstacle detected by the sensor in a certain direction. The sensor has a maximum detection distance limit.
[0029] Furthermore, cellular automata are used to describe the forest fire spread process. The cellular automata divide the task area into closely adjacent square grids, each of which is a cell. The properties of the cell are determined by factors such as the spatial terrain, vegetation, meteorological conditions, ignition time and combustion state inside the cell. The state of the cell can be divided into five stages: unburned, burned, completely burned, unextinguished and completely extinguished, which are represented by F 0 ,F 1 ,F 2 ,F 3 ,F 4 The initial rate of forest fire spread R F0 The formula is:
[0030] R F0 =a*T F +b*V F +c*(100-H F )+d
[0031] Among them, T F is the temperature, V F is the wind force level, H F is the relative humidity, a, b, c, and d are all empirical parameters, and their preferred values are 0.03, 0.05, 0.01, and -0.3, respectively. Further considering the influence of fuel type, terrain slope, and wind vector on the initial spread rate, the corrected forest fire spread rate is R = R F0 *K f *K s *K w , where R is the corrected forest fire spread rate; K f is the correction factor for combustible materials; K s is the terrain slope correction coefficient, K s Describes the effect of different slopes on the spread rate of forest fire; K w is the wind vector correction coefficient, K w The effects of different wind force levels and wind direction conditions on the spread rate are described.
[0032] Furthermore, the coordinate index of each cell is defined as u and v, and the forest fire spread rate of the cell (u, v) is defined as The distance between a cell and its adjacent cells is the side length L of the cell. The state transition rule of the cell is: when an unburned cell is ignited, F 0 Change to F 1 ; F 1 Internal combustion time Then converted to F 2 , ignite the F in the direction parallel to the fire line within Δn time 0 , where R IN is the rate of forest fire spread inside the cell, L is the side length of the cell; F 2 Ignites all surrounding F 0 , when F 2 There is no F around the cell of the state 0 Convert to F 3 , time consuming Δn; F 3 After Δn time, enter F 4 Status. Collection Indicates the area where the forest fire has spread. 1 And the cells of the subsequent state, if the coordinate q of node k k exist , then node k is a failed node.
[0033] Specifically, the mobile node model is explained as follows: Due to different tasks, different mobile nodes have different mobility characteristics. Firefighters and unmanned vehicles as ground nodes have a smaller range of movement during the drone data collection mission. The unmanned vehicle is close to the fire scene, while the firefighters keep a distance from the fire scene. UAVs as aerial nodes have a larger range of activities. Depending on the specific tasks, node drones may have different speeds and directions.
[0034] Furthermore, the ground nodes (including K 2 A driverless car and K 3 The spatial position of the firefighter node) is Since the moving speed of the ground node is relatively small compared to the flight speed of the mission UAV, it is considered as a short-distance slow linear motion; the ground node speed is Direction K 1 The spatial position of the node drone is Speed Heading The node drone used to monitor the fire scene moves along with the changes in the fire line, that is, the speed of the node drone will change according to the mission requirements based on the forest fire spread speed R in the fire area model; the node drone used to monitor a key area adopts a fixed trajectory to cruise over the area, and the cruising trajectory of the node drone is different depending on the size of the area.
[0035] Furthermore, in time slot n, the main ignition points A and B closest to a node are located at (x a ,y a ) and (x b ,y b ), and y b >y a , the edge direction of the fire line formed by A and B is Indicator of the relative position change between the node and the forest fire
[0036] Furthermore, we define the unit data value p of a node k [n] are as follows:
[0037]
[0038] Among them, p k [n] represents the unit data value of node k in the nth time slot, β is the node motion trend attention coefficient, ranging from 1 to 2, preferably 1 in this embodiment, γ is the life safety coefficient of firefighters, γ is 0 when node k is a drone node or an unmanned vehicle node, and is 2 when the node is a firefighter node.
[0039] Furthermore, the node drone may adjust its working altitude at any time according to the needs of the mission. The flight altitude of the node drone is close to h u This brings new safety threats to the flight of the mission UAV. At the same time, considering that there are L non-cooperative UAVs in the same space, the safety distance constraint d between the mission UAV and other UAVs uk [n] is represented by:
[0040] d uk [n]=(x[n]-x k [n]) 2 +(y[n]-y k [n]) 2 +(h u -h k [n]) 2 ≥D u ,
[0041]
[0042] Among them, Du The calculation formula is:
[0043]
[0044] Among them, l mm Indicates the maximum size of the mission drone, l om Indicates the maximum size of other drones, d safe Represents the safety redundancy distance, d safe Preferably it is 1.5m.
[0045] Specifically, the data collection task model is described as follows: The data collection task model is used to establish communication constraints, which include scheduling constraints and data transmission constraints. In the forest fire scenario, the channel between the UAV and the ground mobile node is mainly affected by vegetation. The path loss model PL used in the present invention k The formula for [n] is:
[0046] PL k [n] = 20log 10 (f c )+10n 1 log 10 (θ k [n])+10n 2 log 10 (d k [n])+X σ
[0047]
[0048] Among them, f c represents the carrier frequency, θ k [n] represents the elevation angle of the drone relative to the ground node; X σ Indicates shadow fading, which conforms to Gaussian distribution; d k [n] represents the distance between the drone and node k; the spatial position of the drone is (x[n], y[n], h u ); the position of the kth node is (x k [n],y k [n],h k [n]); n 1 and n 2 They are elevation angle influence factor and path loss exponent respectively.
[0049] Furthermore, the node transmission power is defined as p 0 , the noise power at the receiving end of the drone is p n , the signal-to-noise ratio γ at the drone receiver n [n] is:
[0050] γn [n] = p 0 -PL k [n]-p n
[0051] Furthermore, the communication rate between the UAV and the ground mobile node for:
[0052]
[0053] Wherein, B represents the channel bandwidth.
[0054] The communication channel between the mission UAV and the UAV node can be regarded as line-of-sight propagation, and the channel gain is Then the communication rate between the mission UAV and the node UAV based on free space propagation loss is where β 0 is the channel gain per unit distance.
[0055] Furthermore, in order to support stable data stream transmission, the receiving signal-to-noise ratio γ at the drone should be k [n]≥γ 0 The data collection task starts when γ 0 The signal-to-noise ratio threshold is set manually, preferably in the range of -5dB to -10dB. In this embodiment, -5dB is preferred. The mission drone collects data from at most one node at a time. represents the connection between the UAV and the kth node, and the scheduling constraint is expressed as:
[0056] Furthermore, the UAV will prioritize data transmission tasks with nodes with good and stable communication channels. The formula for the scheduling strategy is:
[0057]
[0058] Among them, active nodes are represented by the signal-to-noise ratio γ k [n]≥γ 0 , otherwise indicates that the signal-to-noise ratio is lower than the threshold γ 0 nodes and failed nodes.
[0059] Furthermore, the data transmission constraint obtained by the amount of data collected by the UAV from node k is expressed as: represents the total amount of data collected from node k at time slot n. The superscript L is used to distinguish and Q K , the total amount of data that can be collected by the drone during the data collection process of all nodes is The maximum mission time is evenly discretized into N time slots, so that the position of the drone in each time slot can be approximately regarded as unchanged. The time slot length of the nth time slot is δ[n] = T max / N, in this embodiment, the number of time slots N is preferably 300.
[0060] S2. Construct optimization problems based on the system model.
[0061] Specifically, the optimization problem is expressed as:
[0062]
[0063] subject to
[0064] c1:
[0065] C2:
[0066] C3:
[0067] C4:
[0068] C5:d uk [n]=(x[n]-x k [n]) 2 +(y[n]-y k [n]) 2 +(h u -h k [n]) 2 ≥D u ,
[0069]
[0070] C6:
[0071] C7:
[0072] C8:T≤T max
[0073] Among them, C1 is the UAV flight speed performance constraint, C2 is the UAV steering performance constraint, C3 is the uncertainty constraint of environmental wind force, C4 is the obstacle avoidance constraint, C5 is the safety distance constraint between UAVs, C6 is the scheduling constraint, C7 is the data transmission constraint, and C8 is the task time constraint.
[0074] S3. Use a deep reinforcement learning model to solve the optimization problem; and obtain an intelligent agent for planning an obstacle avoidance path based on the solution results.
[0075] Furthermore, the sequential decision-making problem is expressed by a Markov decision process, that is, formulated, using a tuple to represent, and the specific meaning of the decision-making process is as follows:
[0076] represents the state space, and the environmental information vector S[n] = (S u [n], S o [n], S s [n], S c [n], S t ); where S u [n] is the state information vector of the UAV, v[n] represents the flight speed of the UAV, and the formula is v[n] = v u [n] + λv w [n], v u [n] is the ground speed of the UAV in a windless environment, λ is the discount factor affected by the wind speed, and the preferred range is 0.005 to 0.01. In this embodiment, it is preferably 0.01; S o [n] = (d 1 [n], d 2 [n], …, d 8 [n]), where d i [n], i ∈ {1, 2, …, 8} represents the distance between the UAV and the obstacle detected by the sensor in a certain direction. The sensor has a maximum detection distance limit, and S o [n] represents the environmental information vector obtained by the UAV detecting the environmental information through the on-board distance sensor; S s [n] is the global node information vector, s k [n] represents the information vector of a single node k; S c [n] is the information vector of non-cooperative UAVs. Since the input of the neural network is the environmental information vector, and the dimensional change of S c [n] will affect the dimension of the environmental information vector, its dimension is restricted to C, so there is S c [n] = (S l [n]), l ∈ C, C < L. When the number of non-cooperative UAVs near the mission UAV is less than C, the corresponding input dimension is filled with 0; S t is the time consumed for task execution.
[0077] represents the action. The action of the agent is restricted by the mechanical performance of the UAV and can only take an executable action pair sampled from the legal action set and define safety rules. When the next action of the UAV will fly towards an obstacle, cancel this action and give a penalty.
[0078] Represents the transition of state in the entire system.
[0079] represents the reward function set according to the optimization problem.
[0080] γ represents the current value discount of future rewards and its value range is [0,1].
[0081] The dual-depth Q network reinforcement learning model is used to solve the optimization problem. When solving, the settings are as follows:
[0082] Data collection rewards are where α 1 Represents the data collection amount scaling factor, which is used to reduce the data collection amount to an appropriate range. In this embodiment, α 1 Preferably 10 -6 , p k [n] represents the unit data value of the kth node, Represents the total amount of data collected from node k in time slot n+1.
[0083] The remaining task time reward is α 2 Represents the time scaling factor, which is used to reduce the time to a suitable range. In this embodiment, α 2 Preferably it is 0.1.
[0084] The dynamic obstacle avoidance penalty is α 3 Represents the obstacle avoidance penalty coefficient, which is used to enlarge the shortest detection distance to an appropriate range. In this embodiment, α 3 Preferably, it is 10, d represents the minimum value of the distance between the drone and the obstacle detected by the sensor in all directions, D s Indicates the longest detection distance of the drone from the sensor.
[0085] There is a static obstacle collision penalty for known obstacles. α 4 Represents the collision penalty value. In this embodiment, α 4 Preferably 1.
[0086] The safety penalty for multi-drone flight environments is α 5 Indicates the safety distance keeping coefficient of the drone. In this embodiment, α 5 Preferably it is 0.5.
[0087] In order to reduce invalid actions, the invalid action penalty is R ia =-α 6 , α 6 Indicates the invalid action penalty value. In this embodiment, α6 Preferably 1.
[0088] The global reward function is R = R dc +R tk +R uo +R so +R uf +R ia .
[0089] Specifically, the experimental parameter table of this embodiment is as follows:
[0090]
[0091]
[0092] Furthermore, to solve the problem of overestimating the action value, the target Q-value function in the dual-depth Q network is designed as follows:
[0093]
[0094] Among them, S[n+1] is the judgment condition of the final state: all The corresponding node signal disappears, or the task execution time T is greater than the maximum task time constraint T max .
[0095] Furthermore, the steps of using the deep reinforcement learning model to solve the optimization problem are as follows:
[0096] Step 1: Set the signal-to-noise ratio threshold γ 0 , maximum task time T max , the maximum flight speed of the drone v max and maximum steering angle
[0097] Step 2: Initialize the replay buffer D, estimate the network ξ, and initialize the target network ξ by copying the estimated network parameters - .
[0098] Step 3: By v max , Compute the legal action space.
[0099] Step 4: Initialize the experimental environment and set all environmental parameters to initial values.
[0100] Step 5: Linearly decay the learning rate to improve the convergence performance of the model.
[0101] Step 6: Obtain the environmental information vector S[n], and perform state normalization on S[n] to obtain S′[n].
[0102] Step 7: Input S′[n] into the estimation network to obtain action a[n], the agent interacts with the environment and obtains rewards And the new environment information vector S[n+1].
[0103] Step 8: Normalize S[n+1] and convert the tuple Added to the replay buffer D.
[0104] Step 9: Sample from D, use the loss function to calculate the gradient, and update the network parameters.
[0105] Step 10: Every N r The update copies the parameters in ξ to ξ - In. r is a threshold value of the number of training times of the neural network set manually. In this embodiment, N r Preferably 600.
[0106] Step 11, determine whether S[n+1] is the final state; if so, jump to step 4; if not, jump to step 5.
[0107] Step 12: If the data collection rate and the path planning success rate reach the expected values, the training cycle is stopped to obtain the optimal neural network parameters, and the optimal parameters are set as the initial values of each parameter in the neural network to obtain a trained agent. The expected value of the data collection rate is preferably 99%, and the expected value of the path planning success rate is preferably 95%.
[0108] S4. Deploy the intelligent agent to the drone. The drone obtains real-time environmental data. The intelligent agent gives the drone action instructions for the next moment based on the real-time environmental data. The drone performs the data collection task according to the action instructions.
[0109] Further, Figure 3 The convergence process of the proposed algorithm and the comparison algorithm after smoothing is demonstrated. Specifically, the algorithm proposed in the present invention is a drone data collection obstacle avoidance trajectory planning algorithm based on a dual deep Q network, and the comparison algorithm is the basic Dueling DQN algorithm. It can be seen that the reward value of the basic Dueling DQN algorithm continues to fluctuate violently after rising during the training process and the reward value does not reach the maximum. The performance of the proposed algorithm is significantly improved, and the fluctuation is also small during the convergence process, showing stability in the forest fire environment with dynamic obstacles.
[0110] Figure 4The total data collection rate of the drone from all nodes in each mission under different algorithms is demonstrated. Affected by environmental factors, fire spread process, aerial nodes and non-cooperative drones, the collection rate of the basic Dueling DQN algorithm hovered around 90% after the initial rise, and then dropped to around 75%; the algorithm proposed in the present invention reached 100% after the initial rise, and then stabilized at a level close to 100%, demonstrating the effectiveness of the proposed algorithm in mobile node data collection tasks under the influence of dynamic obstacles.
[0111] Figure 5 The data collection rate collected by the drone from the aerial nodes under different algorithms is shown. It can be seen that due to the existence of a good line-of-sight propagation path, the mission drone has achieved a 100% data collection rate under different algorithms, and at the beginning of the training process, the drone has achieved 100% collection of aerial node data.
[0112] Figure 6 The data collection rates of drones for ground nodes including unmanned vehicles and firefighters under different algorithms are demonstrated. Compared with aerial nodes, ground nodes have larger communication channel losses. The mobility of nodes requires the mission drone to flexibly adjust its own trajectory during data collection. It can be seen that the collection rate of the basic Dueling DQN algorithm hovered around 90% after the initial rise, and then dropped to around 70%; the algorithm proposed in the present invention reached 100% after the initial rise, and then stabilized at a level close to 100%.
[0113] Figure 7 The collection rate of some data of high-value nodes by drones under different algorithms is shown. The data with the top 30% of unit data value in the figure is high-value node data. It can be seen that the basic Dueling DQN algorithm fluctuates violently around 90% at the beginning, and then drops to 70%; the algorithm proposed in this invention quickly converges to around 100%, showing a preference for high-value data.
[0114] Figure 8 The success rate of safe trajectory planning of drones under different algorithms is demonstrated. It can be seen that although the same method of environmental perception is adopted, under the condition of the simultaneous existence of static and dynamic obstacles in the environment, the trajectory planning success rate of the basic Dueling DQN algorithm increases initially and then continues to decrease and fluctuates violently, which cannot guarantee the flight safety of the drone itself; the success rate of the algorithm proposed in the present invention converges to more than 95% after a slow increase, demonstrating the adaptability of the proposed algorithm to the complex flight environment of forest fires.
[0115] Finally, it should be noted that the above only describes some embodiments of the present invention. For those skilled in the art, it is conceivable that various changes, modifications, substitutions and deformations may be made to these embodiments without departing from the principles and spirit of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents, and the above-mentioned actions should be covered within the scope of protection of the present invention.
Claims
1. A method for collecting data from a mobile node using an obstacle avoidance drone during a forest fire, characterized in that: include: S1. Acquire the data of the ground mobile nodes and the data of the UAV nodes, and establish a system model according to the data of the ground mobile nodes and the data of the UAV nodes; S2, construct optimization problems based on the system model; S3. Solve the optimization problem using a deep reinforcement learning model; obtain an intelligent agent for planning an obstacle avoidance path based on the solution result; S4. Deploy the intelligent agent to the drone. The drone obtains real-time environmental data. The intelligent agent gives the drone action instructions for the next moment based on the real-time environmental data. The drone performs the data collection task based on the action instructions.
2. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 1 is characterized in that: Ground mobile nodes include unmanned vehicle nodes and firefighter nodes.
3. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 1 is characterized in that: The system model includes system architecture, fire area model, mobile node model and data collection task model.
4. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 3 is characterized in that: The system architecture consists of K wireless sensor nodes, including K1 drone nodes, K2 unmanned vehicle nodes and K3 firefighter nodes, K = K1 + K2 + K3.
5. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 3 is characterized in that: The fire area model uses cellular automata to describe the forest fire spread process.
6. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 3 is characterized in that: The mobile node model calculates the unit data value of the node based on the data of the ground mobile node and the data of the drone node, and the unit data value is used to construct the optimization problem.
7. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 6 is characterized in that: The formula for unit data value is: Among them, p k [n] represents the unit data value of node k in the nth time slot, β represents the node movement trend attention coefficient, It indicates the relative position change between the node and the forest fire, and γ indicates the life safety factor of the firefighters.
8. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 1, characterized in that: The constraints of the optimization problem include: UAV flight speed performance constraints, UAV steering performance constraints, environmental wind uncertainty constraints, obstacle avoidance constraints, safe distance constraints between UAVs, scheduling constraints, data transmission constraints, and task time constraints.
9. The method for obstacle avoidance trajectory planning for unmanned aerial vehicle data collection for mobile nodes according to claim 1, characterized in that: A dual deep Q-network reinforcement learning model is used to solve the optimization problem.
Citation Information
Patent Citations
Forest unmanned cooperative fire fighting system and cooperative ad hoc hybrid network establishment method
CN114257955A
Mountain fire helicopter rescue flight path dynamic planning method
CN114625170A
Aero-engine residual life prediction method based on deep learning
CN115510740A
Forest fire spreading prediction method and system
CN116415754A
Method for predicting microcosmic holes of aluminum alloy product and influence of microcosmic holes on macroscopic service performance
CN117995331A
Cited By
Unmanned aerial vehicle data collection method in forest fire scene
CN120010547A
A method for unmanned aerial vehicle data collection in a forest fire scenario
CN120010547B
Obstacle-avoidance unmanned aerial vehicle-based data collection method for mobile node in forest fire
WO2026170997A1