A Real-time Path Planning and Obstacle Avoidance Method, Device, Medium and Equipment for Unmanned Aerial Vehicles

Through real-time path planning and obstacle avoidance methods based on machine learning, the problem of unsafe flight of drones in complex environments is solved, and efficient and secure autonomous navigation of drones in dynamic environments is achieved.

CN119987409BActive Publication Date: 2025-06-17MIANYANG ZHONGYAN ABRASIVES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510481308.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-06-17
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing drone path planning algorithms are difficult to meet the requirements of real-time and flexibility when facing dynamic environments or obstacles, resulting in unsafe flight of drones in complex environments.

Method used

Real-time path planning and obstacle avoidance methods based on machine learning are adopted, and by obtaining environmental perception information, terrain and meteorological information and its own status information, preprocessing and building path planning and obstacle avoidance models, using multimodal spatiotemporal data fusion module, path generation layer and obstacle avoidance decision layer to generate the optimal flight path and give the best obstacle avoidance mode.

Benefits of technology

Significantly improve the autonomous navigation capabilities of drones in complex and dynamic environments, ensuring that drones can complete path planning and obstacle avoidance tasks safely and efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987409B_ABST
    Figure CN119987409B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, medium and equipment for real-time path planning and obstacle avoidance of an unmanned aerial vehicle. The method includes: obtaining environmental perception information, terrain and meteorological information, and its own state information during the flight of the unmanned aerial vehicle; preprocessing the environmental perception information, terrain and meteorological information, and its own state information; constructing an unmanned aerial vehicle path planning and obstacle avoidance model and training the model; inputting the preprocessed environmental perception information, terrain and meteorological information, and its own state information into the trained unmanned aerial vehicle path planning and obstacle avoidance model to plan the best flight path of the unmanned aerial vehicle and give the best obstacle avoidance mode of the unmanned aerial vehicle. This application can significantly improve the autonomous navigation ability of the unmanned aerial vehicle in complex and dynamic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of unmanned aerial vehicles, and particularly relates to a method, device, medium, and equipment for real-time path planning and obstacle avoidance of unmanned aerial vehicles. Background Art

[0002] With the development of unmanned aerial vehicle technology, its application scope is becoming more and more extensive, including but not limited to fields such as logistics distribution, agricultural monitoring, and film shooting. However, how to ensure the safe flight of unmanned aerial vehicles in complex environments, avoid collisions, and efficiently complete tasks is an urgent problem to be solved currently. Traditional path planning algorithms such as A* and Dijkstra often fail to meet the requirements of real-time performance and flexibility when facing dynamic environments or obstacles. Therefore, developing a real-time path planning and obstacle avoidance method based on machine learning has important research value and practical significance. Summary of the Invention

[0003] The main purpose of this application is to provide a method, device, medium, and equipment for real-time path planning and obstacle avoidance of unmanned aerial vehicles, aiming to improve the autonomous navigation ability of unmanned aerial vehicles when facing dynamic environments or obstacles.

[0004] To achieve the above objectives, this application provides the following technical solutions:

[0005] A method for real-time path planning and obstacle avoidance of an unmanned aerial vehicle, the method includes: obtaining environmental perception information, terrain and meteorological information, and its own state information during the flight of the unmanned aerial vehicle; preprocessing the environmental perception information, terrain and meteorological information, and its own state information; constructing an unmanned aerial vehicle path planning and obstacle avoidance model, and training the model; inputting the preprocessed environmental perception information, terrain and meteorological information, and its own state information into the trained unmanned aerial vehicle path planning and obstacle avoidance model to plan the best flight path of the unmanned aerial vehicle and give the best obstacle avoidance mode of the unmanned aerial vehicle.

[0006] Optionally, preprocessing the environmental perception information, terrain and meteorological information, and its own state information includes: cleaning and denoising the environmental perception information, terrain and meteorological information, and its own state information; synchronizing the data and aligning the timestamps of the environmental perception information, terrain and meteorological information, and its own state information after cleaning and denoising; fusing the environmental perception information, terrain and meteorological information, and its own state information after data synchronization and timestamp alignment.

[0007] Optionally, the UAV path planning and obstacle avoidance model includes: a multi-modal spatio-temporal data fusion module, a path generation layer, and an obstacle avoidance decision layer. Among them, the multi-modal spatio-temporal data fusion module is used to integrate the pre-processed environmental perception information, terrain and meteorological information, and its own state information, and perform multi-scale feature extraction; the path generation layer is used to generate an optimal flight path based on the multi-scale features; the obstacle avoidance decision layer is used to process the obstacles in real time during the flight of the UAV along the optimal flight path and give a reasonable obstacle avoidance decision.

[0008] Optionally, the multi-modal spatio-temporal data fusion module includes: an input layer and a spatio-temporal encoder. Among them, the input layer is used to receive the pre-processed environmental perception information, terrain and meteorological information, and its own state information; the spatio-temporal encoder is used to perform multi-scale feature extraction on the pre-processed environmental perception information, terrain and meteorological information, and its own state information, and plan the UAV flight path based on the multi-scale features.

[0009] Optionally, the spatio-temporal encoder includes: an adaptive multi-scale feature extraction module, a lightweight spatio-temporal attention mechanism, a hierarchical dynamic graph neural network, and a feedback enhancement module connected in sequence.

[0010] Optionally, the UAV path planning and obstacle avoidance model is trained through the following steps: collecting different sensor data and pre-processing it to obtain a model training data set; expanding the data set and dividing the expanded data set into a training set and a validation set; setting training parameters and training the model through the training set. During the training process, when the cross-entropy loss function converges, the model training ends.

[0011] Validating the trained model through the validation set. If the mean absolute error, mean square error, and root mean square error, which are used as model performance evaluation indicators, are all less than the threshold, the model validation passes; otherwise, adjust the training parameters or expand the training set samples to retrain the model until the model validation passes.

[0012] This application also provides a UAV real-time path planning and obstacle avoidance device. The device includes: an acquisition module, which is used to obtain environmental perception information, terrain and meteorological information, and its own state information during the flight of the UAV based on different sensors; a pre-processing module, which is used to pre-process the environmental perception information, terrain and meteorological information, and its own state information; a model construction and training module, which is used to construct a UAV path planning and obstacle avoidance model and train the model; a path planning module, which is used to input the pre-processed environmental perception information, terrain and meteorological information, and its own state information into the trained UAV path planning and obstacle avoidance model to plan the best flight path of the UAV and give the best obstacle avoidance mode of the UAV.

[0013] Optionally, the preprocessing module includes: a cleaning and denoising sub-module for cleaning and denoising the environmental perception information, terrain and meteorological information, and its own state information; a synchronization and alignment sub-module for synchronizing data and aligning timestamps of the cleaned and denoised environmental perception information, terrain and meteorological information, and its own state information; and a fusion sub-module for fusing the environmental perception information, terrain and meteorological information, and its own state information after data synchronization and timestamp alignment.

[0014] The present application also provides a storage medium, which includes instructions that, when running on a computer, cause the computer to execute the method described in any of the previous items.

[0015] The present application also provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method described in any of the previous items when executing the program.

[0016] Compared with the prior art, the present application can bring the following beneficial effects: Through the constructed UAV path planning and obstacle avoidance model, the present application can significantly improve the autonomous navigation ability of the UAV in complex and dynamic environments. The present application can not only maximize the comprehensive utilization efficiency of various sensor data, thereby ensuring that the UAV can complete the path planning and obstacle avoidance tasks more safely and efficiently. Description of the Drawings

[0017] Figure 1 is a schematic flowchart of a UAV real-time path planning and obstacle avoidance method provided by an embodiment of the present application;

[0018] Figure 2 is a schematic structural diagram of a UAV real-time path planning and obstacle avoidance model provided by another embodiment of the present application;

[0019] Figure 3 is a schematic structural diagram of a UAV real-time path planning and obstacle avoidance device provided by another embodiment of the present application;

[0020] Figure 4 is a schematic structural diagram of a storage medium provided by another embodiment of the present application;

[0021] Figure 5 is a schematic structural diagram of an electronic device provided by another embodiment of the present application. Detailed Embodiments

[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0023] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will change accordingly.

[0024] In the present application, unless otherwise clearly defined and limited, the terms "connected", "fixed", etc. should be understood in a broad sense. For example, "fixed" can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0025] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution that satisfies both A and B at the same time. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0026] Figure 1 As shown in the flowchart of a real-time path planning and obstacle avoidance method for an unmanned aerial vehicle according to an embodiment of the present application, Figure 1 as shown, the method includes the following steps:

[0027] S100: Obtain environmental perception information, terrain and meteorological information, and its own state information during the flight of the drone based on different sensors. Among them, the environmental perception information includes static obstacle information (such as buildings, utility poles, trees, etc.) and dynamic obstacle information (such as birds, airborne objects, and other moving drones). The terrain and meteorological information includes terrain information (such as the height changes of the ground) and meteorological information (such as factors like wind speed, wind direction, temperature, humidity, etc.). The own state information includes the attitude angle, position coordinates, and velocity vector of the drone, etc.

[0028] S200: Preprocess the environmental perception information, terrain and meteorological information, and its own state information;

[0029] S300: Construct a drone path planning and obstacle avoidance model and train the model;

[0030] S400: Input the preprocessed environmental perception information, terrain and meteorological information, and its own state information into the trained drone path planning and obstacle avoidance model to plan the best flight path of the drone and give the best obstacle avoidance mode of the drone.

[0031] In this application, by constructing a drone path planning and obstacle avoidance model, it can not only enhance the comprehensive utilization efficiency of various sensor data, but also calmly handle unknown obstacles and environmental changes, thereby ensuring that the drone can safely and efficiently complete the path planning and obstacle avoidance tasks.

[0032] In another exemplary embodiment, in step S200, preprocessing the environmental perception information, terrain and meteorological information, and its own state information includes the following steps:

[0033] S201: Clean and denoise the environmental perception information, terrain and meteorological information, and its own state information;

[0034] In this step, a digital filter (such as a low-pass filter) is used to remove the high-frequency noise in each piece of information, such as the position and velocity data from the IMU or GPS. In addition, it is also necessary to identify and remove data points that deviate significantly from the normal range through statistical methods (such as the 3σ principle). Further, for the missing data points in each piece of information, interpolation methods (linear interpolation, spline interpolation, etc.) can be used for filling.

[0035] S202: Synchronize the data and align the timestamps of the environmental perception information, terrain and meteorological information, and its own state information after cleaning and denoising;

[0036] In this step, it is necessary to correct the timestamps of each sensor so that they are based on the same time reference to avoid errors caused by clock drift. Additionally, it is also necessary to match data streams with different sampling frequencies to ensure that key information is correctly updated at the same moment.

[0037] S203: Fuse the environment perception information, terrain and meteorological information, and its own state information after data synchronization and timestamp alignment.

[0038] In this step, various information obtained based on different sensors can be fused through technologies such as Kalman filtering and particle filtering to improve the position estimation accuracy of the UAV and the reliability of the environmental model.

[0039] In another exemplary embodiment, in step S300, as Figure 2 shown, the UAV path planning and obstacle avoidance model includes: a multi-modal spatio-temporal data fusion module, a path generation layer, and an obstacle avoidance decision layer. Among them, the multi-modal spatio-temporal data fusion module is used to integrate the preprocessed environment perception information, terrain and meteorological information, and its own state information and perform multi-scale feature extraction; the path generation layer is used to generate an optimal flight path based on the extracted multi-scale features; the obstacle avoidance decision layer is used to process obstacles in real time during the flight of the UAV along the optimal flight path and give reasonable obstacle avoidance decisions.

[0040] In this embodiment, the multi-modal spatio-temporal data fusion module includes an input layer and a spatio-temporal encoder. The input layer is used to receive the preprocessed environment perception information, terrain and meteorological information, and its own state information. The spatio-temporal encoder includes an adaptive multi-scale feature extraction module, a lightweight spatio-temporal attention mechanism, a hierarchical dynamic graph neural network, and a feedback enhancement module connected in sequence. Among them, the adaptive multi-scale feature extraction module introduces a gating mechanism and can dynamically select the scale of the convolution kernel according to the real-time environmental complexity. For example, when the UAV is in a dense obstacle area, small-scale convolutions (such as 3×3) are preferentially activated to capture details; while in an open area, it switches to large-scale convolutions (such as 7×7) to improve computational efficiency. In addition, the adaptive multi-scale feature extraction module also includes a feature pyramid network to perform cross-layer fusion on the multi-scale features extracted based on the gating mechanism to enhance the semantic understanding of complex terrains (such as hills and canyons) and reduce redundant features at the same time.

[0041] The lightweight spatio-temporal attention mechanism adopts a hierarchical attention design, including a local attention layer and a global attention layer. Among them, the local attention layer is based on a sliding window mechanism and only focuses on obstacles within the current field of view of the drone (such as within a 50-meter radius range) to reduce the computational complexity of the model. The global attention layer uses a sparse Transformer to model long-range dependencies only for key dynamic obstacles (such as fast-moving birds), balancing accuracy and real-time performance. The lightweight spatio-temporal attention mechanism also includes an embedded lightweight LSTM module to record the historical obstacle trajectories and predict the movement trends within the next 3 seconds, providing forward-looking decision support for dynamic obstacle avoidance.

[0042] The hierarchical dynamic graph neural network adopts a two-layer graph structure, including a global graph layer and a local graph layer. Among them, the global graph layer uses terrain elevation and meteorological data as nodes to generate a global optimal path skeleton. The local graph layer uses real-time obstacles (static + dynamic) as nodes and dynamically updates the connection weights to achieve refined obstacle avoidance. In addition, the hierarchical dynamic graph neural network also introduces a graph pruning strategy, that is, it can prune irrelevant nodes (such as obstacles beyond the sensor range) according to the real-time position and speed of the drone, thereby improving the computational efficiency of the model.

[0043] The feedback enhancement module feeds back the output of the hierarchical dynamic graph neural network (such as obstacle avoidance decision confidence or importance score) to the multi-scale feature extraction module through a feedback mechanism to guide the focus of feature extraction for the next frame of data. Specifically, the feedback enhancement module includes an input layer, a filter, and an integration layer connected in sequence. Among them, the input layer is used to receive the output information from the hierarchical dynamic graph neural network, including but not limited to obstacle position, type, predicted movement direction, and obstacle avoidance decision confidence, etc. The filter filters the received information according to predefined criteria (such as a confidence threshold) to determine which important information needs to be fed back to the multi-scale feature extraction module, which can reduce the unnecessary computational burden of the model and improve the running efficiency of the model. The integration layer is used to integrate the output information from the hierarchical dynamic graph neural network with the original input data (i.e., environmental perception information, terrain and meteorological information, and the drone's own state information) and feed the integrated information back to the multi-scale feature extraction module to help the multi-scale feature extraction module dynamically adjust the focus of feature extraction. For example, when encountering high-confidence obstacles or complex terrains, the feature extraction weights of these areas can be increased to analyze the information of these key areas more carefully, which helps to improve the response speed and accuracy of the drone to environmental changes.

[0044] The path generation layer uses the following algorithm to generate the flight path of the drone. This algorithm adjusts the path cost by introducing terrain elevation and static obstacle distribution to generate a more practical drone flight path. The specific algorithm is as follows:

[0045]

[0046] Among them, represents the total cost through node and is used to determine which path is optimal; represents the path length, that is, the actual distance from the current node to the next node; represents the terrain risk; represents the energy consumption of the UAV; represents the wind influence index; represents the load sensitivity; represents the time sensitivity; , , , , and respectively represent weight parameters and can be adjusted according to specific task requirements.

[0047] Next, the present application will specifically describe how to generate a UAV flight path based on the above algorithm:

[0048] First, set the starting point as the current node and mark all other nodes as unexplored. Create a priority queue to store the nodes to be explored and their corresponding estimated total costs.

[0049] Secondly, take out the node with the lowest estimated total cost from the priority queue as the current node. For each neighbor node of the current node, comprehensively consider information such as impassable areas, static obstacle distributions, and dynamic obstacle predictions. Calculate the actual cost of reaching each neighbor node through the current node and update the total cost of the neighbor node using the above more complex cost function. According to real-time data (such as wind speed and wind direction changes), dynamically adjust the relevant parameters in the cost function (such as and ) to reflect the latest environmental conditions.

[0050] Finally, repeat step 2 until the target node is found or the open list is empty. Once the target node is found, trace back to the starting point by tracking the predecessor nodes of each node to obtain the complete optimal path.

[0051] The above algorithm can provide a more accurate and flexible path planning solution by integrating multiple environmental factors and a dynamic adjustment mechanism.

[0052] The obstacle avoidance decision layer adopts a spatio-temporal attention reinforcement learning network, which specifically includes a state space layer, an action space layer, a reward function layer, and an attention mechanism.

[0053] The state space layer defines the set of states that the UAV is in at any given moment, which typically includes a series of variables that can describe the current situation of the UAV. Specifically, in the local obstacle avoidance scenario of the UAV, the state space can include the following aspects of information:

[0054] Position information: The three-dimensional coordinates of the UAV's current position ( x, y, z ).

[0055] Attitude information: Including pitch angle ( pitch ), roll angle ( roll ), and yaw angle ( yaw ), these angles determine the direction of the UAV.

[0056] Velocity vector: Forward velocity ( vx ), lateral velocity (vy ), and vertical velocity ( vz ).

[0057] Sensor data: Information from different sensors, such as the distance, direction, etc. of surrounding obstacles provided by lidar ( LiDAR ).

[0058] Environmental characteristics: Factors affecting flight stability, such as wind speed (windSpeed), wind direction (windDirection), etc.

[0059] Task-related information: Factors that may affect decision-making, such as remaining battery level (batterylevel), payload status, etc.

[0060] The state space layer can be represented as:

[0061]

[0062] Where, represents the state space, and respectively represent the distance to the nearest obstacle and its direction relative to the UAV.

[0063] The action space layer defines all possible actions that the UAV can perform, aiming to help the UAV avoid obstacles and continue moving towards the target. According to the capabilities and application scenarios of the UAV, the action space can include the following types of actions:

[0064] 1. Translational motion: Move forward ( ), move backward ( ), move left or right ( ), rise or fall ( );

[0065] 2. Rotational motion: Rotation around the X-axis (changing the pitch angle), rotation around the Y-axis (changing the roll angle),

[0066] rotation around the Z-axis (changing the yaw angle) or staying at the current position ( ).

[0067] 3. Composite actions: Combining translation and rotation to achieve more complex maneuvers, such as ascending while turning.

[0068] The action space layer can be represented as:

[0069]

[0070] where, represents the action space. When the UAV needs to quickly change direction to avoid collisions, by turning left ( ) and turning right ( ), it selects an appropriate turning angle (d yaw ) according to the position of the obstacle.

[0071] The reward function layer defines the feedback obtained by the UAV after taking a certain action, so as to guide the UAV to learn how to make the best decision to achieve the goal. The specific feedback includes:

[0072] Giving a positive reward when the UAV successfully avoids obstacles; increasing the reward value as the UAV gradually approaches the final target position; giving a negative reward if the UAV collides with an obstacle; encouraging the adoption of an action sequence with lower energy consumption to extend the flight time.

[0073] The reward function layer is represented as follows:

[0074]

[0075] where, represents the reward for approaching the target, represents the collision penalty, represents the energy consumption optimization reward, represents the attitude stability reward.

[0076] By setting the reward function in this application, it can not only effectively guide the UAV to avoid obstacles, but also ensure that it efficiently and economically completes the predetermined task. In this way, the UAV can autonomously learn and optimize its flight path in a complex real-world environment.

[0077] The attention mechanism includes a spatio-temporal feature extraction module, an adaptive fusion layer, and a reinforcement learning decision-making layer. Among them, the spatio-temporal feature extraction module includes a spatial attention branch and a temporal attention branch. The spatial attention branch uses a convolutional neural network (CNN) to extract spatial features from sensor data, such as the distance and direction of obstacles, and generates a spatial attention map to highlight the areas that require special attention. The temporal attention branch uses a recurrent neural network (RNN) to process historical data, predict the future movement trajectory of obstacles, and form temporal attention weights to emphasize the obstacles that may have a significant impact on the flight path in the future.

[0078] The adaptive fusion layer is used to combine the outputs of the spatial attention branch and the temporal attention branch and form a comprehensive attention vector (this process can be completed, for example, by weighted summation, aiming to ensure that the final attention distribution can reflect both the current spatial layout and predict future change trends). In addition, the adaptive fusion layer also introduces an adaptive adjustment factor to allow the attention mechanism to adjust the importance ratio of each component according to real-time environmental feedback. For example, in a rapidly changing environment, the proportion of temporal attention is increased, while in a relatively stable environment, more emphasis is placed on spatial feature analysis.

[0079] Exemplarily, assuming that the outputs of the spatial attention branch and the temporal attention branch are respectively and , this application uses the method of weighted summation and introduces an adaptive adjustment factor to adjust the weight distribution, specifically as follows:

[0080]

[0081] Among them, represents the comprehensive attention vector, and the adaptive adjustment factor is an adjustment factor dynamically adjusted according to the Environmental Dynamics Index (EDI), which is expressed as:

[0082]

[0083] When is close to 0, that is, the environment is very unstable, is close to 1, meaning that more reliance is placed on the temporal attention branch to predict the future movement trend of obstacles.

[0084] When is close to 1, that is, the environment is very stable, is close to 0, and at this time, the decision-making mainly relies on the immediate information provided by the spatial attention branch.

[0085] In summary, by introducing an adaptive adjustment factor in this application , the UAV can flexibly adjust its focus when facing different environments, thereby improving the quality and efficiency of obstacle avoidance decision-making.

[0086] The reinforcement learning decision-making layer takes the comprehensive attention vector generated by the adaptive fusion layer as input to help the model better understand which factors are the most critical, so that the model can formulate the most effective obstacle avoidance strategy.

[0087] Next, the mechanism of action of the reinforcement learning decision-making layer in this application is introduced as follows:

[0088] First, the state space is combined with the comprehensive attention vector to form an extended state representation , that is ;

[0089] Second, define the function , which is used to estimate the expected cumulative reward that can be obtained by taking the action under the extended state . Among them, represents the weight parameter.

[0090] Finally, use the TD error to define the objective function to update the weight parameter . The objective function is expressed as follows:

[0091]

[0092] Among them,

[0093]

[0094] represents the immediate reward; represents the average of the distributions of the state , the action , the immediate reward and the next state ; represents the target Q value, which is composed of the immediate reward and the estimate of the future reward ; represents the discount factor, which is between 0 and 1; represents for a given next state , select the action that maximizes the Q value; represents the parameters of the target network, which are updated to a copy of the current network parameters after a certain number of steps; Represents the Q-value estimation of the current network, which represents the cumulative reward expected to be obtained after taking action A in the extended state ; Represents the square of the TD error, which is used to measure the gap between the target Q-value and the Q-value predicted by the current network.

[0095] Through the above introduction, the reinforcement learning decision-making layer not only depends on the current state and action, but also needs to combine future reward predictions and environmental feedback for learning and optimization, so that the model can help the drone make more intelligent and effective decisions in complex environments.

[0096] By combining the reward function as shown above, the attention mechanism can further optimize the attention allocation strategy, ensuring that the drone can not only avoid the immediate obstacles, but also anticipate potential risks and take actions in advance.

[0097] In this application, by introducing the attention mechanism into the obstacle avoidance decision-making layer, the model can focus on the most important parts when processing data. For example, in a complex environment, not all detected objects need to be equally emphasized. By introducing the attention mechanism, the most urgent obstacles to be avoided can be highlighted according to the current flight state. In addition, through learning historical data, the attention mechanism can help the model predict the future movement direction of the obstacles, so as to make avoidance decisions in advance. Further, based on the changes in the surrounding environment, the attention mechanism allows the drone to adjust its obstacle avoidance strategy in real time to ensure that the optimal path is always selected.

[0098] In another exemplary embodiment, in step S300, the drone path planning and obstacle avoidance model is trained through the following steps:

[0099] S301: Collect different sensor data and perform preprocessing to obtain the model training dataset;

[0100] S302: Augment the dataset, for example, by adding noise, changing the lighting conditions, simulating different weather conditions, etc., and divide the augmented dataset into a training set and a validation set, and the division ratio can be 7:3 for example;

[0101] S303: Set the training parameters. For example, the booster defaults to select the tree-based model gbtree, the learning_rate is set to 0.01, and the max_depth is set to 3. Train the model through the training set. During the training process, when the cross-entropy loss function converges, the model training ends;

[0102] S304: Validate the trained model using the validation set. If the mean absolute error (MAE), mean squared error (MSE), and root mean squared error (RMSE), which are used as model performance evaluation metrics, are all less than the thresholds (the threshold for MAE is set to 0.05, and the threshold for MSE is set to 0.001), then the model passes the validation; otherwise, adjust the training parameters (for example, adjust learning_rate to 0.005 and max_depth to 5) or expand the training set samples (for example, adjust the division ratio to 8:2) and retrain the model until the model passes the validation.

[0103] In this embodiment, the present application proposes an original loss function that combines dynamic weight adjustment and uncertainty estimation. The function is expressed as follows:

[0104]

[0105] where, represents the total number of samples; represents the true label; represents the corresponding value in the probability distribution predicted by the model; represents the dynamic weight function, which adaptively adjusts the weight of each sample according to the current spatial feature and temporal feature ; represents the uncertainty estimation, which is used to measure the prediction confidence of the model under a given input; represents the balance coefficient, which is used to adjust the influence of the uncertainty estimation on the total loss.

[0106] In addition, to further improve the robustness of the model, the present application further introduces an adversarial sample generator during the training process to simulate environmental changes in the worst-case scenario. The adversarial sample generator is expressed as follows:

[0107]

[0108] where, represents the perturbation generated by the adversarial sample generator, and the goal is to find the perturbation that maximizes the loss function.

[0109] Then the final cross-entropy loss function is expressed as the weighted sum of the original loss and the adversarial loss:

[0110]

[0111] where, represents the weight of the adversarial loss.

[0112] The cross-entropy loss function shown in this application can significantly improve the adaptability and robustness of the UAV path planning and obstacle avoidance model by introducing mechanisms such as dynamic weight adjustment, uncertainty estimation, and adversarial learning. It can not only process complex spatio-temporal data more accurately but also effectively cope with the uncertainties and potential threats in the environment, thus ensuring that the UAV can perform tasks safely and efficiently under various conditions.

[0113] Next, this application will compare and explain the method described in this application with traditional methods by defining specific scenarios.

[0114] The scenario is defined as follows:

[0115] Environmental perception information: The UAV is flying over a forest. Its sensors detect a 30-meter-high hill 50 meters ahead, and at the same time, there is a flock of birds moving northeast 200 meters to its right.

[0116] Terrain information: The ground is undulating, with multiple hills and valleys. The highest hill has an altitude of 100 meters.

[0117] Meteorological information: The current wind speed is 5 m / s, coming from the southwest. The temperature is moderate, and the humidity is high but does not affect flight.

[0118] UAV's own state information: The current position coordinates of the UAV are (0, 0, 50) m (relative to the takeoff point), the target position is (500, 500, 50) m, the current speed is 10 m / s, and the optimal cruising altitude without side wind influence is 50 m.

[0119] Based on the above information, this application processes the environmental perception information, terrain and meteorological information, and its own state information through a multi-modal spatio-temporal data fusion module, extracts features, and selects a path slightly deviating from a straight line for the UAV, that is, first ascending to a height of 70 meters to cross the hill ahead, and then adjusting the heading according to the predicted movement trend of the birds to ensure a safe distance. Regarding the A* algorithm, if the three-dimensional space is simplified to a two-dimensional plane graph and the positions of static obstacles are known, this algorithm can find a shortest path from the starting point to the ending point. However, for dynamic obstacles (such as birds), this algorithm needs to frequently update the node weights, which will lead to a huge increase in the amount of calculation and make it difficult to achieve real-time response. The Dijkstra algorithm does not use a heuristic function to guide the search process, which means it is less efficient in large and complex environments. For situations with a large number of obstacles and dynamic changes, Dijkstra cannot quickly find a feasible solution.

[0120] In addition, this application uses a spatio-temporal attention reinforcement learning network to analyze the surrounding environment in real time and gives the best obstacle avoidance strategy based on the current state. For example, when detecting a bird approaching, the drone's heading or altitude is changed in advance to avoid collision. In the case of emergencies (such as a bird approaching suddenly), the A* algorithm needs to rely on an external system to detect obstacles and then manually re-plan the path, with a relatively long reaction time and unable to avoid dynamic obstacles in time. Similarly, the Dijkstra algorithm lacks the ability to react immediately when facing dynamic obstacles and usually requires additional mechanisms to detect and handle obstacles.

[0121] In summary, since this method takes dynamic obstacle avoidance into account and can react within a short time, the overall flight time is relatively short (e.g., 51 seconds). In contrast, the A* and Dijkstra algorithms may have a slightly longer flight time (e.g., 55 seconds) due to the need to recalculate the path, resulting in delays. In addition, this method improves safety by predicting the behavior patterns of obstacles and reduces the risk of collision; while traditional algorithms are difficult to achieve such fine adjustments without external intervention and have a higher risk of collision.

[0122] In another exemplary embodiment, this application also provides a real-time path planning and obstacle avoidance device for drones, as Figure 3 shown, the device includes: an acquisition module 100, configured to obtain environmental perception information, terrain and meteorological information, and its own state information during the flight of the drone based on different sensors; a preprocessing module 200, configured to preprocess the environmental perception information, terrain and meteorological information, and its own state information; a model construction and training module 300, configured to construct a drone path planning and obstacle avoidance model and train the model; a path planning module 400, configured to input the preprocessed environmental perception information, terrain and meteorological information, and its own state information into the trained drone path planning and obstacle avoidance model to plan the best flight path of the drone and give the best obstacle avoidance mode of the drone.

[0123] Optionally, the preprocessing module 200 includes: a cleaning and denoising sub-module, configured to clean and denoise the environmental perception information, terrain and meteorological information, and its own state information; a synchronization and alignment sub-module, configured to perform data synchronization and timestamp alignment on the cleaned and denoised environmental perception information, terrain and meteorological information, and its own state information; a fusion sub-module, configured to fuse the environmental perception information, terrain and meteorological information, and its own state information after data synchronization and timestamp alignment.

[0124] Based on the above embodiments, refer to Figure 4 to describe the computer-readable storage medium of the exemplary embodiments of this application. Please refer to Figure 4, which shows that the computer-readable storage medium is an optical disc 40, on which a computer program (i.e., program product) is stored. When the computer program is run by a processor, it will implement the steps recorded in the above method embodiments. For example, based on different sensors, environmental perception information, terrain and meteorological information, and its own state information during the flight of the drone are obtained; the environmental perception information, terrain and meteorological information, and its own state information are preprocessed; a drone path planning and obstacle avoidance model is constructed and the model is trained; the preprocessed environmental perception information, terrain and meteorological information, and its own state information are input into the trained drone path planning and obstacle avoidance model to plan the best flight path of the drone and give the best obstacle avoidance mode of the drone. The specific implementation methods of each step will not be repeated here.

[0125] It should be noted that the computer-readable storage medium includes, but is not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated one by one here.

[0126] Based on the above embodiments, the present application also provides an electronic device. The following will refer to Figure 5 to describe the electronic device for file downloading according to the exemplary embodiments of the present application.

[0127] Figure 5 shows a block diagram of an exemplary electronic device 50 suitable for implementing the embodiments of the present application. The electronic device 50 may be a computer system or a cloud server. Figure 5 The shown electronic device 50 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0128] As Figure 5 shown, the electronic device 50 includes, but is not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 connecting different system components (including the system memory 502 and the processing unit 501).

[0129] The electronic device 50 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 50, including volatile and non-volatile media, removable and non-removable media.

[0130] System memory 502 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 5021 and / or cache memory 5022. Electronic device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 5023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown in the figure, commonly referred to as a "hard disk drive"). Although not shown in Figure 5 the figure, a disk drive for reading and writing removable non-volatile disks (such as "floppy disks"), and an optical disk drive for reading and writing removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) may be provided. In these cases, each drive may be connected to bus 503 through one or more data media interfaces. System memory 502 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present application.

[0131] A program / utility 5025 having a set (at least one) of program modules 5024 may be stored, for example, in system memory 502, and such program modules 5024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data, and an implementation of a network environment may be included in each or some combination of these examples. Program modules 5024 generally perform the functions and / or methods in the embodiments described in the present application.

[0132] Electronic device 50 may also communicate with one or more external devices 504 (such as a keyboard, a pointing device, a display, etc.). Such communication may be through an input / output (I / O) interface 505. Further, electronic device 50 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 506. As Figure 5 shown, network adapter 506 communicates with other modules (such as processing unit 501, etc.) of electronic device 50 through bus 503. It should be understood that although not shown in Figure 5 the figure, other hardware and / or software modules may be used in conjunction with electronic device 50.

[0133] The processing unit 501 executes various functional applications and data processing by running the programs stored in the system memory 502. For example, it obtains environmental perception information, terrain and meteorological information, and its own status information during the flight of the drone based on different sensors; preprocesses the environmental perception information, terrain and meteorological information, and its own status information; constructs a drone path planning and obstacle avoidance model and trains the model; inputs the preprocessed environmental perception information, terrain and meteorological information, and its own status information into the trained drone path planning and obstacle avoidance model to plan the best flight path of the drone and give the best obstacle avoidance mode of the drone. The specific implementation methods of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the file concurrent download device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0134] In the description of the present application, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0135] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0136] In several embodiments provided in the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0137] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, can exist physically alone for each unit, or two or more units can be integrated into one unit.

[0139] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a cloud server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0140] The above embodiments are only used to illustrate the technical concept and characteristics of the present application, and the purpose is to enable those who are familiar with this technology to understand the content of the present application and implement it accordingly. It cannot be used to limit the protection scope of the present application. All equivalent changes or modifications made according to the spirit and essence of the present application should be covered within the protection scope of the present application.

Claims

1. A real-time path planning and obstacle avoidance method for an unmanned aerial vehicle, characterized in that: The method comprises: Obtain environmental perception information, terrain and weather information, and the drone's own status information during flight; Preprocessing the environmental perception information, terrain and meteorological information, and self-state information includes: Cleaning and denoising the environmental perception information, terrain and meteorological information, and self-state information; Performing data synchronization and timestamp alignment on the cleaned and denoised environmental perception information, terrain and meteorological information, and self-state information; The environmental perception information, terrain and weather information, and self-status information after data synchronization and timestamp alignment are integrated; a UAV path planning and obstacle avoidance model is constructed, and the model is trained; The UAV path planning and obstacle avoidance model includes: Multimodal spatiotemporal data fusion module, path generation layer and obstacle avoidance decision layer, among which, The multimodal spatiotemporal data fusion module is used to integrate the pre-processed environmental perception information, terrain and meteorological information, and self-state information and perform multi-scale feature extraction; The path generation layer is used to generate an optimal flight path based on multi-scale features; The obstacle avoidance decision layer is used to process obstacles encountered by the UAV in real time while it is flying along the optimal flight path, and to make reasonable obstacle avoidance decisions; The multimodal spatiotemporal data fusion module includes: Input layer and spatiotemporal encoder, where The input layer is used to receive the pre-processed environmental perception information, terrain and meteorological information and self-state information; The spatiotemporal encoder is used to extract multi-scale features from the pre-processed environmental perception information, terrain and meteorological information, and self-state information, and plan the flight path of the UAV based on the multi-scale features; The space-time encoder comprises: The adaptive multi-scale feature extraction module, lightweight spatiotemporal attention mechanism, hierarchical dynamic graph neural network and feedback enhancement module are connected in sequence; The UAV path planning and obstacle avoidance model is trained through the following steps: Collect different sensor data and preprocess them to obtain model training data sets; Expand the data set and divide the expanded data set into a training set and a validation set; Set the training parameters and train the model using the training set. During the training process, when the cross entropy loss function converges, the model training ends. The trained model is verified through the validation set. If the mean absolute error, mean square error, and root mean square error, which are the model performance evaluation indicators, are all less than the threshold, the model is verified. Otherwise, the training parameters are adjusted or the training set samples are expanded to retrain the model until the model is verified. The pre-processed environmental perception information, terrain and weather information, and self-state information are input into the trained UAV path planning and obstacle avoidance model to plan the optimal flight path of the UAV and provide the optimal obstacle avoidance mode of the UAV.

2. A real-time path planning and obstacle avoidance device for an unmanned aerial vehicle for implementing the method as claimed in claim 1, characterized in that: The device comprises: The acquisition module is used to obtain the environmental perception information, terrain and weather information, and the drone's own status information during flight based on different sensors; A preprocessing module, used for preprocessing the environmental perception information, terrain and meteorological information, and self-state information; Preprocessing the environmental perception information, terrain and meteorological information, and self-state information includes: Cleaning and denoising the environmental perception information, terrain and meteorological information, and self-state information; Performing data synchronization and timestamp alignment on the cleaned and denoised environmental perception information, terrain and meteorological information, and self-state information; Fusing the environmental perception information, terrain and meteorological information, and self-state information after data synchronization and timestamp alignment; Model building and training module, used to build the UAV path planning and obstacle avoidance model and train the model; The UAV path planning and obstacle avoidance model includes: Multimodal spatiotemporal data fusion module, path generation layer and obstacle avoidance decision layer, among which, The multimodal spatiotemporal data fusion module is used to integrate the pre-processed environmental perception information, terrain and meteorological information, and self-state information and perform multi-scale feature extraction; The path generation layer is used to generate an optimal flight path based on multi-scale features; The obstacle avoidance decision layer is used to process obstacles encountered by the UAV in real time while it is flying along the optimal flight path, and to make reasonable obstacle avoidance decisions; The multimodal spatiotemporal data fusion module includes: Input layer and spatiotemporal encoder, where The input layer is used to receive the pre-processed environmental perception information, terrain and meteorological information and self-state information; The spatiotemporal encoder is used to extract multi-scale features from the pre-processed environmental perception information, terrain and meteorological information, and self-state information, and plan the flight path of the UAV based on the multi-scale features; The space-time encoder comprises: The adaptive multi-scale feature extraction module, lightweight spatiotemporal attention mechanism, hierarchical dynamic graph neural network and feedback enhancement module are connected in sequence; The UAV path planning and obstacle avoidance model is trained through the following steps: Collect different sensor data and preprocess them to obtain model training data sets; Expand the data set and divide the expanded data set into a training set and a validation set; Set the training parameters and train the model using the training set. During the training process, when the cross entropy loss function converges, the model training ends. The trained model is verified through the validation set. If the mean absolute error, mean square error, and root mean square error, which are the model performance evaluation indicators, are all less than the threshold, the model is verified. Otherwise, the training parameters are adjusted or the training set samples are expanded to retrain the model until the model is verified. The path planning module is used to input the pre-processed environmental perception information, terrain and meteorological information, and its own state information into the trained UAV path planning and obstacle avoidance model to plan the optimal flight path of the UAV and provide the optimal obstacle avoidance mode of the UAV.

3. The real-time path planning and obstacle avoidance device for unmanned aerial vehicles according to claim 2, characterized in that: The preprocessing module comprises: A cleaning and denoising submodule, for cleaning and denoising the environmental perception information, terrain and meteorological information, and self-state information; A synchronization and alignment submodule, for performing data synchronization and timestamp alignment on the cleaned and denoised environmental perception information, terrain and meteorological information, and self-state information; The fusion submodule is used to fuse the environmental perception information, terrain and meteorological information, and self-state information after data synchronization and timestamp alignment.

4. A storage medium, characterized in that: The method comprises instructions which, when executed on a computer, cause the computer to perform the method of claim 1 .

5. An electronic device, characterized in that: The electronic device comprises: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to claim 1 is implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle task trajectory optimization method and device based on deep learning

    CN118644028A