Unmanned aerial vehicle autonomous obstacle avoidance optimization method based on laser radar

By building a drone obstacle avoidance system through lidar and reinforcement learning algorithm, the problems of sensors being affected by the environment and dynamic obstacle recognition are solved, and high-precision obstacle avoidance and stable flight of drones are achieved in complex environments.

CN120742935AActive Publication Date: 2025-10-03GUANGZHOU YOUFEI INTELLIGENT EQUIP CO LTD

Patent Information

Application Number
CN202511203126.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-03
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

In existing drone autonomous obstacle avoidance technologies, visual sensors are easily affected by lighting and weather, obstacle recognition accuracy is low, ultrasonic sensors have limited detection distance and difficulty obtaining three-dimensional position information, the motion state of dynamic obstacles is not fully considered, path planning lacks real-time response, and fixed model parameters are difficult to adapt to dynamic environmental changes, resulting in a decline in obstacle avoidance performance.

Method used

LiDAR is used to obtain real-time point cloud data streams, build an obstacle spatial distribution model, and combine reinforcement learning algorithms to predict collision risk values, generate obstacle avoidance path optimization instructions, and update model parameters through real-time flight trajectory feedback to achieve dynamic optimization.

Benefits of technology

It improves the obstacle avoidance accuracy and stability of drones in complex environments, enables them to respond to environmental changes in a timely manner, adapt to multi-obstacle scenarios, and enhances flight safety and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120742935A_ABST
    Figure CN120742935A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle autonomous obstacle avoidance, and discloses an unmanned aerial vehicle autonomous obstacle avoidance optimization method based on a laser radar. According to the method, a real-time point cloud data stream in a flight environment of the unmanned aerial vehicle is obtained through a laser radar scanning module, wherein the real-time point cloud data stream comprises three-dimensional position information and a dynamic movement track of an obstacle; and constructing an obstacle space distribution model based on the real-time point cloud data stream to represent the geometrical shape and the motion state of the obstacle. A reinforcement learning algorithm is adopted, a collision risk value is calculated based on the position information and the motion trail in the model, and an obstacle avoidance path optimization instruction is generated according to the risk value and used for adjusting the flight direction and speed of the unmanned aerial vehicle. And when the instruction is executed, flight path feedback data is recorded in real time, and parameters of the obstacle space distribution model are updated based on the feedback data. According to the method, the adaptability and reliability of autonomous obstacle avoidance of the unmanned aerial vehicle in a complex environment are improved, and the flight safety is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous obstacle avoidance for unmanned aerial vehicles (UAVs), and in particular to an autonomous obstacle avoidance optimization method for UAVs based on laser radar. Background Art

[0002] With the rapid development of drone technology, its application in aerial photography, logistics and transportation, power inspections, agricultural plant protection, and other fields is becoming increasingly widespread. In these applications, drones often need to fly autonomously in complex and changing environments, and autonomous obstacle avoidance is a core element in ensuring drone flight safety. Currently, autonomous obstacle avoidance technology for drones primarily relies on various sensors for environmental perception, coupled with algorithms for obstacle identification and path planning.

[0003] In existing technologies, some drones use visual sensors for environmental perception, using cameras to collect image data and analyze it to identify obstacles. However, visual sensors are easily affected by environmental factors such as lighting and weather conditions. In scenes such as strong light, backlight, rain, or haze, the clarity of image data can be significantly reduced, resulting in reduced obstacle recognition accuracy and even missed or false detections. Other drones use ultrasonic or infrared sensors. While these sensors can detect obstacles at short distances, their detection range is limited, and it is difficult to accurately obtain the three-dimensional position information of obstacles, making them unable to meet the obstacle avoidance requirements of drones during medium and long-distance flight.

[0004] When it comes to obstacle modeling and risk assessment, traditional methods often build models for static obstacles, focusing primarily on the geometric shape and location of the obstacles, while paying insufficient attention to the motion state of dynamic obstacles. In actual flight environments, there are a large number of dynamic obstacles, such as birds, other drones, and moving ground vehicles. The motion trajectories of these obstacles are uncertain. If collision risk assessment is performed solely based on static models, it is very likely that the assessment results will deviate significantly from the actual situation, increasing the risk of collisions. Furthermore, existing obstacle avoidance path planning is mostly based on preset algorithmic models, and the generation of path optimization instructions lacks dynamic response to real-time environmental changes. When the number of obstacles in the environment increases or the motion state suddenly changes, problems such as unreasonable obstacle avoidance paths and delayed adjustments are likely to occur.

[0005] Existing systems often keep model parameters fixed after initial setup, lacking effective real-time feedback mechanisms. During flight, the environment is constantly changing, with new obstacles constantly appearing and existing obstacles changing their motion. Fixed model parameters struggle to adapt to these dynamic changes, leading to a gradual decline in the system's obstacle avoidance performance and an inability to consistently guarantee drone flight safety. Summary of the Invention

[0006] The purpose of the present invention is to provide a laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles to solve the problems raised in the above background technology.

[0007] To achieve the above objectives, the present invention provides a laser radar-based autonomous obstacle avoidance optimization method for a UAV, the method comprising: The laser radar scanning module is used to obtain a real-time point cloud data stream in the UAV flight environment, wherein the real-time point cloud data stream includes the three-dimensional position information and dynamic motion trajectory of obstacles; Building an obstacle spatial distribution model based on the real-time point cloud data stream, wherein the obstacle spatial distribution model is used to characterize the geometric shape and motion state of the obstacle; A reinforcement learning algorithm is used to predict a collision risk value between the drone and an obstacle, where the collision risk value is calculated based on the position information and motion trajectory in the obstacle spatial distribution model; generating an obstacle avoidance path optimization instruction according to the collision risk value, wherein the obstacle avoidance path optimization instruction is used to adjust the flight direction and speed of the UAV; Executing the obstacle avoidance path optimization instructions to control the UAV's drive system and record flight trajectory feedback data in real time; Parameters of the obstacle spatial distribution model are updated based on the flight trajectory feedback data.

[0008] Preferably, the step of acquiring a real-time point cloud data stream in the UAV flight environment by using a laser radar scanning module includes: Configure the scanning frequency and resolution parameters of the LiDAR sensor to ensure that the point cloud collection covers the preset monitoring area; Perform noise filtering on the original point cloud data to remove environmental interference signals; Convert the filtered point cloud data into a standardized 3D coordinate sequence.

[0009] Preferably, the step of constructing the obstacle spatial distribution model includes: Dividing the three-dimensional coordinate sequence into a plurality of spatial grid units; Perform point cloud clustering analysis within each spatial grid cell to identify obstacle boundary features; Generate obstacle geometry description vector and motion speed estimation based on clustering results; The outputs of all spatial grid cells are integrated to form the complete structure of the obstacle spatial distribution model.

[0010] Preferably, the method of predicting the collision risk value between the drone and the obstacle using a reinforcement learning algorithm includes: Inputting the obstacle geometry description vector and motion velocity estimate into the state space of the reinforcement learning algorithm; Design a reward function mechanism based on the relative distance between the drone's current position and the obstacle; The reinforcement learning policy network is updated through iterative training to output the probability distribution of the collision risk value.

[0011] Preferably, the generating of the obstacle avoidance path optimization instruction includes: Analyzing the probability distribution of the collision risk value to determine high-risk obstacle areas; Calculating a candidate safe flight path using a dynamic path planning algorithm, wherein the candidate path includes a plurality of path nodes; Evaluate the feasibility score of each path node and select the path solution with the highest feasibility score as the obstacle avoidance path optimization instruction; The obstacle avoidance path optimization instruction includes a specific direction angle adjustment value and a speed change.

[0012] Preferably, the driving system for controlling the drone includes: Converting the direction angle adjustment value and the speed change into a motor control signal; Sending the motor control signal to the drone propeller via the drive interface module; Synchronously collect thruster response data and real-time position data; The thruster response data and real-time position data form part of the flight trajectory feedback data.

[0013] Preferably, updating the parameters of the obstacle space distribution model based on the flight trajectory feedback data includes: Comparing the real-time location data with the deviation value of the predicted path node; If the deviation value exceeds a preset tolerance threshold, adjusting the division rule of the spatial grid unit; Re-performing point cloud clustering analysis to update the obstacle geometric shape description vector; The updated obstacle spatial distribution model is used for the next round of collision risk value prediction.

[0014] Preferably, the method further comprises performing data preprocessing before constructing the obstacle spatial distribution model: Extracting timestamp information from the three-dimensional coordinate sequence to align the time dimension of the point cloud data; Compensate for the scanning delay error of the lidar and generate a time-synchronized point cloud dataset.

[0015] Preferably, the method further comprises optimizing the training process of the reinforcement learning algorithm, specifically: Initialize the weight parameters of the policy network; Use historical flight data sets to simulate environmental interactions and calculate policy gradient updates; The policy gradient update value is used to adjust the weight of the reward function mechanism.

[0016] Preferably, the method further includes an adaptive mechanism for handling dynamic obstacles, specifically: The changing trend of the motion speed estimate is monitored. If the changing trend exceeds a stable interval, a reassessment of the path node is triggered, and a supplementary obstacle avoidance instruction is generated. The supplementary obstacle avoidance instruction is merged into the obstacle avoidance path optimization instruction.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This method uses a LiDAR scanning module to acquire real-time point cloud data streams. LiDAR, unaffected by lighting and weather conditions, stably outputs data containing obstacle 3D position information and dynamic motion trajectories, enabling drones to more timely and comprehensive perceive their flight environment. Compared to methods that rely on visual or ultrasonic sensors, this data acquisition method reduces the impact of environmental factors on perception results and more accurately captures obstacle position changes and motion trends, providing reliable foundational data support for subsequent obstacle analysis and risk assessment.

[0018] When constructing a spatial obstacle distribution model, this method comprehensively considers both the geometry and motion of obstacles, transcending the limitations of traditional models that focus solely on static features. Geometric shape information reflects static attributes such as the obstacle's size and outline, while motion information captures dynamic characteristics such as its speed and direction. The combination of these two allows the model to more comprehensively characterize the actual presence of obstacles in space. This comprehensive model description helps more accurately assess the potential impact of obstacles on the drone's flight path, avoiding risks that could be misjudged due to incomplete descriptions of obstacle status.

[0019] The use of a reinforcement learning algorithm to predict collision risk values ​​fully leverages the algorithm's ability to learn and adapt in dynamic environments. Reinforcement learning continuously adjusts risk assessment strategies and parameters based on real-time data from the obstacle spatial distribution model, enabling collision risk calculations to dynamically adapt to environmental changes. In complex multi-obstacle scenarios or when obstacles suddenly change in motion, the algorithm rapidly responds and updates risk assessment results, ensuring the risk value accurately reflects the potential for collision in the current flight environment and providing timely decision-making for path optimization.

[0020] Obstacle avoidance path optimization commands generated based on collision risk values ​​can specifically adjust the drone's flight direction and speed. These adjustments aren't based on a fixed obstacle avoidance pattern. Instead, they incorporate the real-time movement of obstacles and the changing trends in collision risk, making the drone's obstacle avoidance more flexible and targeted. For example, when an obstacle approaches rapidly, commands can adjust flight direction to avoid the collision path. When an obstacle moves slowly, commands can optimize flight speed to reduce unnecessary detours, improving flight efficiency while ensuring safety.

[0021] After executing the obstacle avoidance path optimization command, the flight trajectory feedback data is recorded in real time and used to update the obstacle space distribution model parameters, forming a closed-loop dynamic optimization mechanism. The flight trajectory feedback data contains environmental interaction information during the drone's actual obstacle avoidance process, which can reflect the deviations and shortcomings of the model in practical applications. By incorporating feedback data into the model parameter update process, the obstacle space distribution model can continuously learn new environmental characteristics and obstacle behavior patterns, continuously optimizing the accuracy of obstacle descriptions and risk prediction capabilities. This continuous update mechanism enables the system to adapt to different flight scenarios, whether it is urban buildings, natural terrain, or areas with dense dynamic obstacles, maintaining good obstacle avoidance performance, thereby enhancing the stability and safety of the drone's autonomous flight in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a working principle diagram of the laser radar-based UAV autonomous obstacle avoidance optimization method described in the present invention; Figure 2 Flowchart constructed for the obstacle spatial distribution model; Figure 3 This is the cluster analysis diagram of the lidar point cloud; Figure 4 Flowchart generated for obstacle avoidance path optimization instructions; Figure 5 Path planning and collision risk assessment map; Figure 6 Flowchart for updating obstacle spatial distribution model parameters. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] See also Figure 1The present invention provides a laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles, the method comprising: The LiDAR scanning module continuously collects a point cloud data stream from the drone's flight environment. This data stream contains the three-dimensional position information and dynamic motion trajectories of obstacles. Based on this data stream, an obstacle spatial distribution model is constructed to characterize the geometric shape and motion state of obstacles. A reinforcement learning algorithm is used to predict the collision risk between the drone and the obstacle. This risk is calculated from the model's coordinate position and motion trajectory. Based on the risk value, obstacle avoidance path optimization instructions are generated to adjust the drone's flight direction and speed. These instructions are executed to control the drone's drive system and record flight trajectory feedback data in real time. The parameters of the obstacle spatial distribution model are updated based on this feedback data, achieving adaptive optimization.

[0025] Example 1: See Figure 2 The configuration and data processing of the lidar scanning module are the basic links of the drone's autonomous obstacle avoidance optimization method. This module uses high-precision sensors to collect point cloud data streams in the drone's flight environment in real time, providing raw input for subsequent obstacle detection and path planning. The scanning frequency of the lidar sensor is set to an adjustable range of 50Hz to 200Hz to adapt to the environmental perception needs at different flight speeds. The scanning angle covers 360 degrees horizontally, and the vertical field of view is set to 30 degrees to ensure all-round monitoring of the space around the drone. The resolution parameters are adjusted according to the actual application scenario. The typical configuration is a 0.1-degree angle step, and a lateral resolution of approximately 1.7 cm can be achieved at a distance of 10 meters.

[0026] After collecting the raw point cloud data, preprocessing is required to eliminate environmental noise. A Gaussian filter algorithm is used to smooth the point cloud data, using a filter window size of 3×3 pixels and a standard deviation parameter of 0.5. This process effectively suppresses outliers caused by factors such as atmospheric scattering and sensor noise, while preserving the geometric characteristics of real obstacles. The filtered point cloud data is unified into a coordinate system centered on the drone body through a coordinate conversion module. This coordinate system defines the X-axis as the drone's forward direction, the Y-axis as the right direction, and the Z-axis as the vertical downward direction. The 3D coordinate values ​​of each point are stored in floating-point format with millimeter-level accuracy.

[0027] The standardized three-dimensional coordinate sequence enters the spatial grid division stage. The monitoring space is divided into uniform cubic grid units, and the unit size is set to 10 cm × 10 cm × 10 cm. This size selection takes into account both computational efficiency and environmental representation accuracy, and can accurately describe the spatial distribution of obstacles while ensuring real-time performance. Each grid unit maintains an independent point cloud data index to facilitate subsequent clustering analysis and processing. After the grid division is completed, the system performs density clustering analysis on the point cloud within each unit. Clustering is performed using the improved DBSCAN algorithm, which can automatically identify point cloud clusters with spatial continuity while filtering out discrete noise points. The algorithm parameters are set to a neighborhood search radius of 0.5 meters, and the minimum point number threshold for forming a cluster is 5 points.

[0028] Each point cloud cluster output by cluster analysis represents a potential obstacle entity. Geometric features are extracted from these point cloud clusters to generate a description vector for the obstacle. The description vector contains three parts: size features, surface curvature features, and boundary features. Size features are obtained by calculating the maximum span of the point cloud cluster in the X, Y, and Z directions; surface curvature features are obtained by fitting a local plane to the point cloud and calculating the rate of change of the normal vector; boundary features are described by extracting the convex hull vertices of the point cloud cluster. At the same time, the movement speed of the same obstacle is calculated by comparing the position changes of two adjacent frames of point cloud data. Velocity estimation uses the least squares method to fit the displacement-time relationship and output a three-dimensional velocity vector.

[0029] The clustering results of all grid cells are integrated to form a complete obstacle spatial distribution model. This integration process includes establishing a global spatial index and topological relationships. Each obstacle entity is assigned a unique identifier and its grid cell location is recorded. The model uses an octree data structure to enable fast spatial queries and updates. For dynamic obstacles, the model also maintains historical trajectory data to predict their future locations.

[0030] The obstacle spatial distribution model is updated synchronously with the LiDAR scans. Each time new scan data arrives, the system first matches the obstacle locations predicted by the current model with the actual detection results. Successfully matched obstacles have their geometry and motion updated; unmatched predictions are considered lost; and newly detected point cloud clusters are added to the model as new obstacles. This dynamic update mechanism enables the model to accurately reflect real-time changes in the environment.

[0031] During the model building process, the system employs a multi-level verification mechanism to ensure data reliability. Primary verification checks the geometric plausibility of point cloud clusters, filtering out detection results that do not conform to physical laws. Intermediate verification assesses the authenticity of obstacles by tracking consistency across multiple frames. Advanced verification, based on the drone's motion state, assesses the consistency between detection results and inertial measurement data. Only obstacles that pass all verification steps are included in the final spatial distribution model.

[0032] Temporal and spatial alignment of LiDAR data is crucial for model accuracy. The system uses a hardware timestamp synchronization mechanism to ensure strict correspondence between point cloud data and the drone's pose information. A motion compensation algorithm corrects for point cloud distortion caused by the drone's motion during scanning. This compensation algorithm utilizes high-frequency inertial measurement unit data to reconstruct the precise acquisition time and spatial position of each LiDAR scan point.

[0033] The output of the obstacle spatial distribution model provides environmental perception data for the subsequent path planning module. The model not only contains the spatial location information of static obstacles but also accurately describes the motion state of dynamic obstacles. For each obstacle, the model provides complete information, including its geometric dimensions, surface shape, velocity, and predicted trajectory. This data is organized in a structured format, supporting efficient spatial queries and collision detection operations.

[0034] The point cloud data processing process is accelerated using a parallel computing architecture. LiDAR data reception, filtering, and coordinate conversion are performed on dedicated hardware processing units to ensure real-time performance. Computationally intensive tasks such as spatial meshing and cluster analysis are distributed to multi-core processors for parallel processing. The system employs a task pipeline design, with data transferred between different processing stages via a ring buffer to minimize processing latency.

[0035] The accuracy of environmental modeling directly impacts the quality of obstacle avoidance decisions. The system optimizes modeling through an adaptive parameter adjustment mechanism. When complex scenes are detected, the system automatically reduces the grid cell size and cluster radius to improve modeling accuracy. In open environments, these parameters are appropriately increased to reduce computational load. This adaptive mechanism enables the system to maintain optimal modeling performance in a variety of flight environments.

[0036] Calibration and verification of lidar sensors are essential for ensuring data quality. During startup, the system automatically performs a sensor calibration process, including internal parameter calibration and external mounting position calibration. This process utilizes targets with known geometric features to optimize calibration parameters by minimizing reprojection errors. During daily operation, the system continuously monitors sensor status and indicates maintenance requirements when performance degradation is detected.

[0037] The obstacle spatial distribution model also features incremental update capabilities. For continuously tracked obstacles, the model uses a Kalman filter algorithm to smooth their trajectory and reduce the impact of measurement noise. The filter's state vector contains position, velocity, and acceleration information, and the system dynamically adjusts process noise parameters to accommodate different motion patterns. This process significantly improves the accuracy of dynamic obstacle state estimation.

[0038] The effective use of point cloud data is also reflected in the recognition of unique scenarios. By analyzing point cloud distribution patterns, the system can identify obstacles such as glass curtain walls and wire mesh that are difficult for traditional sensors to detect. For these special obstacles, the model annotates their unique attributes, enabling the decision-making module to perform targeted processing. The system can also identify ground features and ceiling structures, providing a three-dimensional spatial positioning reference for the drone.

[0039] The entire environmental modeling process forms a closed-loop quality control mechanism. The system monitors the output metrics of each processing step in real time and automatically triggers data reprocessing or parameter adjustments when anomalies are detected. Quality control metrics include dozens of dimensions, such as point cloud density, clustering success rate, and trajectory continuity, ensuring the reliability of modeling results. This closed-loop mechanism enables the system to adapt to operational needs in a variety of complex environmental conditions.

[0040] See Figure 3 , which shows the entire process of lidar point cloud processing, including four key analysis views: Raw Point Cloud Data Distribution: Displays the raw point cloud collected by the LiDAR, including ground points, obstacle points, and noise points. Point Cloud Clustering and Boundary Recognition: Shows the results of DBSCAN clustering, with different grayscales representing different obstacles. 3D Point Cloud Distribution: Shows the spatial distribution of the point cloud from a 3D perspective, including height information. Obstacle Spatial Modeling and Motion Trajectory: Displays the meshed obstacle modeling results and the motion trajectory of dynamic obstacles. This demonstrates the system's ability to identify static and dynamic obstacles in complex environments, providing an accurate environmental model for autonomous obstacle avoidance.

[0041] Example 2: See Figure 4 The collision risk prediction process of the reinforcement learning algorithm receives an obstacle description vector from the environment modeling module as its core input. This description vector contains geometric features and motion state data. The geometric features include size, surface curvature, and boundary vertex coordinates, while the motion state includes a three-dimensional velocity vector and acceleration estimates. The system normalizes these features into a numerical sequence of uniform dimensions and combines them to form a multidimensional state vector. The dimensionality of the state vector is dynamically adjusted based on the complexity of the scene, with a typical configuration containing 15 to 30 feature dimensions.

[0042] The Deep Q-Network framework constructs a multi-layer neural network to process state inputs. The network structure adopts a fully connected design, with the number of input layer nodes matching the dimensionality of the state vector. The intermediate hidden layer has two processing layers, each containing 64 neurons, using the ReLU nonlinear activation function. The number of nodes in the output layer corresponds to the discretized action space, which includes eight basic flight directions and four speed levels, for a total of 32 possible action combinations. The neural network weights are initialized using a randomized small value strategy to avoid prediction bias in the initial state.

[0043] The reward function is designed based on the spatial relationship between the drone and the obstacle. The system calculates the minimum Euclidean distance from each obstacle surface to the drone's hull in real time. The reward value is set with a negative gradient mechanism: a high penalty is imposed when the minimum distance falls below 0.5 meters; a linearly decreasing penalty is set between 0.5 and 2 meters; and no reward or penalty is applied when the distance exceeds 5 meters. An additional time penalty factor is introduced, applying a small penalty for each decision cycle in which the target point is not reached. The reward signal is transmitted to the network weight update process using a temporal difference algorithm.

[0044] Policy network training utilizes an experience replay mechanism. The system establishes a fixed-capacity memory bank to store state transition records. Each record contains the current state, the action taken, the reward received, and the state after the transition. A batch of historical records is randomly sampled during each network update to mitigate the influence of temporal correlations between data. Network weights are adjusted based on the Bellman optimality equation to calculate the target Q value. The difference between the target value and the predicted value is calculated using the mean squared error to calculate the loss function. The optimization algorithm uses an adaptive moment estimation method to adjust the learning rate and control the amplitude of parameter updates.

[0045] The collision risk value is output as a probability distribution. The system calculates the Q-values ​​corresponding to 32 action options and converts them into a probability distribution using a Softmax function. These probability values ​​are then mapped to a three-dimensional physical space to form a risk heat map. The heat map divides the space into regions centered around the drone's current location. High-risk areas are defined as areas in physical space where the probability peak exceeds a preset threshold. The system then marks the boundary coordinates of these regions and the direction of the risk gradient.

[0046] The path planning module activates a dynamic obstacle avoidance algorithm based on the risk heat map. First, a three-dimensional path search space is constructed, discretized into a grid of cubes with sides of 1 meter. The search starts at the drone's current location and ends at the pre-set flight target. The system assesses the cost of each grid node, a composite of the base path length, the risk heat value, and the energy consumption. The risk heat value is weighted dynamically based on the peak value of the heat map; higher peak values ​​are weighted more heavily.

[0047] The search algorithm operates in a three-dimensional grid space. It maintains two sets of nodes, open and closed, and iteratively expands neighboring nodes starting from the starting point. A heuristic function, using the Euclidean distance metric, guides the search toward the target. A node selection mechanism comprehensively evaluates the actual cost from the starting point to the node and the estimated cost to the end point, prioritizing nodes with the lowest total cost. Grid cells marked as high-risk areas are skipped during the search. When necessary to traverse high-risk areas, the path node density is automatically increased to improve accuracy.

[0048] After candidate paths are generated, the feasibility assessment phase begins. The system performs segment-by-segment verification on each candidate path: calculating whether the radius of curvature at each turning point meets the drone's maneuverability requirements; verifying whether the path's altitude exceeds flight airspace restrictions; and checking whether the path's total length is within the available energy resources. Each path segment is scored on a 100-point scale, with path safety accounting for 40%, energy efficiency for 35%, and time cost for 25%. The scoring process considers the probability of the spatiotemporal intersection of the predicted trajectory of dynamic obstacles with the path's timeline.

[0049] Once the optimal path is determined, it's converted into specific control commands. The system calculates the direction vector for the initial segment of the path and converts it into a yaw angle relative to the drone's nose, with an output accuracy of 0.1 degrees. Speed ​​control is automatically adjusted based on the path curvature: straight segments use the maximum cruising speed, while turns are limited by the curvature radius. The resulting commands include instantaneous azimuth adjustment and acceleration commands. The azimuth adjustment range covers an omnidirectional range of -180 degrees to +180 degrees, and the speed adjustment range is continuously variable between 0 and 5 meters per second squared.

[0050] The real-time control layer uses a sliding window mechanism to execute the path. The system continuously tracks the deviation between the drone's actual position and the planned path, maintaining a 2-meter tolerance threshold. If the deviation exceeds the threshold or an unexpected obstacle is detected, the path planning module interrupts the current command sequence and immediately restarts the risk assessment and path search process. Under normal conditions, the system refreshes the control instructions for the next path segment every 200 milliseconds, achieving high-frequency dynamic response.

[0051] A regular update mechanism for the reinforcement learning policy network supports online learning. The system collects state-action-reward data sequences from actual flights and replays them in batches during nighttime maintenance. A conservative learning rate parameter is set for network weight updates, with each update not exceeding 1% of the original weights to maintain a smooth transition in policy evolution. New training data undergoes multiple validation steps: verifying the consistency of state data with the environmental model; verifying that action selections meet physical constraints; and confirming that reward calculations are consistent with real-time monitoring results.

[0052] The accuracy of collision risk predictions is monitored using three-dimensional trajectory reconstruction technology. The system records the complete flight path after each obstacle avoidance decision and, in an offline phase, reprojects the path onto a spatial obstacle distribution model to verify that the actual flight path meets the expected minimum distance between the obstacle surface and the actual flight path. In cases where predictions deviate significantly, the system reverse-engineers the neural network's processing of the obstacle's feature vectors to pinpoint the root cause of the perception misjudgment or decision-making flaw.

[0053] The path optimization module's specialized scenario processing capabilities cover common flight conditions. For dense obstacle clusters, the system automatically switches its search strategy to random sampling, uniformly generating path samples within the feasible space instead of a global search. For narrow corridors, position control accuracy requirements are increased and flight speed is reduced. For high-speed moving obstacles, predicted position interpolation technology is used to construct a time-space composite risk model. Each scenario is identified based on a specific combination of obstacle distribution characteristics and motion patterns, triggering the corresponding parameter configuration.

[0054] The control command output format is compatible with various UAV platforms. The azimuth angle adjustment value conforms to the standard heading angle definition and is consistent with the heading reference of mainstream flight control systems. Speed ​​commands provide both relative speed change and absolute speed output modes, adapting to different manufacturers' interface protocols. The system is compatible with both pulse-width modulation and serial data bus physical transmission methods, with a data transmission cycle of less than 50 milliseconds, meeting the real-time control requirements of highly dynamic environments.

[0055] The decision-making system's state monitoring incorporates a multi-layered exception handling mechanism. During operation, the system continuously monitors input data validity: confirming that the environmental model is regularly updated, verifying that the number of obstacles is within a reasonable range, and monitoring whether the neural network's output probability distribution satisfies normalization requirements. Upon detecting an abnormal interruption in the decision-making module's input data, the system immediately switches to a simplified obstacle avoidance strategy: forward flight is halted, a rotational scan is performed, and the main path planning process is re-entered once environmental data continuity is restored.

[0056] See Figure 5 , showing the results of collision risk assessment and path planning based on reinforcement learning: Collision risk heat map and path planning: showing the collision risk distribution in the environment (the darker the color, the higher the risk) and the planned safe path; Three-dimensional path planning view: showing the relationship between the planned path of the drone and the risk terrain from a three-dimensional perspective, verifying the effectiveness of the reinforcement learning algorithm in real-time risk assessment and path planning in complex dynamic environments.

[0057] Example 3: See Figure 6The execution of the obstacle avoidance path optimization instruction begins with the conversion process of the control signal. The azimuth angle adjustment value is input into the proportional-integral-differential controller, which contains three independent channels to process the pitch angle, roll angle, and yaw angle respectively. The proportional link directly responds to the angle deviation, the integral link accumulates historical deviations, and the differential link predicts trend changes. The current adjustment value corresponds to the speed change input and is mapped to the motor control quantity through the voltage-thrust curve. The controller outputs two signals: the steering channel outputs a pulse width modulation waveform with a duty cycle proportional to the angle correction value; the speed channel outputs an analog voltage signal with an amplitude that matches the acceleration requirement.

[0058] The drive interface module uses a serial communication protocol to transmit control signals. The protocol frame structure consists of a start bit, address code, data field, and check bit. The data field is divided into 32 bytes, storing a 16-bit pulse width value and a 16-bit voltage value, respectively. The transmission baud rate is fixed at 115200 bps, with a complete command frame transmitted every 50 milliseconds. The interface hardware is electrically isolated to prevent electromagnetic interference from the thrusters from affecting the control core. The module also has a built-in response detection mechanism, triggering redundant retransmissions if no device confirmation is received.

[0059] The drone's propulsion system receives commands and executes them. Four brushless motors receive their own pulse-width modulated signals, which are converted into three-phase drive currents by electronic speed regulators. The motor speed change rate is limited to 2000 revolutions per second to prevent mechanical shock. Propeller response data is collected synchronously: Each motor is equipped with a Hall effect sensor to measure actual speed, a current sensor to monitor phase load, and a temperature sensor to record operating status. Response data is sampled with 12-bit accuracy and a sampling frequency of 1kHz and transmitted back in real time via a dedicated data bus.

[0060] The position tracking system operates independently of the feedback loop. The onboard high-precision global positioning system module, model MS-1521, continuously outputs three-dimensional coordinates with a 10Hz update frequency and a horizontal error radius of 0.1 meter. An inertial measurement unit (IMU) assists in motion compensation and includes a three-axis micro-electromechanical (MEMS) gyroscope and accelerometer, with a data fusion frequency of 200Hz. Real-time position data is integrated with multiple sources through a Kalman filter to output the center coordinates and attitude angles of the aircraft, with attitude measurement accuracy reaching 0.1 degrees. Thruster response data is combined with real-time position data into a flight trajectory feedback packet, with the transmission cycle synchronized with the control system clock.

[0061] The parameter update of the obstacle spatial distribution model is triggered by the trajectory deviation. The system compares the spatial difference between the actual flight position and the predicted path node and calculates the deviation value using the three-dimensional Euclidean distance:

[0062] in, Indicates the position deviation value, are the actual coordinates, is the predicted node coordinate. When three consecutive sampling cycles The model reconstruction process starts when the deviation threshold is less than 0.3 meters. The deviation threshold setting has dynamic adjustment capabilities and can be reduced to 0.3 meters in complex environments.

[0063] The spatial grid cell division rules are adjusted using a gradual refinement strategy. The initial grid is divided from 10 cm cubes into 5 cm cells, with a segmentation depth of three or fewer. Each subdivided grid inherits the point cloud index of its parent cell, and four computational threads are concurrently launched for parallel processing. An automatic grid density balancing algorithm allocates computing resources based on point cloud distribution characteristics: the minimum cell size is used in areas with dense obstacles, while a coarse-grained division is retained in open airspace. The grid topology maintains spatial continuity by updating the adjacency table in real time.

[0064] The algorithm parameters for point cloud clustering analysis have been upgraded during the re-execution process. The neighborhood search radius ε has been reduced from the baseline value of 0.5 meters to 0.3 meters, and the minimum number of cluster points has been increased from 5 to 10. A normal vector consistency check has been introduced, requiring that the angle between point cloud normals within the same cluster be less than 30 degrees. Boundary integrity checks have been added to clustering results to exclude point cloud clusters with internal holes. The generation of new geometric shape description vectors strengthens surface continuity features, adding a new surface concavity indicator based on the original size and curvature characteristics. This indicator is achieved by detecting the rate of change in point cloud density through ray-guided inspection.

[0065] A version management mechanism has been established for the updated obstacle spatial distribution model. Each model snapshot records the timestamp and parameter configuration, with historical versions stored up to the last five iterations. Model switching utilizes a smooth transition algorithm to avoid sudden changes in geometric descriptions: the new model weight increases linearly from 0.2 to 1.0 with a transition period of 200 milliseconds. When reconstructing the motion trajectory of dynamic obstacles, velocity trend data from historical models is integrated, and sudden acceleration changes are smoothed using second-order derivatives.

[0066] The model accuracy verification system operates in parallel. Virtual detection points are evenly distributed across the model space. For each point, the theoretical distance to the nearest obstacle surface is calculated and compared to the actual value from the point cloud match. The verification report automatically generates a defect heat map, highlighting areas where spatial position deviations exceed 10%. These outlier areas are prioritized for analysis in the next round of data processing, with repeated clustering performed by increasing the grid density to 1 cm accuracy.

[0067] The control execution process utilizes a closed-loop monitoring architecture. An independent actuator status watchdog circuit monitors the delay between motor control signals and response data. Dynamic gain adjustment of the control loop is triggered when the time difference between the steering command and the vehicle's attitude change exceeds 80 milliseconds, or when the acceleration command and actual speed do not converge for 0.3 seconds. This adjustment automatically updates the integral time constant of the proportional-integral-derivative controller based on real-time load inertia.

[0068] Adaptive update cycle management utilizes a dual-time-base system. The base update cycle is locked to the LiDAR scan interval, forcing synchronization with data acquisition. The emergency update channel features preemptive capabilities. When continuous position deviation alarms or sudden obstacle warnings occur, regular processing is interrupted to initiate a priority update. All model updates utilize a resource pre-allocation strategy, distributing computational tasks across four processing cores for balanced load, reducing single update times to less than 8 milliseconds.

[0069] A prediction-correction mechanism maintains the spatiotemporal consistency of model data. After each model update, the extrapolated trajectory of each obstacle is reconstructed. Trajectory predictions are based on the current velocity vector and compensate for the drone's acceleration disturbances collected by the inertial measurement unit. Prediction results are stored in a spatial index tree and used for collision pre-checks before the next decision cycle. If a potential collision is detected during pre-checks, path replanning is initiated in advance, forming a closed-loop response chain from model update to decision optimization.

[0070] Real-time performance monitoring of the point cloud processing system includes 22 metrics. Key indicators include mesh processing latency, clustering completion rate, and feature extraction errors. These metrics are recorded in a ring buffer over the last 100 cycles, automatically generating an operating baseline. When a single-cycle metric deviates by 30% from the baseline, a diagnostic analysis process is triggered: hardware interrupt response time is checked, memory bandwidth utilization is verified, and algorithm branch prediction failure rates are analyzed to generate optimization parameter adjustment recommendations.

[0071] A coordinated mechanism for controlling commands and model updates prioritizes them. Model updates receive medium-priority computing resources during normal flight; control command execution during obstacle avoidance maneuvers receives the highest priority; and model reconstruction tasks are restricted to idle processor periods. The resource allocation manager implements a hard real-time scheduling strategy, ensuring critical control cycles adhere to a precise 50-millisecond time window.

[0072] Example 4: During the preprocessing phase of LiDAR point cloud data, time synchronization is crucial for ensuring accurate environmental modeling. When the system acquires raw point cloud data from the LiDAR hardware interface, each point cloud frame is accompanied by a microsecond-accurate timestamp. This timestamp is generated based on the hardware clock used by the Global Positioning System (GPS) and synchronized with the drone's inertial navigation system. Time alignment of point cloud data utilizes a three-level buffering mechanism: a raw data buffer queue stores unprocessed point cloud frames; a time calibration module annotates each frame with an offset relative to the system's unified time base; and a synchronized output queue reorganizes the point cloud sequence in chronological order.

[0073] Scanning delay compensation for lidar requires precise measurement of the delay characteristics of each link. The system maintains a delay parameter table that records the delay components of the entire link, from laser emission to data output. This table includes key parameters such as internal sensor processing delay, data transmission delay, and coordinate transformation calculation delay. Each delay component is measured through specialized calibration experiments, and compensation curves are fitted at different operating temperatures. Typical delay parameters are shown in Table 1.

[0074] Table 1: Typical delay parameters.

[0075]

[0076] The time-synchronized point cloud dataset enters the reinforcement learning training process. The policy network's weights are initialized using a layered randomization strategy: input layer weights follow a uniform distribution, hidden layer weights are initialized using a truncated normal distribution, and output layer weights are set to small random values. The network's input layer node count dynamically matches the state vector dimensions, and the hidden layers use residual connections to prevent vanishing gradients. The historical flight dataset is constructed from millions of state-action-reward records, each of which stores a 128-dimensional state vector, a 32-dimensional action selection vector, and a scalar reward value.

[0077] Interactive training in the simulated environment utilizes an asynchronous parallel architecture. Eight compute nodes simultaneously run the environment simulator, with each simulator instance loaded with different scenario configuration parameters. Scenario parameters include variables such as obstacle density, velocity range, and spatial distribution pattern. During training, the simulator generates state transition data that is passed to the learning algorithm via a shared memory queue. Policy gradient updates are calculated using a sliding window averaging method, with samples drawn from the most recent 1000 training batches.

[0078] The reward function mechanism's weight adjustments are based on policy performance evaluation. The system defines ten evaluation metrics to monitor policy evolution, including average obstacle avoidance distance, path efficiency index, and motion smoothness. After every 500 training iterations, the validation module runs the current policy on an independent test set and generates a policy performance report. The weight adjustment algorithm analyzes the contribution of each reward component to the final policy and dynamically adjusts the coefficient ratios of different reward components. An early stopping mechanism is implemented during training, automatically terminating the current training phase after three consecutive validation cycles without significant improvement.

[0079] The network weight update process implements strict numerical stability controls. The norm of the gradient value is checked before each parameter update, and gradient clipping is automatically enabled if it exceeds a threshold. The learning rate uses a cosine annealing schedule, with an initial moderate value that decays gradually as training progresses. Network parameters are saved using a rolling backup mechanism, retaining the last ten training checkpoints, allowing for rapid recovery to a previously stable version in the event of performance regressions.

[0080] Training data quality monitoring involves multiple verification steps. Raw state data must pass rationality checks: checking that the value range is within the sensor's measurement range; verifying the physical consistency of the data across all dimensions; and confirming the continuity of the time series. Action selection records must be reviewed for compliance with the drone's dynamic constraints, including maximum steering angular velocity and acceleration limits. Reward calculations must be rechecked against the raw observation data to ensure a precise correspondence with the environmental conditions at the time of real-time decision-making.

[0081] The simulation environment configuration emphasizes scene diversity and physical realism. The basic scene library includes twenty typical environment templates, covering diverse spatial structures such as open airspace, urban canyons, and indoor corridors. Each template can be adjusted to generate scene variations based on ten parameters, including obstacle material properties, lighting conditions, and atmospheric disturbance intensity. The physics engine accurately simulates the aerodynamic characteristics of the drone, taking into account parameters such as body size, weight distribution, and thrust curve.

[0082] Resource management during the training process utilizes a dynamic allocation strategy. Compute nodes allocate GPU resources based on task priority. Core training tasks are exclusively allocated to high-performance computing cards, while data preprocessing tasks share mid-range computing units. Memory usage is managed using paging, with frequently accessed training data resident in memory and historical checkpoints stored on high-speed solid-state drives. Network communication utilizes a hybrid protocol, with remote direct memory access (RDMA) technology used for large-scale data transmission. Control commands are exchanged quickly via the User Datagram Protocol (UDP).

[0083] Pre-deployment validation testing includes comprehensive functional checks. Thousands of randomized scenarios are run in a simulation environment, recording key metrics such as collision rate, path deviation, and control stability. Hardware-in-the-loop testing deploys the policy network to actual flight control hardware, simulating sensor inputs through signal injection to verify real-time responsiveness. The final deployment package generates differentiated policy versions, automatically selecting the appropriate network structure simplification solution for hardware platforms with varying computing power.

[0084] The online learning mechanism is designed with practical operational constraints in mind. Newly collected flight data is rigorously screened before being added to the training pool, with priority given to samples containing rare scenarios. The incremental training process limits computing resource utilization to ensure it does not impact real-time control tasks. Policy updates utilize a grayscale release model, piloting new policies on a limited number of drone nodes before verifying their stability and rolling them out to the public. A version rollback function remains readily available, allowing for a quick switch to the previous stable version if performance degradation is detected.

[0085] Continuous monitoring of time synchronization accuracy utilizes a hardware-assisted approach. A high-precision time interval meter is installed on the drone body to directly measure the time offset between lidar data acquisition and inertial measurement unit sampling. The measurement results are fed back to the time calibration module, dynamically adjusting compensation parameters. A temperature sensor network monitors the operating temperature of each hardware module in real time, predicting latency trends based on temperature-delay characteristic curves.

[0086] Spatial calibration and temporal synchronization of point cloud data are performed in conjunction. Each point cloud frame includes aircraft pose data, and a coordinate transformation matrix is ​​used to transform the raw point cloud into a unified world coordinate system. The transformation matrix calculation takes into account temporal interpolation to accurately match the drone's pose at the time of point cloud acquisition. The calibrated point cloud data is stored in an octree structure, supporting efficient spatial range queries and nearest neighbor searches. Data is stored in a compressed format, minimizing memory usage while maintaining accuracy.

[0087] Reproducibility of the reinforcement learning training environment is achieved through seed control. The random number generator seed value is fixed for each training cycle, ensuring consistent training trajectories for identical inputs. The simulator run configuration is saved as a description file, recording all parameter settings that affect policy performance. The training log contains complete snapshots of the environment state and decision records, enabling precise replay and analysis of any training step. This design facilitates identifying policy flaws and optimizing the training process.

[0088] Example 5: The operating basis of the dynamic obstacle adaptation mechanism is to continuously monitor the changing characteristics of the obstacle's motion state. The system periodically extracts the three-dimensional velocity vector data of each identified obstacle from the obstacle space distribution model. The sampling interval is strictly synchronized with the lidar scanning period, with a typical value of 10 milliseconds. The velocity estimation sequence of each obstacle is first processed to remove outliers: the moving standard deviation of five consecutive sampling points is calculated, and the instantaneous data exceeding three times the standard deviation is regarded as noise interference and discarded, and linear interpolation is used to fill the missing values. The preprocessed velocity sequence is input into the sliding average filter, the filter window size is fixed to five sampling periods, and the smoothed velocity trend curve is output.

[0089] A first-order difference algorithm is used to quantitatively analyze speed trends. The system calculates the rate of change of the vector modulus of two adjacent filtered speed values ​​and takes the weighted average of three consecutive difference values ​​as the final trend indicator. When the absolute value of this indicator exceeds 2 m / s², the system is deemed unstable and the dynamic response process is triggered. The decision logic uses a two-stage confirmation mechanism: a 100-ms observation period is initiated after the initial limit violation, during which three consecutive samples exceed the limit before the final state transition is confirmed. Specially marked volatile obstacles (such as flying birds or swaying objects) use independent decision thresholds, with the rate of change threshold relaxed to 4 m / s².

[0090] The path node reassessment process immediately interrupts the current path tracking process. The system saves all node data for the current path plan to a buffer stack, freezing the drone in hover at its current position (altitude error controlled to ±0.3 meters). When the dynamic path planning algorithm is reactivated, the updated spatial distribution model is loaded, focusing on expanding the search space around obstacles with sudden speed changes. High-risk area weighting is adjusted using an incremental overlay strategy: a 20% weight coefficient is added to the original weight. The adjustment range is generated in concentric spheres centered around the obstacle, with the high-risk core area within 1 meter of the obstacle and the buffer transition area between 1 and 3 meters. The search grid density in the core area is tripled, and the grid edge length is reduced to 0.3 meters.

[0091] The supplementary obstacle avoidance command generation mechanism includes three levels of correction logic. The primary correction calculates the intersection angle between the extended line of the obstacle's motion trend and the current heading, outputting a directional angle fine-tuning value within a range of ±5 degrees. The intermediate correction predicts the time window for the drone to pass through the danger zone, and calculates the collision time margin based on the sudden change in velocity. When the margin is less than 1 second, a speed correction command of ±0.5 m / s is triggered. The advanced correction establishes a local cost map of the optimal avoidance path for group obstacle scenarios and solves the optimal avoidance direction. The correction parameters are verified for dynamic feasibility: the directional adjustment value is checked to see if it exceeds the maximum turning rate limit; the speed adjustment value is verified to see if it is within the feasible range of motor thrust.

[0092] Instruction merging and execution adopts a priority override strategy. The main obstacle avoidance path optimization instruction queue maintains its original structure, and supplementary instructions are added through an insertion mechanism and executed before the first instruction. Smooth transitions are set between instructions: direction adjustment uses a trapezoidal acceleration and deceleration curve, limiting angular acceleration to no more than 30 degrees / second²; speed changes follow a parabolic gradient model, with the acceleration rate of change controlled within 2 meters / second³. All merged instructions are timestamped when injected into the execution system to ensure that the execution order meets the expected timeline. Execution status is monitored in real time: the direction adjustment completion accuracy standard is within ±0.5 degrees; the speed adjustment convergence tolerance is ±0.1 meters / second.

[0093] The sudden obstacle response module features cross-module collaboration. When the lidar detects a new obstacle during reassessment, the model update module immediately suspends the current cluster analysis and prioritizes feature extraction for the new obstacle. The description vector for the new obstacle is generated in under 8 milliseconds and directly inserted into the existing spatial distribution model. After obtaining the updated model, the reinforcement learning decision module recalculates the collision probability distribution only for the local state space of the affected area. This partial update mechanism keeps overall response latency under 50 milliseconds.

[0094] The dynamic scene adaptation mechanism includes scene type recognition. The system analyzes the spatial distribution of obstacles with sudden velocity changes: emergency avoidance is activated for single, sudden velocity obstacles; collaborative path planning is initiated for group-coordinated obstacles (such as flocks of birds); and an oscillation model is used to predict periodic motion obstacles. Recognition is based on five characteristic parameters: obstacle density gradient, velocity vector correlation, trajectory periodicity, spatial distribution symmetry, and historical behavior matching. Parameter preset templates are activated for each scene type, automatically configuring optimal path search constraints.

[0095] A self-healing mechanism ensures system stability. If the position deviation persists for more than 1 meter after executing supplementary instructions, or if environmental conditions trigger reassessment twice in a row, the system automatically resets to a safe state: clearing all planned path nodes, climbing to a safe altitude (set to 20 meters above the current position), and initiating a 360-degree omnidirectional environmental scan. After updating the model with the scan data, the path planning is reinitialized and the previously failed area is marked as a special attention area. Subsequent path searches in this area adopt an ultra-conservative strategy, with an additional 50% risk assessment weighting.

[0096] Dynamic quota management is implemented for operational resource allocation. Under normal circumstances, dynamic obstacle handling tasks consume 15% of computing resources; this quota increases to 50% when a response is triggered, achieved by suspending non-critical background tasks (such as historical data archiving and redundancy check calculations). Processing tasks are divided into four concurrent sub-threads: trend monitoring and threshold determination, path space reparameterization, instruction correction generation, and execution status tracking. Thread priorities are dynamically adjusted at the millisecond level, prioritizing latency-sensitive tasks.

[0097] The historical behavior learning system optimizes response parameters. After each response event, an event report is automatically generated, documenting the triggering conditions, response strategy, and actual execution results. The event database utilizes a multidimensional index structure, storing data in clusters based on obstacle movement patterns, environmental complexity, and response effectiveness. Once sufficient samples have been accumulated, the offline analysis module analyzes the correlation between parameter settings and obstacle avoidance success rates, automatically optimizing the judgment threshold and correction range. New parameter configurations are verified through shadow testing and then incrementally updated to the online system.

[0098] Hardware acceleration modules improve response speed. A programmable gate array (FPGA) is configured with dedicated circuitry for motion trend prediction. A differential calculation unit supports parallel processing of six sets of velocity data, while threshold decision logic achieves nanosecond response. Path search space generation is accelerated using a graphics processor (GPU). Meshing calculations utilize a parallel rasterization architecture, and cost map construction utilizes 256 parallel processing cores. These hardware modules compress environmental dynamic response times to milliseconds.

[0099] A staged handover strategy is used to restore normal operations. After three consecutive sampling periods, when the obstacle's speed trend returns to a stable range, the system assesses the current route's risk. If the risk value is below 0.05, the system switches directly back to the primary path. For medium-risk conditions (0.05-0.2), two or three transition paths are generated, tested for stability, and then the primary path is returned. For high-risk conditions, supplementary mechanisms are retained until the danger zone is cleared. Track smoothness is monitored throughout the handover process, with heading angle deviation controlled within 0.5 degrees per meter of path.

[0100] Mechanism performance is monitored in a full-cycle closed-loop manner. The state tracker maintains 40 operational indicators, ranging from motion trend false alarm rate to command merging delay time, from path replanning frequency to obstacle avoidance maneuver completion rate. Indicator data is streamed to form an operational baseline, and anomaly warnings are implemented based on statistical control chart principles. The maintenance and diagnostic interface provides deep debugging capabilities, allowing for the reproduction of system response sequences within any time window for fault analysis.

[0101] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0102] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A laser radar-based autonomous obstacle avoidance optimization method for UAVs, characterized in that: include: The laser radar scanning module is used to obtain a real-time point cloud data stream in the UAV flight environment, wherein the real-time point cloud data stream includes the three-dimensional position information and dynamic motion trajectory of obstacles; Building an obstacle spatial distribution model based on the real-time point cloud data stream, wherein the obstacle spatial distribution model is used to characterize the geometric shape and motion state of the obstacle; A reinforcement learning algorithm is used to predict a collision risk value between the drone and an obstacle, where the collision risk value is calculated based on the position information and motion trajectory in the obstacle spatial distribution model; generating an obstacle avoidance path optimization instruction according to the collision risk value, wherein the obstacle avoidance path optimization instruction is used to adjust the flight direction and speed of the UAV; Executing the obstacle avoidance path optimization instructions to control the UAV's drive system and record flight trajectory feedback data in real time; Parameters of the obstacle spatial distribution model are updated based on the flight trajectory feedback data.

2. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 1, characterized in that: The method of obtaining a real-time point cloud data stream in the UAV flight environment through the laser radar scanning module includes: Configure the scanning frequency and resolution parameters of the LiDAR sensor to ensure that the point cloud collection covers the preset monitoring area; Perform noise filtering on the original point cloud data to remove environmental interference signals; Convert the filtered point cloud data into a standardized 3D coordinate sequence.

3. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 2, characterized in that: The constructing of the obstacle spatial distribution model comprises: Dividing the three-dimensional coordinate sequence into a plurality of spatial grid units; Perform point cloud clustering analysis within each spatial grid cell to identify obstacle boundary features; Generate obstacle geometry description vector and motion speed estimation based on clustering results; The outputs of all spatial grid cells are integrated to form the complete structure of the obstacle spatial distribution model.

4. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 3, characterized in that: The method of using a reinforcement learning algorithm to predict the collision risk value between a drone and an obstacle includes: Inputting the obstacle geometry description vector and motion velocity estimate into the state space of the reinforcement learning algorithm; Design a reward function mechanism based on the relative distance between the drone's current position and the obstacle; The reinforcement learning policy network is updated through iterative training to output the probability distribution of the collision risk value.

5. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 4, characterized in that: The generating of the obstacle avoidance path optimization instruction comprises: Analyzing the probability distribution of the collision risk value to determine high-risk obstacle areas; Calculating a candidate safe flight path using a dynamic path planning algorithm, wherein the candidate path includes a plurality of path nodes; Evaluate the feasibility score of each path node and select the path solution with the highest feasibility score as the obstacle avoidance path optimization instruction; The obstacle avoidance path optimization instruction includes a specific direction angle adjustment value and a speed change.

6. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 5, characterized in that: The driving system for controlling the UAV includes: Converting the direction angle adjustment value and the speed change into a motor control signal; Sending the motor control signal to the drone propeller via the drive interface module; Synchronously collect thruster response data and real-time position data; The thruster response data and real-time position data form part of the flight trajectory feedback data.

7. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 6, characterized in that: The updating of the parameters of the obstacle space distribution model based on the flight trajectory feedback data includes: Comparing the real-time location data with the deviation value of the predicted path node; If the deviation value exceeds a preset tolerance threshold, adjusting the division rule of the spatial grid unit; Re-performing point cloud clustering analysis to update the obstacle geometric shape description vector; The updated obstacle spatial distribution model is used for the next round of collision risk value prediction.

8. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 7, characterized in that: The method further includes performing data preprocessing before constructing the obstacle spatial distribution model: Extracting timestamp information from the three-dimensional coordinate sequence to align the time dimension of the point cloud data; Compensate for the scanning delay error of the lidar and generate a time-synchronized point cloud dataset.

9. The laser radar-based autonomous obstacle avoidance optimization method for unmanned aerial vehicles according to claim 8, characterized in that: The method also includes optimizing the training process of the reinforcement learning algorithm, specifically: Initialize the weight parameters of the policy network; Use historical flight data sets to simulate environmental interactions and calculate policy gradient updates; The policy gradient update value is used to adjust the weight of the reward function mechanism.

10. The laser radar-based autonomous obstacle avoidance optimization method for UAVs according to claim 9, characterized in that: The method also includes an adaptive mechanism for handling dynamic obstacles, specifically: The changing trend of the motion speed estimate is monitored. If the changing trend exceeds a stable interval, a reassessment of the path node is triggered, and a supplementary obstacle avoidance instruction is generated. The supplementary obstacle avoidance instruction is merged into the obstacle avoidance path optimization instruction.

Citation Information

Patent Citations

  • Unmanned aerial vehicle power inspection autonomous flight obstacle avoidance method and device based on laser radar, and storage medium

    CN119717864A

  • Automatic obstacle avoidance point selection and obstacle avoidance method for photovoltaic station polled by unmanned aerial vehicle

    CN119937623A

  • Unmanned aerial vehicle flight control system and method with precise positioning and autonomous obstacle avoidance

    CN120255563A

Cited By

  • Intelligent control system and method based on intelligent indication board

    CN121053809A

  • Unmanned aerial vehicle atmosphere data anomaly detection and correction system based on deep learning

    CN121456777A

  • Omnibearing intelligent anti-collision method and system for trackless equipment

    CN121541677A

  • Safety control method and system of unmanned equipment and storage medium

    CN121541698A

  • Distributed Nash equilibrium collision avoidance control method for low-altitude logistics aircraft cluster operation

    CN121704516A