Unmanned aerial vehicle adaptive flight control system based on multi-modal data fusion

By combining RTK positioning and LiDAR with multimodal data fusion technology using binocular cameras, adaptive flight control of UAVs is achieved, solving the problems of low obstacle avoidance recognition accuracy and response lag, improving obstacle avoidance accuracy and efficiency, and reducing collision risk.

CN121918595APending Publication Date: 2026-04-24GUANGZHOU ICLOUDSTAR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing drone obstacle avoidance technology is susceptible to external environmental influences, has low obstacle recognition accuracy, slow response, low flight efficiency, and lacks the ability to predict the movement trend of obstacles, resulting in high collision risk and insufficient safety margin.

Method used

The system employs an RTK positioning unit that works synchronously with a lidar unit, and combines lidar and binocular cameras for multimodal data fusion. Through target localization, tracking and monitoring, motion direction prediction, and flight control modules, it achieves adaptive flight control of the UAV.

Benefits of technology

It improves the obstacle avoidance accuracy and response efficiency of drones in dynamic environments, ensuring a balance between safety and flight efficiency, and reducing the risk of collisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121918595A_ABST
    Figure CN121918595A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle adaptive flight control system based on multi-modal data fusion, and relates to the field of unmanned aerial vehicle flight control, and the method comprises the steps: obtaining the real-time position of a target unmanned aerial vehicle in real time through an RTK positioning unit, carrying out the scanning monitoring of the surrounding environment of the target unmanned aerial vehicle through a laser radar unit, and determining a target obstacle; performing tracking detection on the target obstacle to obtain a first detection sequence and a second detection sequence, and performing fusion processing to obtain a fusion detection sequence; performing motion direction prediction according to the fusion detection sequence, determining a plurality of predicted motion directions, and obtaining the motion confidence of each predicted motion direction; and on the basis of the real-time position of the unmanned aerial vehicle, generating an obstacle avoidance path based on the plurality of predicted motion directions and the motion confidence of each predicted motion direction, and performing adaptive flight control. The problems that in the prior art, unmanned aerial vehicle obstacle avoidance is prone to being affected by the external environment, obstacle avoidance recognition precision is low, obstacle avoidance response lags behind, and flight efficiency is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) flight path planning, and more specifically to an adaptive flight control system for UAVs based on multimodal data fusion. Background Technology

[0002] With the rapid development and widespread application of drone technology, drones are playing an increasingly important role in fields such as power line inspection, logistics delivery, emergency rescue, and agricultural plant protection. However, existing drone obstacle avoidance technologies mainly rely on single sensors or simple combinations of multiple sensors. Although they can provide distance information and identify obstacle types and texture features, their depth perception accuracy is relatively low, and they are easily affected by changes in lighting conditions, rain, fog, and other weather conditions, resulting in a sharp decline in performance.

[0003] Meanwhile, most existing obstacle avoidance technologies employ passive reaction strategies, lacking the ability to predict obstacle movement trends. This results in delayed obstacle avoidance predictions for dynamic obstacles, leading to extremely high collision risks. Furthermore, the lack of defined safe distances in the spatial dimension for obstacle avoidance prevents differentiated protection based on threat levels, reducing flight efficiency and resulting in insufficient safety margins in high-risk directions. Summary of the Invention

[0004] This application provides an adaptive flight control system for unmanned aerial vehicles (UAVs) based on multimodal data fusion, which addresses the problems in existing technologies such as UAV obstacle avoidance being easily affected by the external environment, low obstacle avoidance recognition accuracy, delayed obstacle avoidance response, and low flight efficiency.

[0005] In view of the above problems, this application provides an adaptive flight control system for unmanned aerial vehicles based on multimodal data fusion, comprising: The target positioning module is used to obtain the real-time position of the target UAV through the RTK positioning unit, and at the same time scan and monitor the surrounding environment of the target UAV through the lidar unit to identify target obstacles. The tracking and monitoring module is used to track and detect the target obstacle based on the lidar unit and the binocular camera unit, obtain a first detection sequence and a second detection sequence, and perform fusion processing on the first detection sequence and the second detection sequence to obtain a fused detection sequence. The motion direction prediction module is used to predict the motion direction of the target obstacle based on the fused detection sequence, determine multiple predicted motion directions, and obtain the motion confidence of each predicted motion direction. The flight control module is used to generate an obstacle avoidance path for the target UAV based on the real-time position of the UAV, the multiple predicted motion directions and the motion confidence of each predicted motion direction, and to perform adaptive flight control on the target UAV.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application first achieves centimeter-level high-precision positioning of the UAV itself and reliable initial screening of dynamic obstacles in the surrounding environment by synchronously working the RTK positioning unit and the lidar unit, providing a stable and accurate self-state reference and initial threat targets for the entire system. Second, based on the lidar unit and the binocular camera unit, target obstacles are tracked and fused, overcoming the performance limitations of a single sensor in specific environments. Light intensity perception is introduced to dynamically adjust the fusion weights, enhancing the environmental stability of target state estimation. Third, based on the high-quality fused detection sequence, an obstacle motion prediction engine is used to predict the motion direction and output the confidence level, providing a refined decision-making basis for risk assessment. Finally, based on the UAV's real-time position, predicted direction, and confidence level, differentiated obstacle avoidance paths are generated and adaptive control is executed, achieving a two-way balance between safety and flight efficiency. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a schematic diagram of the structure of an adaptive flight control system for unmanned aerial vehicles based on multimodal data fusion, as proposed in this application.

[0009] Figure 2 This is a schematic diagram of the process of activating the obstacle motion prediction engine and obtaining the total number of predictors in an UAV adaptive flight control system based on multimodal data fusion according to this application.

[0010] In the attached diagram, the components represented by each number are as follows: Target positioning module 11, tracking and monitoring module 12, motion direction prediction module 13, flight control module 14. Detailed Implementation

[0011] This application provides an adaptive flight control system for unmanned aerial vehicles (UAVs) based on multimodal data fusion, which solves the problems in the prior art where UAV obstacle avoidance is easily affected by the external environment, with low obstacle avoidance recognition accuracy, delayed obstacle avoidance response, and low flight efficiency.

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0013] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.

[0014] The present invention will now be described in detail with reference to the accompanying drawings.

[0015] In the embodiments, such as Figure 1 As shown, this application provides an adaptive flight control system for unmanned aerial vehicles (UAVs) based on multimodal data fusion, the system comprising: The target positioning module 11 is used to obtain the real-time position of the target UAV through the RTK positioning unit, and at the same time scan and monitor the surrounding environment of the target UAV through the lidar unit to identify target obstacles. The tracking and monitoring module 12 is used to track and detect the target obstacle based on the lidar unit and the binocular camera unit, obtain a first detection sequence and a second detection sequence, and perform fusion processing on the first detection sequence and the second detection sequence to obtain a fused detection sequence. The motion direction prediction module 13 is used to predict the motion direction of the target obstacle based on the fused detection sequence, determine multiple predicted motion directions, and obtain the motion confidence of each predicted motion direction. The flight control module 14 is used to generate an obstacle avoidance path for the target UAV based on the real-time position of the UAV, the multiple predicted motion directions and the motion confidence of each predicted motion direction, and to perform adaptive flight control on the target UAV.

[0016] The target positioning module 11 is used to obtain the real-time position of the target UAV through the RTK positioning unit, and at the same time scan and monitor the surrounding environment of the target UAV through the lidar unit to identify target obstacles. In this embodiment, the RTK positioning unit is a carrier phase differential positioning device, a high-precision satellite positioning technology. The lidar unit is an active remote sensing sensor that emits a laser beam and receives the signal reflected back from the surface of an object, measures the time of flight of light, and thus calculates the precise distance from the sensor to the object surface. The real-time position of the UAV is the precise spatial coordinates of the target UAV in the global or local navigation coordinate system at the current moment. The scanning monitoring involves the lidar continuously and repeatedly sampling the space within its field of view at a certain angular velocity and scanning mode to acquire three-dimensional point cloud data of the environment.

[0017] Specifically, the RTK positioning unit continuously outputs the UAV's real-time position, while the LiDAR unit performs high-speed, high-resolution scanning of the space surrounding the UAV, generating a dense 3D point cloud. Then, by processing continuous frames of point cloud data, moving objects are distinguished from the static background. When the speed, trajectory, or distance of a moving object is determined to pose a potential collision risk to the UAV, it is identified as a target obstacle requiring focused attention and intervention.

[0018] Specifically, the real-time position of the target drone is acquired through the RTK positioning unit, and the surrounding environment of the target drone is scanned and monitored by the lidar unit to identify target obstacles, including: The scanning system uses a lidar unit to scan the surrounding environment within a preset range around the target UAV to obtain environmental point cloud time sequence, which includes environmental point cloud data from multiple consecutive frames. The environmental point cloud data of the continuous multi-frames is analyzed by a time-series analysis system to determine the point cloud change areas. The target determination system calculates the movement speed of the point cloud change region and identifies the point cloud change region with a movement speed exceeding a preset movement threshold as a target obstacle.

[0019] In this embodiment, the surrounding environment within a preset range around the target UAV is first scanned by a scanning system using a lidar unit to obtain an environmental point cloud time sequence, which includes multiple consecutive frames of environmental point cloud data. Specifically, the lidar unit performs a 360-degree scan of the surrounding environment within the preset range around the target UAV to obtain an environmental point cloud time sequence, which includes multiple consecutive frames of environmental point cloud data. The preset range is the boundary of the spatial area that the lidar needs to focus on monitoring; scanning is the process by which the lidar sensor emits a laser beam and receives the echo according to its inherent physical mechanism, systematically traversing the space within its field of view; each complete traversal is called a frame; the environmental point cloud time sequence is a collection of multiple consecutive frames of environmental point cloud data arranged in chronological order; and the environmental point cloud data is the collection of all valid measurement points acquired by the lidar within a single scanning cycle.

[0020] Based on the UAV's flight speed, braking capability, and mission requirements, a reasonable preset monitoring range is set. Subsequently, after the lidar unit is activated, it begins periodic scanning of the defined spatial range at a fixed frequency. After each scan, a frame of environmental point cloud data is generated. Spatial scanning continues, and the continuously generated multiple frames of point cloud data are saved and arranged in timestamp order to form a coherent environmental point cloud time sequence.

[0021] Secondly, a time-series analysis system is used to perform time-series analysis on multiple consecutive frames of environmental point cloud data to determine the point cloud change regions. The time-series analysis involves comparing and calculating multiple data frames of environmental point cloud data to analyze patterns, trends, or anomalies in data changes over time. The point cloud change region is a three-dimensional spatial sub-region in which the spatial distribution of point sets changes significantly within multiple consecutive frames of point cloud data.

[0022] Specifically, to perform temporal analysis on environmental point cloud data of multiple consecutive frames, the inter-frame difference method can be used to spatially register the point cloud data of two adjacent frames, and then calculate the point density change within each tiny voxel. By calculating, the set of local point clouds that have changed can be selected from the global point cloud containing a large number of static background points, and potential point cloud change areas can be identified in each set as the location of subsequent moving candidate targets.

[0023] Finally, the target determination system calculates the velocity of the point cloud change region and identifies those regions with velocities exceeding a preset velocity threshold as target obstacles. Velocity is the movement rate of the object represented by the point cloud change region in three-dimensional space; it is a vector containing magnitude and direction. The preset velocity threshold is a pre-defined threshold value for velocity magnitude, used to distinguish between threatening motion and harmless minor movements or noise. Target obstacles are dynamic objects that, after velocity screening, are determined to pose a substantial threat to the flight safety of the UAV.

[0024] Specifically, after identifying the point cloud change region, the movement velocity of that region is first calculated: a short-term motion trajectory is fitted using the three-dimensional centroid coordinates from the previous few frames, and its average speed is calculated. Then, the movement velocity is compared with a preset motion threshold. This preset motion threshold is set based on the UAV's own speed, maneuverability, and mission safety standards. Point cloud change regions whose movement velocity exceeds the preset motion threshold are identified as target obstacles for the entire tracking and obstacle avoidance process.

[0025] For example, for a point cloud change region, the centroid displacement over a period of 0.3 seconds from time t1 to t4 is calculated to be 2.4 meters, with a velocity of 8 meters per second. The preset motion threshold is 1.5 meters per second. Since 8 meters per second > 1.5 meters per second, the motion velocity of the point cloud change region exceeds the preset motion threshold. The object represented by the moving point cloud cluster is then identified as the target obstacle.

[0026] In this embodiment, by analyzing the temporal sequence of multiple consecutive frames of point cloud data, the static background and moving foreground in the environment are effectively distinguished, static interference is filtered out, and the false alarm rate is significantly reduced. Subsequently, by calculating the motion speed of the point cloud change region and comparing it with a preset threshold, low-speed harmless motion and high-speed threatening motion are distinguished, ensuring the adaptation of subsequent tracking and computing resources. This provides a high-quality data foundation for subsequent tracking, prediction, and obstacle avoidance, improving the system's response efficiency and security in dynamic environments.

[0027] The tracking and monitoring module 12 is used to track and detect the target obstacle based on the lidar unit and the binocular camera unit, obtain a first detection sequence and a second detection sequence, and perform fusion processing on the first detection sequence and the second detection sequence to obtain a fused detection sequence. In this embodiment, the binocular camera unit is a vision system composed of two cameras placed parallel to each other at a certain baseline distance; tracking detection is the continuous locking and measurement of the state of the identified target obstacle in continuous time series data; the first detection sequence is the observation data generated when tracking and detecting the target obstacle; the second detection sequence is the observation data generated by the binocular camera unit when tracking and detecting the target obstacle; fusion processing is the process of aligning, associating, complementing and integrating the first detection sequence and the second detection sequence in time and space.

[0028] Specifically, the lidar unit uses its ranging capability to perform directional scanning of target obstacles, outputting a 3D point cloud of the target's position, forming a first detection sequence containing high-precision distance information. The binocular camera unit performs visual tracking of the target and estimates its depth using stereo vision principles, forming a second detection sequence containing visual features and depth information. These are then fused to generate a final fused detection sequence.

[0029] First, the target obstacle is tracked and detected based on the lidar unit and the binocular camera unit, resulting in a first detection sequence and a second detection sequence. The first and second detection sequences are then fused to obtain a fused detection sequence, including: The target obstacle is continuously scanned by the lidar unit to obtain multi-frame point cloud data of the target obstacle, and multiple point cloud position information of the target obstacle is determined based on the multi-frame point cloud data to form a first detection sequence, wherein each point cloud position information has a data timestamp. The target obstacle is continuously imaged using a binocular camera unit to obtain multi-frame image data of the target obstacle. Based on the multi-frame image data, multiple image position information of the target obstacle is determined to form a second detection sequence, wherein each image position information has a data timestamp.

[0030] In this embodiment, after identifying the target obstacle, the lidar unit first switches from a wide-area monitoring mode to a tracking mode that scans a specific target. It continuously emits a laser beam and receives the echo reflected from the surface of the target obstacle to acquire point cloud data of the target obstacle for each frame. Then, it separates the point cloud data from the background using a clustering algorithm and calculates the three-dimensional centroid coordinates of the point cloud data as the point cloud position information at the current moment, while recording a precise data timestamp. This process is repeated and arranged in chronological order to form a first detection sequence describing the target's spatial trajectory: [(t1,x1,y1,z1),(t2,x2,y2,z2)...].

[0031] Among them, continuous scanning is the process by which the lidar unit focuses its scanning resources on the spatial region where the target obstacle is located after it has been identified, and performs high-frequency, continuous ranging sampling to continuously acquire its latest three-dimensional contour data; multi-frame point cloud data of the target obstacle is the set of point clouds on the surface of the target obstacle acquired by the lidar in multiple consecutive scanning cycles during tracking; point cloud position information is the state quantity of the overall spatial position of the target extracted from the point cloud data of a frame of the target obstacle.

[0032] Secondly, a binocular camera unit is used to acquire continuous frame images of the identified target obstacle. First, a visual detection algorithm is run to identify and outline the target obstacle in the left-eye image, obtaining its image bounding box. Then, feature matching is performed in the right-eye image to find the corresponding region of the target obstacle. Next, based on the pixel coordinate difference between the center of the bounding box or specific feature points in the left and right images, a pre-calibrated binocular camera model is used to obtain the three-dimensional spatial coordinates of the target obstacle relative to the camera, i.e., the image position information. The image acquisition timestamps are also appended. Continuous processing yields the final visual three-dimensional position points with timestamps, which are then arranged in chronological order to form the second detection sequence.

[0033] Among them, continuous frame image acquisition is the simultaneous image stream formed by the binocular camera unit continuously capturing synchronous image pairs of the scene where the target obstacle is located at a fixed frame rate; multi-frame image data is the continuous left and right eye images containing the target obstacle by the binocular camera during the tracking process; image position information is the position of the target obstacle in three-dimensional space determined from the left and right eye image data through target detection and stereo vision matching.

[0034] For example, within the same time frame, the binocular camera unit and the lidar are time-synchronized to acquire continuous frame images of the target obstacle at a certain frequency. In the image data acquired at time t1, the bounding box of the target obstacle is detected in the left image. The disparity with the right image is calculated, and the image position information is determined as (x1, y1, z1). At time t2, the position (x2, y2, z2) is calculated, ultimately forming the second detection sequence [(t1, x1, y1, z1), (t2, (x2, y2, z2), (t3, (x3, y3, z3)...].

[0035] Furthermore, the target obstacle is tracked and detected based on the lidar unit and the binocular camera unit to obtain a first detection sequence and a second detection sequence. The first detection sequence and the second detection sequence are then fused to obtain a fused detection sequence, including: The current light intensity of the target UAV's current flight environment is collected, and the preset standard light intensity of the binocular camera unit is obtained. The preset standard light intensity has a corresponding benchmark weighting coefficient. A weight adjustment factor is determined based on the current light intensity and the preset standard light intensity, and the benchmark weight coefficient is adjusted based on the weight adjustment factor to obtain the second dynamic weight. The first dynamic weight is determined based on the second dynamic weight; By combining the data timestamps, the first detection sequence and the second detection sequence are fused based on the first dynamic weight and the second dynamic weight to obtain a fused detection sequence.

[0036] In this embodiment, the current illumination intensity of the flight environment is first acquired using airborne sensors. Simultaneously, a pre-stored preset standard illumination intensity is retrieved from memory. This preset standard intensity is associated with a reference weight coefficient; when the ambient illumination is K, the position information provided by the binocular camera has a basic reliability of k. Furthermore, the current illumination intensity is the actual illumination level of the surrounding environment at the moment the target UAV performs its flight mission; the preset standard illumination intensity is an ideal or typical illumination condition reference value pre-set for the binocular camera unit during the system design or calibration phase; and the reference weight coefficient is an initial weight value associated with the preset standard illumination intensity, representing the initial reliability of the second detection sequence derived from the binocular camera data during the fusion process.

[0037] Secondly, a transformation function is used to calculate the weight adjustment factor, which is the m-th power of the ratio of the current illumination intensity to the preset standard illumination intensity, where m is 0.5. If the weight adjustment factor is less than 1, it indicates that the current illumination intensity is weaker than the preset standard illumination intensity, and the reliability of the visual data is expected to decrease; conversely, it indicates that the reliability of the visual data is relatively high. Subsequently, the baseline weight coefficient is adjusted according to the weight adjustment factor, where the second dynamic weight = baseline weight coefficient × weight adjustment factor. Furthermore, the weight adjustment factor is a scaling factor or scaling factor calculated based on the ratio or difference between the current illumination intensity and the preset standard illumination intensity; the second dynamic weight is the adjusted actual weight coefficient used in the fusion of the second detection sequence at the current moment.

[0038] For example, assuming the current light intensity is 25,000 Lux, the preset standard light intensity is 50,000 Lux, and the weight adjustment factor is (25,000 / 50,000). 0.5 ≈0.707. If the baseline weight coefficient is 0.5, then the second dynamic weight = 0.5 × 0.707 ≈ 0.35.

[0039] Next, the first dynamic weight is determined based on the second dynamic weight. The first dynamic weight is a real-time weighting coefficient applied to the first detection sequence during fusion, relative to the second dynamic weight. Specifically, after determining the second dynamic weight, the first dynamic weight of the LiDAR data is determined according to a preset weighting relationship rule. Since the first dynamic weight + second dynamic weight = 1, the first preset weight is 1 - the second dynamic weight.

[0040] For example, the first dynamic weight is 1 - 0.35 = 0.65.

[0041] Finally, combining the data timestamps, the first and second detection sequences are fused based on the first and second dynamic weights to obtain the fused detection sequence. The fused detection sequence is a new sequence of location information output after the fusion process.

[0042] Specifically, firstly, the asynchronous LiDAR sequence and stereo camera sequence are time-aligned by combining data timestamps. For the LiDAR position at a certain moment, the two points that are closest in time in the stereo camera sequence are found, and the visual position at that moment is estimated by linear interpolation. Then, based on the first and second dynamic weights, the two positions are fused. Similarly, this process is performed for each time point in the sequence, ultimately resulting in a fused detection sequence with timestamps.

[0043] In this embodiment, detection sequences are generated independently by LiDAR and binocular cameras. Then, by sensing the light intensity in real time, the fusion weights of the LiDAR and binocular camera data are dynamically adjusted, overcoming the problem of single vision systems failing in low light. Combining the two ensures higher-quality fused detection sequences under any lighting conditions, providing solid data support for subsequent motion direction prediction.

[0044] The motion direction prediction module 13 is used to predict the motion direction of the target obstacle based on the fused detection sequence, determine multiple predicted motion directions, and obtain the motion confidence of each predicted motion direction. In this embodiment, motion direction prediction is based on the motion information of the target obstacle over a period of time, using mathematical models or machine learning algorithms to infer its most likely motion trend in the near future; motion direction prediction is to derive the possible future motion orientation of the target obstacle through motion direction prediction; motion confidence is a reliability assessment value for each predicted motion direction, and the higher the value, the more reliable the predicted direction.

[0045] Specifically, after obtaining a high-quality fusion detection sequence, the historical motion patterns of the target obstacles revealed by the fusion detection sequence are analyzed, and multiple predicted motion directions and motion confidence scores are output through algorithms. The motion confidence scores are calculated based on factors such as the uncertainty of the prediction model, the smoothness of the historical trajectory, and prior knowledge of the target type.

[0046] like Figure 2 As shown, specifically, based on the fused detection sequence, the motion direction of the target obstacle is predicted, multiple predicted motion directions are determined, and the motion confidence of each predicted motion direction is obtained, including: Activate the obstacle motion prediction engine, which includes N obstacle motion predictors, and obtain the total number of predictors; The fused detection sequence is input into the obstacle motion prediction engine, and the fused detection sequence is processed by N obstacle motion predictors to predict the motion direction of the target obstacle, resulting in N prediction results. The motion directions of the N prediction results are classified, and the same predicted motion directions are merged into one category to form multiple predicted motion directions; The occurrence frequency of multiple predicted motion directions is counted separately to obtain the prediction frequency of multiple directions. Combined with the total number of predictors, the motion confidence of each predicted motion direction is determined.

[0047] In this embodiment, the obstacle motion prediction engine is first activated. The obstacle motion prediction engine includes N obstacle motion predictors, and the total number of predictors is obtained. The obstacle motion prediction engine is an integrated software module or algorithm structure that predicts the future movement trend of obstacles based on historical trajectory data. Activation involves initializing and starting the prediction engine when a prediction task needs to be performed, putting it into a working state.

[0048] Specifically, the pre-built obstacle motion prediction engine loaded into memory is first activated. Internally, it encapsulates N obstacle motion predictors; if N=5, it contains 5 different prediction algorithms. Simultaneously, the total number of obstacle motion predictors, N, is obtained from the engine's configuration and recorded for subsequent quantization calculations.

[0049] Secondly, the fused detection sequence is input into the obstacle motion prediction engine. N obstacle motion predictors process the fused detection sequence to predict the motion direction of the target obstacle, resulting in N prediction results. These N prediction results are N direction prediction values ​​generated by each of the N predictors working independently. Specifically, after initialization, the fused detection sequence representing the historical trajectory of the target obstacle is input into the activated obstacle motion prediction engine. The engine's N obstacle motion predictors process the fused detection sequence individually. Each predictor predicts the motion direction of the target obstacle, ultimately resulting in a set of N prediction results: {direction A, direction B…}.

[0050] Next, the motion directions of the N predicted results are categorized, and those with the same predicted motion direction are merged into one category, forming multiple predicted motion directions. Here, motion direction categorization is the process of clustering or grouping the N predicted results, which may be continuous or discrete values, according to their similarity; multiple predicted motion directions refer to the set of distinct and representative hypotheses of future motion directions obtained after categorization and merging.

[0051] Specifically, after collecting N raw prediction results, the motion directions of the prediction results are categorized. Each motion direction is compared and classified according to a set classification threshold. Motion directions with similar angular differences are grouped into the same category. Finally, the directions of the N raw prediction results are merged to form several predicted motion directions. For ease of prediction, the 360° scanning space can be divided into eight directions: north, northeast, east, southeast, south, southwest, west, and northwest. Each predictor selects the closest direction from these preset directions as its prediction result.

[0052] For example, if N=5, after classifying the 5 original prediction results, two predicted motion directions are obtained: {direction A, direction B}.

[0053] Finally, the occurrence frequency of multiple predicted motion directions is counted to obtain the prediction frequency for each direction. Combined with the total number of predictors, the motion confidence for each predicted motion direction is determined. The prediction frequency for a direction is obtained by statistically analyzing the number of predictors supporting a given predicted motion direction; the motion confidence is an estimate of the probability of each predicted motion direction occurring in the future.

[0054] Specifically, the N predicted motion directions are categorized, merging identical directions to form several unique predicted motion directions. Then, the frequency of each predicted motion direction in the N predicted directions is counted. The motion confidence score for each predicted motion direction is obtained by calculating the ratio of its frequency to the total number of predictors N. A higher motion confidence score indicates stronger prediction reliability for that motion direction, providing a quantitative basis for subsequent obstacle avoidance strategy development. Motion confidence score = (Number of direction predictions / Total number of predictors (N)). The motion confidence score ranges from 0 to 1, representing the proportion of all predictors that support a particular motion direction.

[0055] For example, if direction A appears 4 times and direction B appears 1 time, then the motion confidence scores for direction A and direction B are 4 / 5 = 0.8 and 1 / 5 = 0.2, respectively.

[0056] Specifically, the construction steps of the obstacle motion prediction engine include: Collect historical obstacle movement records, which include multiple historical obstacle movement data, each of which includes historical lidar detection sequences and historical binocular camera detection sequences; Historical lidar detection sequences and historical binocular camera detection sequences from multiple historical obstacle motion data are fused to construct a sample fusion detection sequence set. Based on the historical obstacle motion records, the motion direction of each fusion detection sequence in the sample fusion detection sequence set is labeled to obtain the sample motion direction set; Construct an architecture with N predictors; Using the sample fusion detection sequence set as input and the sample motion direction set as labels, train the N predictor architectures until convergence to obtain N obstacle motion predictors; The N obstacle motion predictors are integrated to obtain the obstacle motion prediction engine.

[0057] In this embodiment, all relevant historical obstacle motion records are first collected from past flight mission logs. Each record is a historical obstacle motion data, which contains two dimensions of raw tracking data: historical lidar detection sequences generated by lidar and historical binocular camera detection sequences generated synchronously by binocular cameras. Hundreds or thousands of such multimodal paired data are collected to form the raw dataset required for machine learning.

[0058] Among them, historical obstacle motion records are a collection of data archives recording the motion processes of various dynamic obstacles in the actual flight missions of UAVs in the past; historical obstacle motion data are independent motion instance data units that describe the motion of a specific obstacle over a period of time; historical lidar detection sequences are timestamped point cloud position information sequences generated by lidar units tracking the historical obstacle; and historical binocular camera detection sequences are timestamped image position information sequences generated by binocular camera units tracking the same historical obstacle.

[0059] Secondly, historical LiDAR detection sequences and historical binocular camera detection sequences from multiple historical obstacle motion data are fused to construct a sample fusion detection sequence set. Specifically, the data acquisition method for the sample fusion detection sequence set is the same as that for the current fusion detection sequence; the historical LiDAR detection sequences and historical binocular camera detection sequences are fused to obtain the sample fusion detection sequence set.

[0060] Next, based on the historical obstacle motion records, the motion direction of each fusion detection sequence in the sample fusion detection sequence set is labeled to obtain the sample motion direction set.

[0061] Specifically, for each fused detection sequence in the sample fusion detection sequence set, it is matched and compared with historical obstacle movement records for the corresponding time period. By analyzing the actual trajectory changes of the obstacle within that time period, its true direction of movement is determined. To facilitate rapid processing, a preset encoding method is used to discretize and label the movement directions, dividing the 360° scanning space into eight standard directional intervals, corresponding to due north, northeast, due east, southeast, due south, southwest, due west, and northwest, respectively, and assigning a unique numerical code to each direction. Subsequently, by labeling each fused detection sequence in the sample fusion detection sequence set, the confidence level of the resulting sample movement direction set is obtained. By querying historical data, the movement confidence level of the movement direction in the sample movement direction is obtained, and it is mapped to the movement direction of the sample to obtain a sample movement direction set that corresponds one-to-one with each fused detection sequence. This ensures that the prediction model can accurately learn the mapping relationship between obstacle movement patterns and sensor detection data.

[0062] Furthermore, N predictor architectures are constructed. These N predictor architectures represent N different principles or structural models. Specifically, based on the characteristics of the prediction task, N different predictor architectures are constructed. Each architecture is a blank model framework awaiting data filling. When constructing the N predictor architectures, differentiation between predictors is achieved by adjusting key parameters of the algorithm. For the adopted neural network algorithm, different parameter configurations are set for the number of network layers, neurons, learning rate, activation function, etc., to construct N neural network predictors with different feature learning capabilities.

[0063] Furthermore, using the sample fusion detection sequence set as input and the sample motion direction set as labels, N predictor architectures are trained until convergence, resulting in N obstacle motion predictors. Here, the label is the correct answer corresponding to each input, i.e., the direction value in the sample motion direction set; convergence is the state during training where the model's performance on the validation dataset no longer significantly improves, or the loss function value decreases to a stable level.

[0064] Specifically, the training algorithm is run using the sample fusion detection sequence set and the sample motion direction set as input. The training process is iterative, continuously adjusting the weight parameters within the predictor architecture until the model performs stably on the reserved validation set, reaching convergence. This independent training process is repeated separately. Ultimately, N obstacle motion predictors with varying performance are obtained, all of which have completed learning.

[0065] Finally, the N obstacle motion predictors are integrated to obtain an obstacle motion prediction engine. Integration is the process of organizing multiple independently trained base models and combining them into a more powerful unified model or decision system through specific strategies. Specifically, the N obstacle motion predictors are integrated. Internally, the engine calls all N predictors to perform predictions, then collects the prediction results, and packages the code of the N predictors, model parameters, and this integrated decision logic to obtain the complete obstacle motion prediction engine.

[0066] A backpropagation (BP) neural network can be used to build an obstacle motion prediction engine. The BP neural network model is a type of feedforward neural network that is trained through backpropagation of errors and is commonly used to predict continuous values.

[0067] For example, taking a BP neural network as an example, the steps to build an obstacle motion prediction engine are as follows: First, data preparation involves using the sample fusion detection sequence set as input features and the sample motion direction set as labels. The data is then divided into training, validation, and test sets in a 7:2:1 ratio.

[0068] Secondly, the model is constructed, mainly consisting of an input layer, hidden layers, and an output layer. The input layer performs a weighted summation on the sample fusion detection sequence set; the hidden layer performs a nonlinear transformation on the sample fusion detection sequence set through an activation function; and the output layer outputs the prediction results. Based on the characteristics of the N predictor architectures, the number of layers in the neural network is configured appropriately, and the number of neurons in each layer is also configured accordingly. The first architecture has 1 output layer, the second architecture has 2 output layers, and so on, up to the Nth architecture with N output layers. The number of neurons in each layer also increases accordingly. Subsequently, the corresponding learning rate and weights are adopted according to the model structure to ensure that N differentiated predictors are trained, producing differentiated prediction results.

[0069] Next, model training uses the sample motion direction set as the supervision label, and trains N predictor architectures simultaneously. For each of the N predictor architectures, different initial learning rates and weights are set, and weights are assigned. The mean squared error (MSE) function is used to calculate the error between the predicted and actual results. Weight adjustments are made and the calculation is repeated iteratively until the error is minimized. Parameters are generated through forward propagation and updated through backpropagation. Performance is evaluated using a validation set after each training epoch to avoid overfitting. The model is considered converged when the MSE loss on the training set decreases by less than 1e-5 for five consecutive epochs and the MSE loss on the validation set stabilizes below 0.01, resulting in N obstacle motion predictors.

[0070] Finally, the N obstacle motion predictors are integrated to obtain an obstacle motion prediction engine that contains N obstacle motion predictors.

[0071] In this embodiment, multiple obstacle motion predictors are used. By classifying and outputting multiple possible prediction directions and motion confidence levels, the fault tolerance and stability of the prediction are enhanced. Subsequently, by collecting historical multimodal data, constructing a fusion sample set and performing direction labeling, a high-quality training foundation is provided for supervised learning, realizing intelligent prediction and providing a decision basis for downstream path planning, which is conducive to realizing subsequent differentiated and adaptive obstacle avoidance.

[0072] The flight control module 14 is used to generate an obstacle avoidance path for the target UAV based on the real-time position of the UAV, the multiple predicted motion directions and the motion confidence of each predicted motion direction, and to perform adaptive flight control on the target UAV.

[0073] In this embodiment, the obstacle avoidance path is a temporary flight trajectory planned for the UAV, which enables the UAV to safely avoid predicted obstacle threats and return to the original mission route or head to the next target point as smoothly as possible after obstacle avoidance; adaptive flight control is the flight control system that dynamically adjusts the attitude, throttle and actions of each actuator of the UAV according to the obstacle avoidance path or control command generated in real time, so that the UAV can accurately and stably track the flight process of the new path.

[0074] Specifically, the latest acquired real-time location of the UAV is used as the starting point for path planning. Threat modeling and path search are performed based on multiple predicted motion directions and their confidence levels. Using a differentiated risk map, the path planning algorithm calculates an obstacle avoidance path that safely bypasses all threat areas and satisfies the UAV's dynamic constraints. Finally, the obstacle avoidance path is converted into waypoints or velocity commands and sent to the flight control system to execute adaptive flight control.

[0075] Specifically, based on the fused detection sequence, the motion direction of the target obstacle is predicted, multiple predicted motion directions are determined, and the motion confidence of each predicted motion direction is obtained, including: Obtain a preset set of motion directions, which includes multiple standard motion directions; A preset basic safety distance is set, and obstacle avoidance safety distances for multiple standard motion directions are determined by combining multiple predicted motion directions and the motion confidence of each predicted motion direction. Centered on the current position of the target obstacle, an obstacle avoidance threat zone is constructed based on the obstacle avoidance safety distances of multiple standard movement directions; Based on the obstacle avoidance threat area and the real-time position of the UAV, an obstacle avoidance path for the UAV to bypass the obstacle avoidance threat area is generated, and the target UAV is controlled to perform adaptive obstacle avoidance flight according to the UAV obstacle avoidance path.

[0076] In this embodiment of the application, a coordinate framework for analyzing spatial threats is first established, and a predefined set of motion directions is obtained from the internal configuration. This set of motion directions includes several standard motion directions. The predefined set of motion directions does not depend on the current specific obstacle prediction results and is a fixed pre-agreed agreement used to systematically evaluate the threat of obstacles in various possible directions.

[0077] Specifically, in the obstacle avoidance path generation stage of practical applications, eight standard motion directions are also used as the preset motion direction set to ensure that the trained N obstacle motion predictors can output prediction results that completely correspond to the training labels, avoiding prediction errors or system failures caused by inconsistent direction definitions. When the prediction engine outputs multiple predicted motion directions, the eight standard motion directions are used as a subset of the preset motion direction set, while for other directions not predicted in the preset motion direction set, a preset basic safety distance is used as a safety guarantee.

[0078] Secondly, a preset basic safety distance is set, and obstacle avoidance safety distances for multiple standard movement directions are determined by combining multiple predicted movement directions and the movement confidence of each predicted movement direction. Among them, the preset basic safety distance is the minimum distance that the drone should maintain between itself and any static or dynamic obstacle in any direction without considering the movement prediction of any specific obstacle.

[0079] Specifically, a preset base safety distance is first set. Then, multiple predicted motion directions of the target obstacle are combined, and the additional obstacle avoidance safety distance required for each standard motion direction is determined based on the proximity of each standard motion direction to the predicted motion direction and the corresponding motion confidence level. For standard directions that coincide with or are close to high-confidence predicted directions, the obstacle avoidance safety distance is significantly increased; for standard directions that are close to low-confidence directions, the obstacle avoidance safety distance is moderately increased; for standard directions that are unrelated to any predicted direction, the obstacle avoidance safety distance may remain at the base safety distance, thus determining different obstacle avoidance safety distances for each.

[0080] Next, using the current position of the target obstacle as the center, an obstacle avoidance threat zone is constructed based on obstacle avoidance safety distances in multiple standard motion directions. The current position of the target obstacle is its precise three-dimensional spatial coordinates at the latest moment, obtained through multimodal fusion tracking. The boundary of the obstacle avoidance threat zone defines a no-fly zone for the drone, and is a polyhedron or curved surface generated based on differentiated safety distances in each direction, representing the predicted threat distribution.

[0081] Specifically, the current position of the target obstacle is first obtained, and a three-dimensional obstacle avoidance threat region is constructed using this as the center and the acquired obstacle avoidance safety distance value. Starting from the current position of the target obstacle at the center point, the corresponding safety distance is extended outward along each standard direction to obtain an endpoint. Then, all endpoints are connected by curved surfaces or planes, and the enclosed space formed is the obstacle avoidance threat region. In high-threat directions, the safety distance is large, and the obstacle avoidance threat region bulges outward; in low-threat directions, the safety distance is small, and the obstacle avoidance threat region is relatively flat.

[0082] Finally, based on the obstacle avoidance threat area and the real-time position of the UAV, an obstacle avoidance path is generated to bypass the obstacle avoidance threat area. The target UAV is then controlled to perform adaptive obstacle avoidance flight according to the obstacle avoidance path. All waypoints of the obstacle avoidance path must be located outside the obstacle avoidance threat area, spatially bypassing the new flight trajectory formed by the dynamic restricted area.

[0083] Specifically, based on the obstacle avoidance threat area and real-time position of the target obstacle, and under the premise of satisfying the UAV's dynamic constraints, a flight path is generated that starts from the current real-time position, safely detours around the obstacle avoidance threat area, and eventually returns to the original mission route. This path serves as the final UAV obstacle avoidance path. Subsequently, the generated path point sequence is sent to the flight control system, which then controls the UAV to begin real-time, adaptive obstacle avoidance flight according to this new path by adjusting the motor speed and control surface angle.

[0084] Specifically, a preset basic safety distance is set, and obstacle avoidance safety distances for multiple standard motion directions are determined by combining multiple predicted motion directions and the motion confidence of each predicted motion direction, including: Determine the first standard motion direction from the plurality of standard motion directions; Determine whether the first standard motion direction exists among the plurality of predicted motion directions; If the first standard motion direction exists among the plurality of predicted motion directions, then the motion confidence of the corresponding predicted motion direction is obtained as the first motion confidence, and the first distance adjustment coefficient is calculated based on the first motion confidence. If the first standard direction of motion does not exist among the plurality of predicted directions of motion, then the first distance adjustment coefficient is set to 1; Based on the preset basic safety distance and the first distance adjustment coefficient, the obstacle avoidance safety distance corresponding to the first standard movement direction is obtained; Obstacle avoidance safe distances for other standard motion directions are obtained by obtaining the obstacle avoidance safe distances for the first standard motion direction in the same way as obtaining the obstacle avoidance safe distances for the other standard motion directions.

[0085] In this embodiment, a first standard motion direction is first determined from multiple standard motion directions. This first standard motion direction is selected as the first object for calculating the obstacle avoidance safety distance from among multiple standard directions included in a preset set of motion directions.

[0086] Specifically, a preset set of motion directions containing K directions is first obtained. From the set of K directions, by traversing all directions in the set, the first direction is selected in ascending order of azimuth angle and will be used as the first standard motion direction.

[0087] Secondly, it is determined whether the first standard motion direction exists among multiple predicted motion directions. Specifically, after identifying the first standard motion direction, it is determined whether it belongs to one of the multiple predicted motion directions of the prediction engine.

[0088] For example, assuming there is a matching tolerance, if the angle difference between the first standard motion direction and all predicted motion directions is much greater than the matching tolerance, then it is determined that the first standard motion direction does not match any predicted direction, and it is considered that the first standard motion direction does not exist in the predicted motion direction; if the angle difference between the first standard motion direction and one or more predicted motion directions in the predicted motion direction is less than the matching tolerance, then it is considered that the first standard motion direction exists in the predicted motion direction.

[0089] For example, assume a matching tolerance of 22.5 degrees for direction A (270 degrees, confidence level 0.8) and direction B (315 degrees, confidence level 0.2). The first standard motion direction is 0 degrees. The angular difference between the first standard motion direction and direction A is 90 degrees, and the angular difference with direction B is 45 degrees, both much greater than 22.5 degrees. Therefore, the first standard motion direction does not exist in the predicted motion directions.

[0090] Furthermore, if the first standard motion direction exists among multiple predicted motion directions, the motion confidence of the corresponding predicted motion direction is obtained as the first motion confidence, and a first distance adjustment coefficient is calculated based on the first motion confidence. If the first standard motion direction does not exist among multiple predicted motion directions, the first distance adjustment coefficient is set to 1. Here, the first motion confidence is the confidence value carried by the predicted direction that matches the currently processed first standard motion direction; the first distance adjustment coefficient is the calculated adjustment factor. A higher first motion confidence results in a larger adjustment coefficient, indicating that a greater safety extension is needed in that direction.

[0091] Specifically, if the judgment result is that the motion confidence of the predicted motion direction matching the first standard motion direction is obtained and used as the first motion confidence, and a first distance adjustment coefficient is calculated. The first distance adjustment coefficient = 1 + first motion confidence. If the judgment result is that the motion confidence of the predicted motion direction matches the first standard motion direction, the first distance adjustment coefficient is 1.

[0092] For example, assuming that the first standard motion direction exists in the predicted direction and matches direction A, and the motion confidence of direction A is 0.8, then the first distance adjustment coefficient = 1.8.

[0093] Furthermore, based on the preset basic safety distance and the first distance adjustment coefficient, the obstacle avoidance safety distance corresponding to the first standard movement direction is obtained. The obstacle avoidance safety distance in the first standard movement direction is the product of the preset basic safety distance and the first distance adjustment coefficient. Here, the obstacle avoidance safety distance is the final safety distance value calculated for the currently processed standard direction.

[0094] For example, if the preset safety distance is 5 meters and the first distance adjustment coefficient is 1.8, then the obstacle avoidance safety distance in the first standard movement direction is 1.8 × 5 = 9 meters.

[0095] Finally, following the method for obtaining the obstacle avoidance safety distance corresponding to the first standard motion direction, the obstacle avoidance safety distances corresponding to the other standard motion directions are obtained, resulting in obstacle avoidance safety distances for multiple standard motion directions. The other standard motion directions are all standard directions other than the first standard motion direction, within a preset set of motion directions.

[0096] Specifically, the remaining standard motion directions are processed iteratively or cyclically. First, it is determined whether they exist in the prediction set. Then, an adjustment coefficient is calculated based on the determination result. Finally, the base distance is multiplied by the coefficient to obtain the safe distance. Ultimately, after all remaining standard motion directions have been processed, a set of obstacle avoidance safe distances containing K values ​​is obtained.

[0097] In this embodiment, by presetting a set of motion directions, determining the safe distance for each direction, constructing a non-uniform threat area, generating a detour path, and controlling flight, the geometry of the threat area is dynamically shaped using predictive information. Subsequently, by determining whether the standard direction is predicted as a threat direction and converting the confidence level linearly or non-linearly into a distance adjustment coefficient, the obstacle avoidance safe distance is accurately calculated. A significant safety extension is obtained for high-confidence threat directions, a moderate extension is obtained for low-confidence directions, and the basic safe distance is maintained for directions without prediction. Obstacle avoidance is completed with a better path and flight distance deviation, reducing the risk of collision and improving inspection efficiency.

[0098] The embodiments of this application, through the above specific implementation methods, achieve the following technical effects: In this embodiment, the target localization module 11 first performs time-series analysis on multiple consecutive frames of point cloud data to effectively distinguish between static backgrounds and moving foregrounds in the environment, filter out static interference, and significantly reduce the false alarm rate. Subsequently, by calculating the motion speed of the point cloud change region and comparing it with a preset threshold, it distinguishes between low-speed harmless motion and high-speed threatening motion, ensuring the adaptation of subsequent tracking and computing resources, providing a high-quality data foundation for subsequent tracking, prediction, and obstacle avoidance, and improving the system's response efficiency and security in dynamic environments.

[0099] Secondly, the tracking and monitoring module 12 acquires detection sequences independently generated by the LiDAR and binocular camera. Then, by sensing the light intensity in real time, the fusion weights of the LiDAR and binocular camera data are dynamically adjusted, overcoming the problem of single vision systems failing in low light. The combination of these two methods ensures higher-quality fused detection sequences under any lighting conditions, providing a solid data foundation for subsequent motion direction prediction.

[0100] Furthermore, the motion direction prediction module 13 uses multiple obstacle motion predictors to classify and output multiple possible prediction directions and motion confidence, which enhances the fault tolerance and stability of the prediction. Subsequently, by collecting historical multimodal data, constructing a fusion sample set and performing direction labeling, a high-quality training foundation is provided for supervised learning, realizing intelligent prediction and providing a decision basis for downstream path planning, which is conducive to realizing subsequent differentiated and adaptive obstacle avoidance.

[0101] Finally, the flight control module 14 acquires a preset set of motion directions, determines the safe distance for each direction, constructs a non-uniform threat area, generates a detour path, and controls the flight. It dynamically shapes the geometry of the threat area using predictive information. Subsequently, by determining whether the standard direction is predicted as a threat direction and converting the confidence level linearly or non-linearly into a distance adjustment coefficient, it achieves accurate calculation of the obstacle avoidance safe distance. It obtains a significant safety extension for high-confidence threat directions, a moderate extension for low-confidence directions, and maintains a basic safe distance for directions without prediction. It completes obstacle avoidance with a better path and range deviation, reducing collision risk and improving inspection efficiency.

[0102] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0103] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0104] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.

Claims

1. An adaptive flight control system for unmanned aerial vehicles (UAVs) based on multimodal data fusion, characterized in that, The system includes: The target positioning module is used to obtain the real-time position of the target UAV through the RTK positioning unit, and at the same time scan and monitor the surrounding environment of the target UAV through the lidar unit to identify target obstacles. The tracking and monitoring module is used to track and detect the target obstacle based on the lidar unit and the binocular camera unit, obtain a first detection sequence and a second detection sequence, and perform fusion processing on the first detection sequence and the second detection sequence to obtain a fused detection sequence. The motion direction prediction module is used to predict the motion direction of the target obstacle based on the fused detection sequence, determine multiple predicted motion directions, and obtain the motion confidence of each predicted motion direction. The flight control module is used to generate an obstacle avoidance path for the target UAV based on the real-time position of the UAV, the multiple predicted motion directions and the motion confidence of each predicted motion direction, and to perform adaptive flight control on the target UAV.

2. The system according to claim 1, characterized in that, The execution steps of the target localization module include: The scanning system uses a lidar unit to scan the surrounding environment within a preset range around the target UAV to obtain environmental point cloud time sequence, which includes environmental point cloud data from multiple consecutive frames. The environmental point cloud data of the continuous multi-frames is analyzed by a time-series analysis system to determine the point cloud change areas. The target determination system calculates the movement speed of the point cloud change region and identifies the point cloud change region with a movement speed exceeding a preset movement threshold as a target obstacle.

3. The system according to claim 1, characterized in that, The execution steps of the tracking and monitoring module include: The target obstacle is continuously scanned by the lidar unit to obtain multi-frame point cloud data of the target obstacle, and multiple point cloud position information of the target obstacle is determined based on the multi-frame point cloud data to form a first detection sequence, wherein each point cloud position information has a data timestamp. The target obstacle is continuously imaged using a binocular camera unit to obtain multi-frame image data of the target obstacle. Based on the multi-frame image data, multiple image position information of the target obstacle is determined to form a second detection sequence, wherein each image position information has a data timestamp.

4. The system according to claim 3, characterized in that, The first detection sequence and the second detection sequence are fused to obtain a fused detection sequence, including: The current light intensity of the target UAV's current flight environment is collected, and the preset standard light intensity of the binocular camera unit is obtained. The preset standard light intensity has a corresponding benchmark weighting coefficient. A weight adjustment factor is determined based on the current light intensity and the preset standard light intensity, and the benchmark weight coefficient is adjusted based on the weight adjustment factor to obtain the second dynamic weight. The first dynamic weight is determined based on the second dynamic weight; By combining the data timestamps, the first detection sequence and the second detection sequence are fused based on the first dynamic weight and the second dynamic weight to obtain a fused detection sequence.

5. The system according to claim 1, characterized in that, The execution steps of the motion direction prediction module include: Activate the obstacle motion prediction engine, which includes N obstacle motion predictors, and obtain the total number of predictors; The fused detection sequence is input into the obstacle motion prediction engine, and the fused detection sequence is processed by N obstacle motion predictors to predict the motion direction of the target obstacle, resulting in N prediction results. The motion directions of the N prediction results are classified, and the same predicted motion directions are merged into one category to form multiple predicted motion directions; The occurrence frequency of multiple predicted motion directions is counted separately to obtain the prediction frequency of multiple directions. Combined with the total number of predictors, the motion confidence of each predicted motion direction is determined.

6. The system according to claim 1, characterized in that, The steps for building the obstacle motion prediction engine include: Collect historical obstacle movement records, which include multiple historical obstacle movement data, each of which includes historical lidar detection sequences and historical binocular camera detection sequences; Historical lidar detection sequences and historical binocular camera detection sequences from multiple historical obstacle motion data are fused to construct a sample fusion detection sequence set. Based on the historical obstacle motion records, the motion direction of each fusion detection sequence in the sample fusion detection sequence set is labeled to obtain the sample motion direction set; Construct an architecture with N predictors; Using the sample fusion detection sequence set as input and the sample motion direction set as labels, train the N predictor architectures until convergence to obtain N obstacle motion predictors; The N obstacle motion predictors are integrated to obtain the obstacle motion prediction engine.

7. The system according to claim 1, characterized in that, The execution steps of the flight control module include: Obtain a preset set of motion directions, which includes multiple standard motion directions; A preset basic safety distance is set, and obstacle avoidance safety distances for multiple standard motion directions are determined by combining multiple predicted motion directions and the motion confidence of each predicted motion direction. Centered on the current position of the target obstacle, an obstacle avoidance threat zone is constructed based on the obstacle avoidance safety distances of multiple standard movement directions; Based on the obstacle avoidance threat area and the real-time position of the UAV, an obstacle avoidance path for the UAV to bypass the obstacle avoidance threat area is generated, and the target UAV is controlled to perform adaptive obstacle avoidance flight according to the UAV obstacle avoidance path.

8. The system according to claim 7, characterized in that, A preset basic safety distance is set, and obstacle avoidance safety distances for multiple standard motion directions are determined by combining multiple predicted motion directions and the motion confidence of each predicted motion direction, including: Determine the first standard motion direction from the plurality of standard motion directions; Determine whether the first standard motion direction exists among the plurality of predicted motion directions; If the first standard motion direction exists among the plurality of predicted motion directions, then the motion confidence of the corresponding predicted motion direction is obtained as the first motion confidence, and the first distance adjustment coefficient is calculated based on the first motion confidence. If the first standard direction of motion does not exist among the plurality of predicted directions of motion, then the first distance adjustment coefficient is set to 1; Based on the preset basic safety distance and the first distance adjustment coefficient, the obstacle avoidance safety distance corresponding to the first standard movement direction is obtained; Obstacle avoidance safe distances for other standard motion directions are obtained by obtaining the obstacle avoidance safe distances for the first standard motion direction in the same way as obtaining the obstacle avoidance safe distances for the other standard motion directions.