Multi-mode unmanned aerial vehicle perception data fusion analysis method
Through multimodal sensor data fusion and edge computing closed-loop decision-making, the target recognition accuracy and space-time registration problems of the UAV perception system in complex environments are solved, and high-precision target tracking and rapid response are achieved, which are suitable for environmental monitoring and sea area supervision.
Patent Information
- Application Number
- CN202510393393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional drone perception systems have insufficient target recognition accuracy, large spatial and temporal registration errors, poor dynamic target tracking stability, and high decision-making response delays in complex environments, making it difficult to meet the real-time needs of scenarios such as environmental monitoring and sea area supervision.
Multimodal sensor data fusion, Beidou-3 RTK+IMU high-precision spatiotemporal registration, deep learning feature extraction and fusion, focal length-resolution mapping matrix correction, edge computing closed-loop decision-making and other technologies are adopted to achieve high-precision spatiotemporal registration, intelligent multimodal data fusion and low-latency decision-making.
It improves the accuracy of target recognition and tracking stability, shortens the task response time, and enhances the autonomous decision-making ability and system adaptability of the drone in complex environments.
Smart Images

Figure CN120337132A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) perception, and particularly to a multi-modal UAV perception data fusion and analysis method. Background Art
[0002] With the increasing demand for high-precision perception and real-time decision-making in fields such as environmental monitoring, smart cities, and maritime area supervision, multi-modal data fusion and intelligent target perception technologies have become research hotspots. Traditional UAV perception systems usually rely on a single sensor (such as a visible light camera or radar), which are vulnerable to factors such as illumination changes and occlusion interference in complex environments, resulting in limited perception capabilities. At the same time, existing multi-modal fusion methods still have technical bottlenecks in terms of target recognition, tracking, and spatio-temporal registration accuracy in dynamic scenarios.
[0003] Existing technologies mainly adopt perception methods based on single modality or rule matching, such as target recognition through traditional visual detection algorithms or target tracking relying on radar point clouds. However, these methods exhibit insufficient robustness in complex environments. Especially in adverse conditions such as low light, haze, and reflection, the perception capabilities of a single sensor are often insufficient, leading to a decline in target recognition accuracy. In addition, traditional multi-modal data fusion methods often rely on large-scale data annotation and model training, and it is difficult to respond in real time in resource-constrained scenarios.
[0004] In the field of UAV perception, spatio-temporal registration accuracy directly affects the reliability of data fusion. Traditional navigation devices rely on single-frequency GPS or inertial measurement units (IMUs), with limited accuracy, resulting in deviations in the spatio-temporal synchronization between UAV-captured images and geographical coordinates, thus affecting the accuracy of target positioning and tracking. At the same time, the real-time performance and accuracy of existing target detection and tracking technologies are difficult to balance in dynamic scenarios. Especially during the flight of UAVs, the problem of target scale distortion caused by focal length changes has not been effectively solved.
[0005] In addition, most existing decision-making systems rely on cloud computing for task scheduling and data processing. Limited by communication bandwidth and transmission delay, it is difficult to meet the requirements of real-time applications. Although edge computing technology has been explored in some applications, the current UAV perception system has not formed an efficient "perception - analysis - decision" closed-loop mechanism, making it difficult for the system to respond quickly to emergencies.
[0006] Therefore, there is an urgent need for a UAV perception method that integrates high-precision spatio-temporal registration, intelligent multi-modal data fusion, dynamic target tracking, and low-latency closed-loop decision-making to improve the target recognition accuracy and real-time response ability in complex environments and meet the actual needs of scenarios such as environmental monitoring and maritime area supervision. Summary of the Invention
[0007] In view of this, the present invention provides a method for multi-modal UAV perception data fusion and analysis to solve the problems of insufficient target recognition accuracy, large spatio-temporal registration error, poor dynamic target tracking stability, and high decision-making response delay in the existing UAV perception system, and improve the robustness of multi-modal data fusion and the autonomous decision-making ability of the UAV.
[0008] The present invention provides a method for multi-modal UAV perception data fusion and analysis, and the method includes:
[0009] Obtain multi-modal sensor data such as visible light cameras, infrared thermal imaging cameras, SAR radars, and AIS receivers carried by the UAV, and preprocess the data;
[0010] Perform spatio-temporal registration by using Beidou-3 / GPS dual-frequency RTK positioning combined with IMU measurement data, synchronize the image frame timestamps through the Beidou timing module, and standardize the coordinates by using UTM projection transformation to achieve high-precision spatio-temporal alignment;
[0011] Use deep learning methods to extract and fuse features from multi-modal data, including performing target detection through a convolutional neural network (CNN), optimizing target recognition by combining infrared and visible light information, and using an adaptive feature correction algorithm to reduce environmental interference;
[0012] Use a focal length-resolution mapping matrix to correct the distortion of the target size caused by zooming, and combine Kalman filtering and the Hungarian algorithm for target tracking to improve the detection stability of small targets in a dynamic environment;
[0013] Construct an edge computing unit, and use Bayesian optimization and Markov decision process for real-time task scheduling, so that the UAV can complete task decision-making without relying on the cloud, and achieve low-latency closed-loop response;
[0014] Transmit data through a 5G high-speed communication link or Beidou short message to ensure the stability of remote command and dispatch in a network-constrained environment.
[0015] The beneficial effects of the present invention are as follows:
[0016] Through Beidou-3 RTK high-precision positioning and IMU attitude compensation, centimeter-level high-precision spatio-temporal registration is achieved, effectively reducing the synchronization error of UAV perception data and improving the reliability of data fusion.
[0017] Adopt deep learning and multi-modal fusion to improve the accuracy of target recognition in different environments, combine adaptive feature correction technology to optimize the target detection ability under occlusion and illumination changes, and increase the recognition accuracy rate to more than 95%.
[0018] By combining the focal length-resolution mapping matrix and the dynamic target tracking algorithm, the accuracy of target size calculation in the zoom scenario is improved, the size distortion problem is reduced, and the target tracking becomes more stable.
[0019] Through the edge computing dynamic decision-making engine, the autonomous mission planning ability of the UAV is improved, the mission response time is shortened from the 10-second level to within 1 second, the dependence on the ground control center is reduced, and the execution efficiency of emergency missions is improved.
[0020] Adopting a 5G / Beidou short message dual-channel communication scheme to ensure the stability of data transmission, enabling the UAV to perform monitoring tasks in remote or low-bandwidth environments, and improving the adaptability and reliability of the system.
[0021] The present invention solves the problems of large spatio-temporal registration errors, poor target recognition robustness, easy drift in dynamic target tracking, and insufficient real-time decision-making ability of traditional UAV perception systems in complex environments, and provides a high-precision and high-efficiency UAV perception solution for environmental monitoring, sea area supervision, and emergency response. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the flowchart of the method of the present invention;
[0023] Figure 2 is the schematic diagram of the high-precision spatio-temporal registration process of the present invention;
[0024] Figure 3 is the flowchart of the target detection and dynamic correction algorithm of the present invention;
[0025] Figure 4 is the calibration curve graph of the focal length-resolution mapping matrix;
[0026] Figure 5 is the flowchart of the operation of the edge-side closed-loop decision-making engine. DETAILED DESCRIPTION OF THE INVENTION
[0027] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the scope of protection of the present invention more clearly defined.
[0028] The present invention proposes a multi-modal UAV perception data fusion and analysis method, which is mainly applied to scenarios such as environmental monitoring, sea area supervision, and emergency response. By integrating multi-modal sensors, optimizing data fusion technology, improving target tracking accuracy, and combining edge computing to achieve low-latency closed-loop decision-making.
[0029] The core technical solution of the present invention is based on deep learning and sensor data fusion. Through the correction of Beidou-3 RTK + IMU, high-precision spatio-temporal registration is carried out to ensure the spatial and temporal alignment of different modality data, and improve the accuracy of target recognition.
[0030] This solution uses multi-modal data fusion technology, combines visible light cameras, infrared thermal imaging cameras, SAR radars and AIS data to achieve cross-modal feature extraction, and uses deep learning models to optimize target detection and classification, improving the robustness of the perception system.
[0031] The present invention uses a focal length-resolution mapping matrix for scale correction, dynamically adjusts the target size calculation in the zoom scenario, reduces the detection error caused by the change of focal length, and ensures that the tracking accuracy is stable in a high-dynamic environment.
[0032] To improve the stability of target tracking, the present invention combines the Kalman filter and the Hungarian algorithm to achieve multi-frame tracking and matching of the target, improves the small target detection ability of the UAV in a complex environment, and reduces the false detection rate caused by occlusion and noise interference.
[0033] The present invention introduces an edge computing closed-loop decision-making engine, adopts Bayesian optimization and Markov decision-making process to achieve autonomous path planning and task scheduling of the UAV, reduces the dependence on the cloud, improves the real-time performance, and shortens the decision response time to within 1 second.
[0034] The present invention uses 5G / Beidou short message dual-channel communication to ensure that the UAV can still transmit key data in a low-bandwidth environment, improves the stability of remote task execution, and makes the UAV perception system have higher adaptability and reliability.
[0035] The present invention proposes a high-precision spatio-temporal registration method. By combining Beidou-3 RTK, IMU correction and time synchronization technology, centimeter-level accuracy alignment of UAV multi-modal perception data is achieved, and the accuracy of target detection and tracking is improved.
[0036] This method mainly includes three key steps: time synchronization, coordinate projection transformation and attitude compensation to ensure the spatio-temporal consistency of UAV multi-modal sensor data.
[0037] The present invention uses the Beidou-3 timing module to synchronize the timestamps of UAV image frames, controls the error within 1 ms, and ensures the time alignment of multi-modal sensor data.
[0038] Let the time provided by the GPS timing signal of the UAV be T GPS , the timestamp of the optical sensor be T cam , the timestamp of the infrared sensor be T IR , then the time synchronization satisfies:
[0039] T sync = T GPS + ΔT
[0040] where ΔT is the timing error, and the present invention makes it satisfy, through high-precision timing technology:
[0041] |ΔT| < 1 ms
[0042] to ensure that all sensor data is based on the same time reference frame.
[0043] The target coordinates obtained by the UAV sensor are usually recorded in the WGS-84 longitude and latitude format, while actual perception and fusion analysis require the use of a local plane coordinate system (UTM). The present invention uses UTM projection transformation for coordinate conversion.
[0044] Let the geographical coordinates of the target point be (λ, φ, h), where λ is the longitude, φ is the latitude, and h is the altitude. The converted UTM plane coordinates are:
[0045] X = k(λ - λ0)cosφ
[0046] Y = k(φ - φ0)
[0047] where λ0, φ0 are the reference longitude and latitude, and k is the projection scale factor.
[0048] Through high-precision RTK measurement, the present invention ensures that the UTM coordinate conversion error is less than 0.1 meter.
[0049] During the flight of the UAV, pitch, roll, and yaw motions occur, resulting in the deviation of the camera view angle, thus affecting the accuracy of the target position information. The present invention uses IMU sensor data for attitude compensation to ensure the precise alignment of visual data with geographical coordinates.
[0050] Let the point P u (x u , y u , z u ) in the UAV coordinate system have the transformation matrix in the geographical coordinate system as:
[0051] P g = RP u + T
[0052] where P g (x g , y g , z g ) is the point in the geographical coordinate system, R is the rotation transformation matrix, and T is the translation vector.
[0053] The rotation transformation matrix R is calculated from the IMU data:
[0054]
[0055] Among them, ψ is the yaw angle, θ is the pitch angle, and φ is the roll angle.
[0056] The present invention utilizes this attitude compensation matrix to ensure that the data acquired by the sensor is always consistent with the true geographical location, effectively reducing the errors caused by the changes in the flight attitude of the unmanned aerial vehicle.
[0057] In order to verify the high-precision spatio-temporal registration method of the present invention, we use the error analysis method to calculate the spatio-temporal alignment error:
[0058] E total = E time + E coord + E attitude
[0059] Where:
[0060] The time error E time = |ΔT|, which is controlled within 1 ms in the present invention;
[0061] The coordinate error E coord Is determined by the GPS-RTK accuracy and is less than 0.1 m after optimization in the present invention;
[0062] The attitude error E attitude Is calculated by the IMU and is less than 0.05° after optimization in the present invention.
[0063] Through the above method, the present invention effectively reduces the spatio-temporal synchronization error of the perception data of the unmanned aerial vehicle, improves the accuracy of multi-modal data fusion, and makes the target recognition and tracking more stable and reliable.
[0064] After completing the high-precision spatio-temporal registration, the present invention proposes a target detection and dynamic correction method based on multi-modal data fusion to improve the target recognition accuracy and tracking stability of the unmanned aerial vehicle in complex environments.
[0065] This method combines multi-modal sensor data such as visible light, infrared thermal imaging, and SAR radar, and uses a deep learning model for target detection to improve the adaptability and recognition accuracy in different environments.
[0066] After the target detection is completed, the present invention further introduces a scale dynamic correction method to compensate for the target size error caused by the zoom of the unmanned aerial vehicle and ensure the detection accuracy of the target at different focal lengths.
[0067] To improve the tracking stability, the present invention uses the Kalman filter and the Hungarian algorithm for target tracking to ensure the cross-frame consistency in a dynamic environment and reduce the target drift phenomenon.
[0068] This invention uses a Convolutional Neural Network (CNN) for object detection and combines a multi-modal data fusion method to improve detection accuracy and environmental adaptability.
[0069] Let the pixel representation of the input image be I(x, y, c), where (x, y) are the image coordinates and c is the color channel (RGB or infrared thermal imaging). The feature extraction process of the object detection network can be expressed as:
[0070] F = f CNN (I(x, y, c))
[0071] where F represents the extracted feature map and f CNN is the convolutional neural network model.
[0072] For object detection, a multi-modal data fusion method is adopted. Let the visible light image feature be F vis , the infrared image feature be F IR , and the SAR image feature be F SAR . The final fused feature is expressed as:
[0073] F fusion = w vis F vis + w IR F IR + w SAR F SAR
[0074] where w vis , w IR , w SAR are the weight coefficients of each modal data and are determined through training optimization.
[0075] The object detection network outputs the bounding box and class label of the object. Let the object detection result be:
[0076] B = (x min , y min , x max , y max , s)
[0077] where (x min , y min ) and (x max , y max ) are the upper left and lower right coordinates of the object respectively, and s is the object confidence.
[0078] Since the zoom operation of the drone will cause the change of the object size in the image, if not corrected, the detection result may have scale distortion.
[0079] The present invention uses a focal length-resolution mapping matrix for scale dynamic correction to ensure that the target size remains accurate under different focal length conditions.
[0080] Let the focal length of the current camera be f, and the actual size corresponding to a unit pixel be d(f). Then the formula for calculating the actual target size is:
[0081] S actual =S pixel ·d(f)
[0082] Where S pixel is the target size in pixel coordinates, and S actual is the actual physical size.
[0083] The focal length-resolution mapping matrix of the present invention is pre-established through a calibration experiment, and its mapping relationship can be expressed as:
[0084] d(f)=af 2 +bf+c
[0085] Where a, b, and c are coefficients obtained through calibration.
[0086] Through this mapping matrix, the present invention can dynamically correct the target size according to the zoom situation of the drone, improving the accuracy of target detection.
[0087] The present invention adopts a target tracking method that combines the Kalman filter and the Hungarian algorithm to improve the stability of target tracking and reduce the target drift phenomenon.
[0088] Let the state vector of the target be expressed as:
[0089] X k =[x k ,y k ,v x ,v y T
[0090] Where (x k ,y k ) is the target center coordinate, and (v x ,v y ) is the target velocity component.
[0091] The update of the target state follows the Kalman filter prediction model:
[0092] X k+1 =AX k +W k
[0093] Where A is the state transition matrix, and W k is the process noise.
[0094] The target detection results and prediction results are matched by the Hungarian algorithm to ensure the stability of cross-frame ID association.
[0095] Let the set of detected targets be {B i}, and the set of tracking prediction targets be {P j}. The cost matrix of the Hungarian algorithm is defined as:
[0096] C(i,j) = |B i - P j ||
[0097] The target matching is solved by minimizing the cost matrix C to ensure the optimal allocation.
[0098] To optimize the performance of target detection and dynamic correction, the present invention analyzes the error sources and proposes optimization strategies.
[0099] The target detection error E detect is caused by the uncertainty of the deep learning model. The present invention adopts a high-confidence filtering strategy to reduce the detection error:
[0100] E detect = 1 - s max
[0101] where s max is the confidence of the target with the highest confidence.
[0102] The target scale error E scale is caused by the focal length-resolution mapping error. After optimization, the error is controlled within 5%:
[0103] E scale = |S actual - S predicted |
[0104] The target tracking error E track is caused by the cross-frame matching error. After optimization by the present invention, the error is reduced to within 2%:
[0105] E track = ||X tracked - X true ||
[0106] Through the above optimizations, the target detection and dynamic correction scheme of the present invention effectively improves the target recognition and tracking accuracy of the drone in a dynamic environment, and improves the robustness and stability of the system.
[0107] The present invention proposes a calibration method based on the focal length-resolution mapping matrix to compensate for the change in target size during the zooming process of the drone and improve the accuracy of target detection and tracking.
[0108] During the target detection and dynamic correction process, the zoom operation of the drone may cause the size of the target to change in the image. If no correction is made, scale errors may be introduced, affecting the stability and consistency of the detection.
[0109] To solve this problem, the present invention calibrates the focal length-resolution mapping relationship through experiments and establishes a mathematical model to dynamically calculate the actual size of the target.
[0110] The focal length-resolution calibration method of the present invention mainly includes three key steps: data acquisition, model fitting, and error optimization.
[0111] In the experimental environment of the present invention, standard targets at different focal lengths are photographed, and the camera focal length, pixel size, and actual physical size of the target are recorded.
[0112] Let the focal length of the camera be f, and the actual size corresponding to a unit pixel at this focal length be d(f). Then the relationship between the pixel size and the physical size is:
[0113] S actual =S pixel ·d(f)
[0114] Where S actual is the true size of the target, and S pixel is the pixel size in the image.
[0115] By measuring the pixel size of the same target at different focal lengths, a series of (f, d(f)) data points can be obtained, providing training data for model fitting.
[0116] The present invention uses a quadratic curve fitting method to model the relationship between the focal length f and the unit pixel size d(f).
[0117] Let the focal length-resolution mapping relationship be described by the following equation:
[0118] d(f)=af 2 +bf+c
[0119] Where a, b, and c are fitting parameters, calculated by the least squares method.
[0120] Specifically, the goal of the least squares method is to minimize the following error function:
[0121]
[0122] Where N is the number of data points collected in the experiment, f i and d(f i ) are the experimentally measured focal length and the corresponding unit pixel size, respectively.
[0123] By solving for a, b, and c, the present invention obtains a focal length-resolution mapping curve, enabling accurate calculation of the target size under different focal length conditions.
[0124] Due to the existence of experimental measurement errors, there may be certain deviations in the calibration curve. The present invention improves the fitting accuracy through an error optimization method.
[0125] Let the fitted unit pixel size be d pred (f), and the true unit pixel size measured experimentally be d true (f). Then the error is defined as:
[0126]
[0127] The present invention adopts polynomial regularization technology to optimize the fitting curve, reduce overfitting, improve the generalization ability, and make the error E scale controlled within 5%.
[0128] Through this optimization, the present invention can, in practical applications, calculate the target size in real time according to the zoom situation of the drone, improving the accuracy of target detection and tracking.
[0129] The focal length-resolution mapping matrix of the present invention plays a key role in the process of target detection and dynamic correction, enabling the detection system to adapt to the target scale changes under different focal lengths.
[0130] In the drone perception system, this method combines target detection and dynamic correction algorithms to achieve dynamic compensation, enabling the target size error to remain stable under zoom conditions.
[0131] Through the above method, the present invention improves the target detection accuracy of the drone perception system under different focal length conditions and enhances the adaptability of the system in complex scenarios.
[0132] After completing target detection, dynamic correction, and focal length-resolution mapping correction, the present invention proposes a closed-loop decision-making method based on edge computing to improve the real-time performance and autonomous task execution ability of the drone perception system.
[0133] Since the drone may face problems such as limited communication and decision-making delay in complex environments, relying on cloud computing for task scheduling and decision-making may lead to a decline in response speed and affect task execution efficiency.
[0134] To solve this problem, the present invention deploys an intelligent decision-making engine at the edge computing end, combines Bayesian optimization and Markov decision-making process to achieve low-latency closed-loop scheduling, and improves the autonomous decision-making ability of the drone.
[0135] The edge computing closed-loop decision-making method of the present invention mainly includes four key steps: task modeling, path optimization, dynamic task scheduling, and low-latency response.
[0136] The present invention first performs mathematical modeling on the tasks of the unmanned aerial vehicle (UAV) to optimize the decision-making process and improve the computing efficiency.
[0137] Let the UAV task space consist of multiple observed target points {s1, s2,..., s n}, and each target point s i has a task weight w i , indicating its importance level.
[0138] The task set can be expressed as:
[0139]
[0140] The goal of task execution is to minimize the total path cost C total , while maximizing the priority and benefit of task completion:
[0141]
[0142] where λ is a trade-off factor used to balance the task priority and the path length.
[0143] To optimize the task path of the UAV, the present invention uses the Markov decision process (MDP) to solve the optimal path.
[0144] Let the path planning problem of the UAV be represented as an MDP, defined as follows:
[0145]
[0146] where:
[0147] is the set of task space states;
[0148] is the set of actions that the UAV can execute;
[0149] P(s′|s, a) is the state transition probability;
[0150] R(s, a) is the task reward function;
[0151] γ ∈ (0, 1) is the discount factor.
[0152] The task reward function R(s, a) is defined as:
[0153] R(s,a) = w s - λC(s,a)
[0154] Among them, w s is the weight of the current task point, and C(s, a) is the cost required to execute action a (such as path distance, energy consumption, etc.).
[0155] By solving the value function of the MDP:
[0156]
[0157] The optimal path planning strategy can be obtained.
[0158] During the execution of the drone task, the environment may change, so it is necessary to dynamically adjust the task scheduling strategy.
[0159] The present invention adopts the Bayesian optimization algorithm to dynamically update the scheduling strategy according to the task completion situation, improving the task execution efficiency.
[0160] Let the objective function of the task scheduling be:
[0161]
[0162] Among them, x represents the task scheduling parameters, including path adjustment, task priority update, etc.
[0163] Through Bayesian optimization, the objective function is iteratively solved:
[0164]
[0165] Among them, D is the historical task data.
[0166] Through this optimization method, the present invention can dynamically adjust the scheduling strategy during the task execution, enabling the drone to adapt to environmental changes and improving the task execution efficiency.
[0167] In order to ensure the real-time performance of the task execution, the present invention uses the edge computing terminal for fast calculation and reduces the dependence on the cloud.
[0168] Let the task calculation time be T compute Consisting of the local calculation time T edge and the cloud calculation time T cloud :
[0169] T compute = min(T edge , T cloud )
[0170] Since the latency of edge computing is usually much lower than that of cloud computing, that is, T edge << T cloud , therefore, the present invention preferentially uses edge computing for task scheduling.
[0171] Through experimental verification, the edge computing closed-loop decision-making method of the present invention can shorten the task response time to within 1 second, improving the real-time performance compared with the traditional cloud computing-based method.
[0172] The edge computing closed-loop decision-making method of the present invention is combined with object detection, dynamic correction, and focal length-resolution mapping to achieve a complete autonomous task execution process for drones.
[0173] This method can be widely applied to scenarios such as environmental monitoring, sea area supervision, and emergency response, improving the intelligence level and task execution efficiency of drones.
[0174] Through the above method, the edge computing closed-loop decision-making system of the present invention can efficiently complete task scheduling in a low-latency environment, ensuring that drones have strong autonomous decision-making capabilities in complex environments.
[0175] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the technical solution of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-modal UAV perception data fusion and analysis method, characterized in that The method includes: Collecting target area data through multi-modal sensors such as visible light cameras, infrared thermal imaging cameras, SAR radars, and AIS receivers; Performing spatio-temporal registration by combining Beidou-3 / GPS dual-frequency RTK positioning with IMU data to achieve centimeter-level data synchronization accuracy; Implementing data synchronization between the cloud and the edge through a 5G / Beidou short message communication link to ensure low-latency response.
2. The method according to claim 1, wherein The multi-modal data fusion includes: Using a deep learning model to fuse multi-modal data, including feature extraction, weight assignment, and adaptive correction, to improve the accuracy of target recognition; Combining visible light and infrared data to enhance the detection ability of occluded targets; Using an adaptive feature correction algorithm to reduce the impact of complex environments (such as strong light and water surface reflection) on the detection results and improve robustness.
3. The method according to claim 2, characterized in that The target detection and tracking includes: Dynamically correcting the distortion of target size caused by zooming using a focal length-resolution mapping matrix; Combining Kalman filtering and the Hungarian algorithm for target tracking to improve the stability of small targets; Executing dynamic decisions through an edge computing unit, generating task scheduling instructions based on Bayesian optimization and Markov decision process, enabling the drone to autonomously execute tasks.
4. The method according to claim 1, wherein The spatio-temporal registration process includes: Synchronizing the timestamps of drone image frames using a Beidou timing module to make the error less than 1ms; Converting geographical coordinates to local plane coordinates through UTM projection to ensure the matching accuracy of drone sensing data and geographical information; Performing attitude compensation by combining IMU attitude data (pitch angle, roll angle) to reduce data offset caused by drone attitude changes.
5. The method according to claim 2, wherein The target tracking uses a multi-algorithm fusion method, including: Predicting target motion by combining Kalman filtering to reduce tracking drift; Using the Hungarian algorithm for inter-frame target matching to ensure target ID consistency; Dynamically compensating for the change in target size caused by zooming through a focal length-resolution mapping matrix to control the target size error within ±5%.
6. The method according to claim 3, wherein The dynamic decision engine uses Bayesian optimization and Markov decision process, including: Calculating the optimal path based on target type, density distribution, and priority; Optimizing drone task scheduling using reinforcement learning to improve drone autonomy; Implementing low-latency closed-loop control through an edge computing unit to reduce cloud dependence and make the decision response time less than 1 second.
7. The method according to claim 1, characterized in that, The communication link uses a 5G / Beidou short message dual-channel mode, including: 5G high-speed transmission for real-time data synchronization; Beidou short message for key data transmission in low-bandwidth environments to ensure the reliability of remote task scheduling.
8. The method according to claim 1, wherein The method is applicable to environmental monitoring, sea area supervision, and emergency pollution event response scenarios. Through multi-modal data fusion and autonomous decision optimization, the target recognition accuracy is improved to over 95%, and the response time is shortened to within 1 second.
Citation Information
Cited By
Unmanned aerial vehicle dynamic path control method based on multi-source sensing fusion
CN120722928A
Unmanned aerial vehicle positioning method, device and equipment based on sequence observation and medium
CN120831630A
Multi-source information fusion unmanned aerial vehicle offshore target tracking method and system
CN121277220A
Multi-source information fusion unmanned aerial vehicle sea target tracking method and system
CN121277220B