Adaptive fusion multi-mode unmanned aerial vehicle target tracking method and tracking system

Through adaptive fusion of multimodal data and reinforcement learning optimization Transformer technology, high-precision target tracking of drones in complex environments is achieved, solving the shortcomings of tracking accuracy and stability of existing technologies in complex environments.

CN120070501APending Publication Date: 2025-05-30XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510114356.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing UAV tracking methods are difficult to achieve high-precision, high-stability and robust target tracking in complex and changeable environments.

Method used

Adaptive fusion multimodal drone target tracking method is adopted to collect visual, infrared and lidar data through sensors mounted on the drone, perform time synchronization and noise removal, extract environmental feature vectors, and perform multimodal feature fusion through weighted average and Kalman filtering. Then, multi-scale feature fusion is used to use the Transformer encoder to optimize feature fusion methods, and dynamic environment adaptation is achieved by combining reinforcement learning and meta-learning techniques.

Benefits of technology

It significantly improves the target recognition accuracy and tracking robustness of the drone in complex environments, and can quickly adapt to environmental changes and continuously optimize system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070501A_ABST
    Figure CN120070501A_ABST
Patent Text Reader

Abstract

The invention relates to a self-adaptive fusion multi-mode unmanned aerial vehicle target tracking method, which comprises the following steps of 1, performing data information acquisition including visual information, infrared information and laser radar information of a sensor; 2, performing time synchronization on the data information, and removing noise through a filtering algorithm; step 3, acquiring an environment feature vector E of the data information; step 4, fusing multi-modal features to obtain a comprehensive feature vector; and 5, extracting a plurality of features of different scales from the comprehensive feature vector, inputting the features into a Transform encoder to obtain fusion features, and performing target tracking based on the obtained fusion features. According to the adaptive fusion multi-modal unmanned aerial vehicle target tracking method and tracking system, multi-modal feature fusion is carried out in combination with sensor data, the target recognition precision of a small unmanned aerial vehicle in a complex environment can be remarkably improved through modular and hierarchical design, and the environment sensing ability of the system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and a system for tracking an unmanned aerial vehicle (UAV) target, and more particularly to a method and a system for multi-modal UAV target tracking with adaptive fusion. Background Art

[0002] With the rapid development of UAV technology, target tracking has been increasingly widely used in fields such as security monitoring, logistics distribution, and environmental monitoring. There are still many challenges in the existing target tracking methods for dealing with small targets, complex environmental changes, and multi-sensor information fusion: Generally speaking, there are two mainstream methods in the field of UAV tracking, namely the method based on correlation filtering (CF) and the method based on machine learning (ML). The online tracker of the method based on correlation filtering (CF) is widely adopted because of its low computational complexity. Although it is very efficient, the tracker of the method based on correlation filtering is difficult to meet the tracking requirements of UAVs in complex scenarios in terms of accuracy and robustness. Especially under complex conditions such as light changes, target occlusion, and camera movement, the tracking accuracy and real-time performance of traditional models are difficult to meet the actual needs. The method based on machine learning (ML) constructs a model by learning patterns from data, can handle complex non-linear relationships, and can provide high-precision prediction and classification results. However, machine learning methods may require a large amount of data to train the model, and the training and tuning of the model may be complex and time-consuming. Therefore, there is an urgent need for a UAV target tracking method that can have high precision, high stability, and robustness in a complex and changeable environment. Summary of the Invention

[0003] The object of the present invention is to solve the technical problem that the existing UAV tracking methods are difficult to adapt to complex and changeable environments, and to provide a method and a system for multi-modal UAV target tracking with adaptive fusion, which combines sensor data for multi-modal feature fusion, and can significantly improve the target recognition accuracy of small UAVs in complex environments through modular and hierarchical designs, and enhance the environmental perception ability of the system.

[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] A method for multi-modal UAV target tracking with adaptive fusion, which is characterized in that it includes the following steps:

[0006] Step 1, collecting data information through sensors carried by the UAV, where the data information includes visual information, infrared information, and lidar information of the sensors;

[0007] Step 2, synchronizing the visual information, infrared information, and lidar information of the sensors in time, and removing noise through a filtering algorithm;

[0008] Step 3, obtain the environmental feature vector E of the data information:

[0009]

[0010] Among them, E vision , E infrared , E lidar are the feature vectors of visual information, infrared information, and lidar information respectively;

[0011] Step 4, perform feature extraction from the environmental feature vector E, and use the weighted average method to fuse multi-modal features to obtain a comprehensive feature vector;

[0012] Step 5, extract multiple features of different scales from the comprehensive feature vector, input the extracted multi-scale features into the Transformer encoder, perform multi-scale feature fusion to obtain the fused feature, and perform target tracking based on the obtained fused feature.

[0013] Furthermore, the feature vectors E vision , E infrared , E lidar obtained in step 3 are specifically:

[0014]

[0015] Among them are the inputs from visual information, infrared information, and lidar respectively;

[0016] Furthermore, the specific steps of step 4 are:

[0017] Step 4.1, perform multi-modal fusion on the extracted feature vectors through the weighted average method to obtain the multi-modal fusion feature vector X′k:

[0018] X′ k =W 1 E vision +W 2 E lidar +W 3 E infrared

[0019] Among them, W 1 , W 2 , W 3 are the weights of the feature vectors of visual information, lidar information, and infrared information respectively;

[0020] Step 4.2, update the state of the Kalman filter, and use the Kalman filter to filter the multi-modal fusion feature vector to obtain the comprehensive feature vector.

[0021] Furthermore, the specific steps of step 5 are:

[0022] Step 5.1: Select convolution kernels of different sizes from the comprehensive feature vector, extract features of multiple different scales, and generate feature maps F_scaleN of different scales;

[0023] F scaleN = Conv(F, K a×a )

[0024] where K a×a is a convolution kernel of size a×a; N is the number of indexing times, and F is the input;

[0025] Step 5.2: Concatenate the feature maps F scaleN of different scales in the channel dimension to form a tensor F multiscale that contains multi-scale features;

[0026] F multis cale = concat(F scale1 , F scale2 …F scaleN );

[0027] Step 5.3: Input the multi-scale features into the Transformer encoder for multi-scale feature fusion to obtain the fused feature T:

[0028]

[0029] where Q, K, and V are the query vector, key vector, and value vector respectively; d k is a hyperparameter;

[0030] Step 5.4: Perform feature integration. Use the fused feature T output by the Transformer as the unified target feature for object tracking.

[0031] Furthermore, in step 4.3, the specific update of the Kalman filter state is as follows:

[0032] Step 4.3.1: Use the Kalman filter KalmanFilter to obtain the initial position x 0 through feature point matching and motion estimation information:

[0033] x 0 = KalmanFilter(x vision , x infrared , X lidar )

[0034] where x vision is the feature of visual information, x infrared is the feature of infrared information, x lidarDepth information features provided for lidar;

[0035] Step 4.3.2, predict the position of the target at the next time step to obtain the predicted target position

[0036]

[0037] where A is the state transition matrix and Bu k is the input control term;

[0038] Step 4.3.2, compare the actually observed target position with the predicted position, calculate the error and update the state of the Kalman filter:

[0039] K k = P k|k-1 H T (HP k|k-1 H T + R) -1

[0040]

[0041] where K k is the Kalman gain, z k is the observed position, and P k|k-1 is the state covariance matrix;

[0042] H is the initially known observation matrix; R is the initially known observation noise covariance matrix; is the predicted state.

[0043] Furthermore, in step 5.3, during multi-scale feature fusion, continuous iterative trial-and-error learning is achieved through the reinforcement learning function to select the optimal feature fusion method and optimize the fused features.

[0044] Furthermore, in step 5.3, during multi-scale feature fusion, self-supervised loss calculation is achieved through the self-supervised motion compensation function; self-supervised loss calculation is used for motion estimation, and the motion compensation of the target is found by minimizing the loss function; combined with the self-supervised learning results, the target position estimation is adjusted to obtain the motion compensation amount.

[0045] An adaptive fusion multi-modal UAV target tracking system, which is characterized in that it includes a data acquisition and processing module, a multi-modal feature fusion module, and a Transformer processing module based on reinforcement learning; the data acquisition and processing module is used to collect data information sent by sensors carried by the UAV and extract features; the multi-modal feature fusion module performs weighted fusion on the extracted features of the data acquisition and processing module; the Transformer processing module based on reinforcement learning performs fusion of different scales on the fusion features of the multi-modal feature fusion module, and the fused features are transmitted to the multi-modal feature fusion module to output a prediction result.

[0046] Furthermore, it also includes a meta-learning online model update module and a spatio-temporal self-supervised motion compensation module;

[0047] The meta-learning online model update module is used to optimize the feature fusion of the multi-modal feature fusion module and the Transformer processing module based on reinforcement learning; the spatio-temporal self-supervised motion compensation module analyzes and processes the fusion features of the Transformer processing module based on reinforcement learning, calculates the self-supervised loss and gives the motion compensation amount.

[0048] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0049] 1. The adaptive fusion multi-modal UAV target tracking method and tracking system of the present invention combines sensor data for multi-modal feature fusion, and performs multi-scale feature fusion of Transformer optimized based on reinforcement learning, enhancing the environmental perception ability of the system. Especially in complex scenarios such as target occlusion and light change, it greatly improves the robustness and stability of tracking; when the target motion changes violently or the environment changes suddenly, the system can effectively perform dynamic compensation, can also quickly adapt to the new environment and changing scenarios, and continuously optimize the system performance. This mechanism enhances the adaptability and scalability of the model in a dynamic environment;

[0050] 2. The adaptive fusion multi-modal UAV target tracking method and tracking system of the present invention comprehensively tracks the model system through the integration and innovation of multiple technologies such as meta-learning online update and spatio-temporal self-supervised motion compensation, and can significantly improve the target recognition accuracy of small UAVs in complex environments. Description of the Drawings

[0051] Figure 1 It is a schematic diagram of an embodiment of the adaptive fusion multi-modal UAV target tracking method of the present invention. Detailed Embodiment

[0052] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention, and the purpose is not to limit the protection scope of the present invention.

[0053] As Figure 1 shown, an adaptive fusion multi-modal UAV target tracking method of the present invention aims at the position of the target to be tracked in each frame of image, and is represented by coordinates as:

[0054]

[0055] where k is the frame number index, and this model mainly uses visual, infrared and lidar data for target tracking. The specific implementation method is as follows:

[0056] 1. Information acquisition:

[0057] First, data information is collected through sensors carried by the UAV. The UAV is equipped with a high-resolution camera, an infrared sensor and a lidar to obtain information such as images, thermal images and depths of the target respectively. At the same time, environmental parameter data is used by sensors such as light to monitor some parameters of the environment where the UAV is located in real time. In order to ensure the consistency of the data, it is necessary to synchronize the time of different sensor data, and remove noise through algorithms such as filtering, and perform standardization processing to unify the scale, etc.

[0058] 2. Perform multi-modal feature fusion

[0059] Environmental feature vectors such as are generated by collecting data through sensors carried by the UAV:

[0060]

[0061] where E vision 、E infrared 、E lidar 、E other are the feature vectors of visual information, infrared information, lidar information and other information respectively. Other information can be light information, etc.

[0062] Then, features are extracted from the environmental vectors, and a comprehensive feature vector is obtained through weighted fusion.

[0063] First, feature extraction is performed:

[0064]

[0065] where E vision is the visual information feature (edge, color, etc. information) jointly affected by OSNet and StrongSort, and F infrared extracted by a dedicated infrared sensoris the infrared information feature (thermal imaging contour, temperature difference, etc.), F lidar is the depth information feature (shape feature, point cloud data, etc.), where are the inputs from visual information, infrared information, and lidar information respectively.

[0066] After extraction, the weighted average method can be used to fuse multi-modal features. The weights are determined by the sensor noise level σ. The noise level determines the performance of the sensor, and the weights are adjusted according to the sensor noise level to influence the contribution of each sensor data to the final fusion result. Sensors with less noise (i.e., higher accuracy sensors) will have greater weights, thus providing more contributions during the fusion process and improving the accuracy and reliability of the final fused feature vector X′ k .

[0067] The weight calculation method is as follows:

[0068] X′ k = W 1 E vision + W 2 E lidar + W 3 E infrared

[0069] Where:

[0070]

[0071] Among them, i and j represent the weight calculations corresponding to different sensors (vision, lidar, infrared). The role of these two indices is to identify the noise levels of each sensor and how they affect the weighted average. i is the index of the current sensor (such as vision, lidar, infrared), and j is used to sum up the noise levels of all sensors.

[0072] Finally, the Kalman filter is used to filter the fused features. Assume that we have designed a Kalman filter KalmanFilter, and the initial position x 0 and velocity v 0 are obtained through feature point matching and motion estimation:

[0073] x 0 = KalmanFilter(E vision , E lidar , E infrared )

[0074] Predict the position of the target at the next time step by combining the multi-modal features of the current time step through the Kalman filter:

[0075]

[0076] Among them, is the predicted target position, A is the state transition matrix, and Bu k is the input control term;

[0077] Finally, compare the actually observed target position with the predicted position, calculate the error, and update the state of the Kalman filter:

[0078] K k = P k|k-1 H T (HP k|k-1 H T + R) -1

[0079]

[0080] Among them, K k is the Kalman gain, z k is the observed position, P k|k-1 is the state covariance matrix, H is the observation matrix, and R is the noise covariance matrix. H (initially known): The observation matrix, which describes how to map from the state space to the observation space. Usually, it maps the state of the target (such as position, velocity) to the actual observed values (such as sensor measurements).

[0081] R (initially known): The observation noise covariance matrix, indicating the uncertainty of the observed data.

[0082] The predicted state, the estimate of the state based on previous information at time k - 1. This is obtained through the filtering prediction iteration.

[0083] 3. Perform multi-scale feature fusion based on the Transformer optimized by reinforcement learning

[0084] First, extract features of different scales from the results of step 2. Assume that the comprehensive feature vector obtained in step 2 is D-dimensional. Then, select convolution kernels of different sizes, such as 3×3, 5×5, 7×7, to extract features of different scales. Taking 3×3 as an example:

[0085] F scale1 = Conv(F, K 3×3 )

[0086] Among them, K 3×3 is the convolution kernel of size 3×3. Repeat this process and use convolution kernels of 5×5 and 7×7 for feature extraction.

[0087] Subsequently, perform feature concatenation, concatenate the feature maps of different scales on the channel dimension to form a tensor containing multi-scale features:

[0088] F multis cale = concat(F scale1 , F scale2 , F scale3 )

[0089] It is possible to downsample the concatenated features, for example, through a pooling operation, to further reduce the size of the feature map and the computational load:

[0090] F downsampled = MaxPool(F multis cale , poolsze)

[0091] Specific example:

[0092] Suppose that after feature fusion in the second step, the shape of the obtained feature map F is 128×128×64 (i.e., H = 128, W = 128, D = 64).

[0093] Extract features using a 3×3 convolutional kernel to generate the feature map F scale1 :

[0094] F scale1 shape = 128×128×64

[0095] Extract features using a 5×5 convolutional kernel to generate the feature map F scale2 :

[0096] F scale2 shape = 128×128×64

[0097] Extract features using a 7×7 convolutional kernel to generate the feature map F scale3 :

[0098] F scale3 shape = 128×128×64

[0099] After concatenation, a new feature tensor F is obtained multi_scale :

[0100] F multi_scale shape = 128×128×192

[0101] After applying the pooling operation for downsampling, assume that a downsampled feature map with a shape of 64×64×192 is obtained

[0102] F downsampledo

[0103] That is, through the above steps, an updated tensor F containing multi-scale features is obtained downsampled :

[0104] {F 1 , F2 , …, F n}

[0105] Input the multi-scale features into the Transformer encoder for global feature fusion. The calculation method is as follows:

[0106]

[0107] Among them, Q, K, and V are the query, key, and value vectors respectively, and T is the fused feature calculated by the self-attention mechanism. Finally, feature integration is performed, and the fused feature T output by the Transformer is used as the unified target feature representation. K, Q, and V are obtained through simple calculations, which are described in the following section. What is calculated according to the formula is T, the fused feature calculated by the self-attention mechanism, which is a standard mechanism widely used in the Transformer model.

[0108] Q is the query vector, used to perform a "query" operation on the input features:

[0109] Q = F multis cale W Q

[0110] Among them, W Q is the learned weight matrix, and the dimension is usually D×d k , d k is the dimension of the query vector;

[0111] K is the key vector, used to match with the query vector to calculate the correlation:

[0112] K = F multis cale W K

[0113] Among them, W K is the corresponding weight matrix;

[0114] V is the value vector, representing the actual information related to each key:

[0115] V = F multis cale W v

[0116] Among them, W V is the corresponding weight matrix;

[0117] d k is a hyperparameter, usually set when designing the Transformer, which determines the complexity and ability of the model. The selection of d k is generally related to the dimension of the input features or the complexity of the task.

[0118] Combine reinforcement learning to optimize feature fusion:

[0119] In multi-scale feature fusion, it is desired to automatically select the optimal feature fusion method so that features at different scales can maximize the accuracy of target tracking. Reinforcement Learning (RL) can find the optimal policy through continuous iterative trial-and-error learning and can be used to optimize feature fusion.

[0120] Reinforcement learning optimization framework:

[0121] State: The current combination of multi-scale features, such as the feature tensors F extracted from multiple scales multi_scale .

[0122] Action: Select the pending feature fusion method, such as different fusion methods like weighted average, concatenation, etc.

[0123] Reward: Evaluate the selected fusion method according to the performance of the fused features (such as accuracy or loss function in the target tracking task).

[0124] Reinforcement learning optimization steps:

[0125] First, define the state vector s t , s t containing the current multi-scale feature information extracted:

[0126] s t = FeatureVector(F multis cale )

[0127] Set the set of possible actions AT, for example: A 1 : weighted average, A 2 : concatenation, A 3 : select the maximum feature, A 4 : select the minimum feature.

[0128] Subsequently, define the reward function Rt to evaluate the effectiveness of the fused features:

[0129] R t = Accuracy(F fused , y true ) - λ·Loss(F fused , y true )

[0130] where y true is the true label, λ is the weight to balance accuracy and loss, and F fused is the fused feature.

[0131] Then, use the policy gradient method (such as the REINFORCE algorithm) to update the policy, and the policy function π(at |s t ) represents the probability of selecting action a in state s t The update rule, where α is the learning rate, Rt is the reward, and θ is the policy parameter: t

[0132]

[0133] Optimization steps for the fusion extraction of multi-scale features:

[0134] First, initialize the state s 0 , the policy parameter θ, and the action set A T Then, select action a from state s t : t

[0135] a t ~π(a t |s t )

[0136] Adopt the selected fusion strategy to obtain the fused feature F fused , and then calculate the reward R t Perform policy update and update the policy parameter θ according to the reward. Keep repeating the above steps until the predetermined number of training rounds or convergence conditions are reached.

[0137] After training, the finally selected feature fusion strategy will be applied to the actual fusion of multi-scale features. The output after feature fusion can be expressed as:

[0138] F fused =Action(s t )

[0139] That is, the best fusion strategy taken in state s t . During the running process, this process will continue, automatically adjust the feature fusion method according to environmental changes, improve the adaptability and tracking accuracy of the system, and continuously optimize the fusion strategy to improve the accuracy.

[0140] The optimized multi-scale feature fusion result F fused will continue to be passed as input to the Transformer encoder for self-attention mechanism calculation, and the output is the global feature vector for object tracking. Subsequently, the global feature vector for object tracking is used again for Kalman filtering to predict and track the target position trajectory. The steps are the same as those in step 2. The initial position x 0 and velocity v 0 are obtained through feature point matching and motion estimation:

[0141] x 0 =KalmanFilter(T)​​

[0142] Predict the position of the target at the next time step by combining the multi-modal features of the current time step through a Kalman filter:

[0143]

[0144] where, is the predicted target position, A is the state transition matrix, Bu k is the input control term

[0145] Finally, compare the actually observed target position with the predicted position, calculate the error, and update the state of the Kalman filter:

[0146] K k = P k|k-1 H T (HP k|k-1 H T + R) -1

[0147]

[0148] where, K k is the Kalman gain, z k is the observed position, P k|k-1 is the state covariance matrix.

[0149] 4. Online model update through meta-learning

[0150] The features fused in step 3 can also be used to calculate the meta-learning loss and adjust the model parameters in real time. By using the difference between the current tracking result and the true label, the model can continuously learn.

[0151] Adopt a meta-learning algorithm to quickly adjust the model parameter θ M , where the parameters include some parameters involved in the whole model, such as weights, etc.

[0152] Based on the real-time tracking and the updated data collected, calculate the meta-learning loss function L meta :

[0153] L meta = Loss(F fused , y true )

[0154] And use the gradient descent method to update the model parameters. The calculation method is as follows:

[0155]

[0156] where, θ M is the current model parameter, θM * To optimize the model parameters θ after updating M , α is the learning rate, is the gradient of the loss function with respect to the parameters.

[0157] On the premise of maintaining the model stability, new data can be continuously learned, and the model performance can be gradually optimized.

[0158] Comprehensive loss function:

[0159] For the multi-scale feature fusion updated based on reinforcement learning and meta-learning, the comprehensive loss function is defined as:

[0160]

[0161] θ, φ: represent the model parameters in reinforcement learning and meta-learning respectively;

[0162] E: represents the expectation, that is, the expected value of this loss function in this dynamic adjustment process;

[0163] α1, α2: weight coefficients, controlling the relative importance of the reward and penalty terms;

[0164] β1, β2: weight coefficients, controlling the relative importance of the time cost and resource usage;

[0165] R t : The reward term at each time step t of reinforcement learning, representing the task scheduling quality, mainly measuring the task completion, accuracy and speed.

[0166] Penalty t : The penalty term at each time step, reflecting the cost of default or error in task scheduling, such as delay and priority violation.

[0167] TimeCost t : The cost related to the task completion time, and the goal is to minimize the time cost through reinforcement learning.

[0168] Resource Usage t The resource consumption related to task scheduling, and the goal is to minimize the resource usage.

[0169] The final optimization objective function L total combines two parts of reinforcement learning and meta-learning, ensuring the efficiency and adaptive ability of task scheduling:

[0170] L total = L(θ, φ) + λ meta ·L meta (φ)

[0171] Among them, λ meta is a hyperparameter that adjusts the impact of meta-learning on overall optimization. The comprehensive optimization objective can simultaneously improve task efficiency, resource utilization, and system adaptability in multi-agent task scheduling.

[0172] 5. Spatiotemporal Self-Supervised Motion Compensation

[0173] By analyzing and processing the continuously updated model parameters and the optimized extracted features in the above steps, using the motion information between consecutive video frames, the motion law of the target is learned without additional labeled data. Calculate the self-supervised loss, the feature difference between the current frame and the previous frame:

[0174]

[0175] Among them, L self is the self-supervised loss function, f θ (x t ) is the feature extraction function of the current frame x t , t is the time step; then use the self-supervised loss L self to calculate the motion estimation, and the motion compensation of the target can be found by minimizing the loss function:

[0176]

[0177] Specifically, the gradient descent method can also be used to optimize the motion compensation amount:

[0178]

[0179] Among them, α is the learning rate, is the gradient with respect to the motion compensation amount.

[0180] According to the principle of self-supervised learning, the motion compensation amount can be directly calculated as the motion difference of the target between consecutive frames:

[0181] Δx c = f θ (x t ) - f θ (x t-1 )

[0182] When a drastic change in the target trajectory is detected, use the spatiotemporal features obtained by self-supervised learning to compensate for the motion to ensure the continuity and accuracy of tracking.

[0183] Finally, combine the self-supervised learning results to adjust the target position estimation:

[0184] x′ k = x k + Δx c

[0185] where Δx c is the motion compensation amount learned based on spatio-temporal information; x k is the position output by the Kalman filter, i.e., the target position, and x′ k is the target position after adding compensation (motion difference compensation by self-supervised learning). This output is used as the final output of the model, i.e., the finally given target position.

[0186] The adaptive fusion multi-modal UAV target tracking system of the present invention, as Figure 1 shown, includes a data acquisition and processing module, a multi-modal feature fusion module, a Transformer processing module based on reinforcement learning, a meta-learning online model update module, and a spatio-temporal self-supervised motion compensation module; the data acquisition and processing module is used to collect data information sent by sensors carried by the UAV and extract features; the multi-modal feature fusion module performs weighted fusion on the extracted features of the data acquisition and processing module; the Transformer processing module based on reinforcement learning performs fusion at different scales on the fusion features of the multi-modal feature fusion module, and the fused features are transmitted to the multi-modal feature fusion module to output a prediction result.

[0187] The meta-learning online model update module is used to optimize the feature fusion of the multi-modal feature fusion module and the Transformer processing module based on reinforcement learning; the spatio-temporal self-supervised motion compensation module analyzes and processes data such as the fusion features of the Kalman filter and the Transformer processing module based on reinforcement learning, calculates the self-supervised loss, and gives the motion compensation amount. The first extracted features are input into the Kalman filter for the first prediction, and then the features are input into the Transformer processing module based on reinforcement learning. The features obtained by this model are then re-filtered by the Kalman filter to output a prediction result, and the process is iteratively optimized to improve the recognition accuracy; at the same time, the parameters of the Transformer module based on reinforcement learning are also input into the meta-learning model for optimization and update; the self-supervised motion compensation module also continuously learns and obtains a motion compensation based on data errors in the process; the whole process is a dynamic optimization process, and finally the prediction result of the Kalman filter plus the result of the motion compensation is used as the final output of the model, i.e., the predicted and recognized target position.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present invention.

Claims

1. An adaptive fusion multi-modal UAV target tracking method, characterized by: The following steps are involved: Step 1: Collect data information through sensors carried by the drone, the data information includes visual information, infrared information, and lidar information of the sensor; Step 2: synchronize the visual information, infrared information, and lidar information of the sensor, and remove noise through filtering algorithms; Step 3, obtain the environmental feature vector E of the data information: Among them, E vision 、E infrared 、E lidar They are the feature vectors of visual information, infrared information, and lidar information respectively; Step 4: extract features from the environment feature vector E, and use the weighted average method to fuse multimodal features to obtain a comprehensive feature vector; Step 5: extract multiple features of different scales from the comprehensive feature vector, input the extracted multi-scale features into the Transformer encoder, perform multi-scale feature fusion, obtain fused features, and perform target tracking based on the obtained fused features.

2. The adaptive fusion multi-modal UAV target tracking method according to claim 1 is characterized in that: Step 3E vision 、E infrared 、E lidar Specifically: in They are the inputs from visual information, infrared information and lidar respectively.

3. The adaptive fusion multi-modal UAV target tracking method according to claim 2 is characterized in that: The specific steps of step 4 are: Step 4.1: Perform multimodal fusion on the extracted feature vectors by weighted average method to obtain the multimodal fusion feature vector X′ k : X′ k =W1E vision +W2E lidar +W3E infrared Among them, W1, W2, and W3 are the weights of the feature vectors of visual information, lidar information, and infrared information, respectively; Step 4.2, update the state of the Kalman filter, use the Kalman filter to filter the multimodal fusion feature vector to obtain a comprehensive feature vector.

4. The adaptive fusion multi-modal UAV target tracking method according to claim 3 is characterized in that: The specific steps of step 5 are: Step 5.1, from the comprehensive feature vector, select convolution kernels of different sizes, extract features of multiple different scales, and generate feature maps F_scaleN of different scales; F scaleN =Conv(F,K a×a ) Among them, K a×a is a convolution kernel of size a×a; N is the number of indexes, and F is the input; Step 5.2: The feature maps F of different scales are scaleN Concatenate in the channel dimension to form a tensor F containing multi-scale features multiscale ; F multiscale =concat(F scale1 ,F scale2 ……F scaleN ); Step 5.3, input the multi-scale features into the Transformer encoder, perform multi-scale feature fusion, and obtain the fused feature T: Among them, Q, K, and V are query vector, key vector, and value vector respectively; d k is a hyperparameter; Step 5.4: perform feature integration and use the fused feature T output by Transformer as the unified target feature for target tracking.

5. The adaptive fusion multi-modal UAV target tracking method according to claim 4 is characterized in that: In step 4.3, the state of the Kalman filter is updated as follows: Step 4.3.1, the Kalman filter KalmanFilter obtains the initial position x0 through feature point matching and motion estimation information: x0=KalmanFilter(x vision ,x infrared ,x lidar ) where x vision is the feature of visual information, x infrared is the characteristic of infrared information, x lidar Depth information features provided for LiDAR; Step 4.3.2, predict the position of the target at the next time step and get the predicted target position Among them, A is the state transfer matrix, Bu k For input control items; Step 4.3.2, compare the actual observed target position with the predicted position, calculate the error and update the state of the Kalman filter: K k =P k|k-1 H T (HP k|k-1 H T +R) -1 Among them, K k is the Kalman gain, z k is the observed position, P k|k-1 is the state covariance matrix; H is the initial known observation matrix; R is the initial known observation noise covariance matrix; The predicted state.

6. The adaptive fusion multi-modal UAV target tracking method according to claim 5 is characterized in that: In the step 5.3, when multi-scale features are fused, continuous iterative trial and error learning is achieved through the reinforcement learning function, the optimal feature fusion method is selected, and the fusion features are optimized.

7. The adaptive fusion multi-modal UAV target tracking method according to claim 6 is characterized in that: In the step 5.3, when multi-scale features are fused, self-supervised loss calculation is implemented through the self-supervised motion compensation function; the motion estimation is calculated using the self-supervised loss, and the motion compensation of the target is found by minimizing the loss function; combined with the self-supervised learning results, the target position estimate is adjusted to obtain the motion compensation amount.

8. An adaptive fusion multi-modal UAV target tracking system, used to implement the adaptive fusion multi-modal UAV target tracking method according to any one of claims 1 to 7, characterized in that: It includes a data acquisition and processing module, a multimodal feature fusion module, and a Transformer processing module based on reinforcement learning; the data acquisition and processing module is used to collect data information sent by sensors carried by unmanned aerial vehicles and perform feature extraction; the multimodal feature fusion module performs weighted fusion on the extracted features of the data acquisition and processing module; the Transformer processing module based on reinforcement learning performs fusion of fusion features of the multimodal feature fusion module at different scales, and the fused features are transmitted to the multimodal feature fusion module to output prediction results.

9. The adaptive fusion multi-modal UAV target tracking system according to claim 8, characterized in that: It also includes a meta-learning online model update module and a spatiotemporal self-supervised motion compensation module; The meta-learning online model updating module is used to optimize the feature fusion of the multimodal feature fusion module and the Transformer processing module based on reinforcement learning; the spatiotemporal self-supervised motion compensation module analyzes and processes the fusion features of the Transformer processing module based on reinforcement learning, calculates the self-supervised loss and gives the motion compensation amount.

Citation Information

Cited By

  • Unmanned aerial vehicle target detection method and device based on multi-modal detection information fusion

    CN120891493A