Aerial target tracking method based on multi-frame fusion
Through multi-frame fusion technology, combined with adaptive illumination correction, deep convolutional neural network and multimodal data fusion, the problem of feature extraction and motion prediction of multi-target tracking in complex environments is solved, efficient target tracking in dynamic environments is achieved, and the accuracy and continuity of tracking are improved.
Patent Information
- Application Number
- CN202510832826.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing multi-target tracking technologies have difficulty achieving high-quality image processing, accurate multimodal data alignment, and robust feature management in complex environments, resulting in degraded target tracking performance. In particular, targets are prone to loss or misjudgment under conditions such as lighting changes, weather interference, and rapid target movement.
A multi-frame fusion method is adopted to extract target features through adaptive illumination correction and deep convolutional neural network, combine with Kalman filtering to predict motion trajectory, and fuse multimodal sensor data in dynamic environment. The attention mechanism is used to enhance key features, and the long short-term memory network is used for time series analysis. When the target is lost, it is re-identified through the region proposal network, and the tracking parameters are dynamically optimized according to the complexity of the environment.
It improves the accuracy and robustness of multi-target tracking, can effectively cope with challenges such as complex lighting and occlusion, ensures the continuity and reliability of target tracking, and is suitable for fields such as intelligent monitoring and autonomous driving.
Smart Images

Figure CN120707595A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target tracking, and in particular relates to an aerial target tracking method based on multi-frame fusion. Background Art
[0002] Multi-target tracking technology is crucial in fields such as intelligent surveillance, autonomous driving, and drone navigation. Its core objective is to accurately identify and continuously track target objects in complex environments, providing reliable support for system decision-making. However, existing methods often perform poorly in complex scenarios, particularly under conditions such as changing lighting, meteorological interference, and rapid target motion, which can easily lead to target loss or misjudgment. These limitations stem from the lack of adaptability of existing solutions to dynamic environments and the low accuracy of multi-source data fusion, resulting in reduced tracking performance.
[0003] In this field, image processing and multimodal data alignment in dynamic environments have become key challenges. Complex lighting and meteorological conditions can significantly interfere with image quality, making it difficult for traditional image processing methods to accurately extract target features. As a result, the robustness of target features is limited, especially in fast-motion scenes, where inter-frame blur further weakens the reliability of feature extraction. The lack of feature discriminability directly affects the continuity of target tracking, making it difficult for the system to resume tracking after a target is briefly lost. How to achieve high-quality image processing, accurate multimodal data alignment, and robust feature management in a dynamic environment has become a key issue in improving multi-target tracking performance. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes an aerial target tracking method based on multi-frame fusion to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above object, the present invention provides an aerial target tracking method based on multi-frame fusion, comprising:
[0006] Acquiring ambient light intensity data, setting image correction model parameters based on the ambient light intensity data, and preprocessing the image based on the image correction model;
[0007] Obtaining a first feature set based on the preprocessed image and the pre-trained deep convolutional neural network model;
[0008] Predicting the motion of the target object based on the first feature set to obtain a first motion trajectory;
[0009] If the deviation between the first motion trajectory and the target position in the actual frame image exceeds a preset threshold, an auxiliary data set is obtained, and a second feature set is obtained based on the first feature set and the auxiliary data set; based on the second feature set, a feature enhancement algorithm based on the attention mechanism is used to obtain a third feature set;
[0010] Performing a time series analysis on the third feature set to obtain a continuous tracking state of the target, and if the target is lost, using a target re-identification algorithm based on a region proposal network to obtain a restored target tracking state;
[0011] The tracking optimization algorithm based on adaptive threshold is used to optimize the recovered target tracking state to obtain the final tracking result.
[0012] Optional pre-processing steps include:
[0013] Ambient light intensity data is obtained through a sensor, and dynamic configuration parameters of an image correction model are obtained based on the ambient light intensity data; based on the dynamic configuration parameters, an input image is processed to obtain a first image with adjusted brightness and contrast; if the brightness mean of the first image exceeds a preset threshold, the first image is adjusted twice using a histogram equalization algorithm to obtain a second image; based on the pixel distribution of the second image, image features are extracted using a preset convolutional neural network model to obtain a feature vector; the feature vector is matched with a preset standard feature library to determine whether the image meets the normalization standard. If not, the image correction model parameters are adjusted to reprocess the input image to obtain a normalized image.
[0014] Optionally, the process of obtaining the first feature set based on the preprocessed image and the pretrained deep convolutional neural network model includes:
[0015] A preset deep convolutional neural network model is used to extract the features of the target object from the normalized image, and the extracted features are subjected to dimensionality reduction processing through a preset pooling layer to obtain a compressed feature set; a principal component analysis algorithm is used to perform feature selection on the compressed feature set to obtain a key feature set; the feature variance of the key feature set is calculated and it is determined whether it is lower than a preset threshold. If so, the third feature set is optimized through a feature enhancement algorithm to obtain a first feature set; if not, the key feature set is used as the first feature set.
[0016] Optionally, the process of obtaining the first motion trajectory includes:
[0017] Based on the first feature set, a Kalman filter algorithm is used to predict the original motion trajectory of the target object, including the position and velocity of the next frame. The original motion trajectory is segmented through a preset time window to obtain a set of trajectory segments. The mean filter algorithm is used to smooth the set of trajectory segments to eliminate noise interference and obtain the first motion trajectory.
[0018] Optionally, the process of obtaining the second feature set includes:
[0019] The first feature set and the auxiliary data set are temporally and spatially aligned to obtain a fused data set; the fused data set is subjected to dimensionality reduction processing using the principal component analysis algorithm to extract the main eigenvector; if the variance of the eigenvector is lower than a preset threshold, the eigenvector is complemented by a linear interpolation method to obtain a second feature set.
[0020] Optionally, the third feature set is temporally modeled through a long short-term memory network, the historical frame feature sequence is analyzed, the current state of the target is predicted, and a continuous tracking state sequence is obtained.
[0021] Optionally, if the target is lost, a target re-identification algorithm based on a region proposal network is used. The process of obtaining the recovered target tracking state includes:
[0022] If target loss is detected in the tracking state sequence, a region proposal network is used to generate candidate target regions from the current frame to obtain the regional position of the candidate target. Based on the regional position of the candidate target, the feature vectors of each region are extracted to generate a set of regional feature vectors. Using a preset re-identification model, the similarity between the set of regional feature vectors and the historical target features is calculated to obtain the regional feature vector with the highest similarity and determine the re-identified target position. Based on the re-identified target position, the tracking state sequence is updated to obtain the restored target tracking state.
[0023] Optionally, if the matching degree between the restored target tracking state and the historical target features is lower than a preset threshold, the motion trajectory of the target in the current frame is estimated by the optical flow algorithm to obtain an estimated motion vector; according to the estimated motion vector, the re-identified target position is adjusted to generate a corrected target position; according to the corrected target position, the similarity between the regional feature vector and the historical target features is recalculated, and the tracking state sequence is updated.
[0024] Optionally, an adaptive threshold-based tracking optimization algorithm is used to optimize the recovered target tracking state. The process of obtaining the final tracking result includes:
[0025] Based on the state sequence data obtained from the recovered target tracking state, the state feature vector is generated by the feature extraction method to obtain the state feature set; according to the state feature set, the environmental complexity is analyzed, and the complexity coefficient is calculated by the preset complexity evaluation model to determine the environmental complexity level; if the environmental complexity level is higher than the preset threshold, the tracking parameters are adjusted by the adaptive threshold algorithm to generate an optimized parameter set; according to the optimized parameter set, the target tracking state sequence is updated, and the state smoothing method is adopted to obtain the smoothed tracking state; the motion features are extracted from the smoothed tracking state, and the continuity index is generated by the motion feature analysis to determine whether the continuity meets the preset conditions; if the continuity index is lower than the preset conditions, the motion features are optimized by the Kalman filter algorithm to generate a corrected tracking state; according to the corrected tracking state, the state feature vector is recalculated, the tracking state sequence is updated, and the final tracking result is obtained.
[0026] Optionally, it also includes: extracting the spatiotemporal features of the target from the final tracking results, using an association analysis algorithm based on a graph neural network to obtain the interactive relationship between multiple targets, and obtaining a multi-target tracking sequence including interactive information.
[0027] Compared with the prior art, the present invention has the following advantages and technical effects:
[0028] The present invention discloses an aerial target tracking method based on multi-frame fusion in a complex environment. The method extracts target features through adaptive illumination correction and deep convolutional neural network, and combines Kalman filtering to predict motion trajectory. When the prediction deviation is large, multimodal sensor data is fused to improve feature accuracy, and the attention mechanism is used to enhance key features. Long short-term memory network is used for time series analysis to achieve continuous tracking. When the target is lost, it is re-identified through the region proposal network, and the tracking parameters are dynamically optimized according to the complexity of the environment. Finally, the graph neural network is used to analyze the interaction relationship between multiple targets to obtain a tracking sequence containing interaction information. The present invention can effectively cope with challenges such as complex illumination and occlusion, improve the accuracy, robustness and continuity of multi-target tracking, and provide a reliable target tracking solution for intelligent monitoring, autonomous driving and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0030] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0031] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0032] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] Example 1
[0034] like Figure 1 As shown, this embodiment provides an aerial target tracking method based on multi-frame fusion, including:
[0035] Acquiring ambient light intensity data, setting image correction model parameters based on the ambient light intensity data, and preprocessing the image based on the image correction model;
[0036] As a specific implementation method, the preprocessing process includes: obtaining ambient light intensity data through a sensor, and obtaining dynamic configuration parameters of an image correction model based on the ambient light intensity data; processing the input image based on the dynamic configuration parameters to obtain a first image with adjusted brightness and contrast; if the brightness mean of the first image exceeds a preset threshold, performing a secondary adjustment on the first image through a histogram equalization algorithm to obtain a second image; based on the pixel distribution of the second image, using a preset convolutional neural network model to extract image features to obtain a feature vector; matching the feature vector with a preset standard feature library to determine whether the image meets the normalization standard. If not, adjusting the image correction model parameters to reprocess the input image to obtain a normalized image.
[0037] Specifically, the acquisition of ambient light intensity data is the core of adaptive image processing.
[0038] For example, a light sensor collects ambient light intensity in real time. For example, in an indoor setting, the light intensity is 500 lux. The sensor transmits this data to the processing unit, which then provides a basis for the correction model. This approach dynamically adapts to lighting changes, ensuring that the image processing results are accurate for the actual scene. The preset adaptive light correction model adjusts parameters based on light intensity.
[0039] For example, the model includes a built-in mapping table between light intensity and parameters. When the light intensity is 500 lux, the model automatically selects a brightness gain factor of 1.2 and a contrast adjustment coefficient of 1.1. These dynamically configured parameters are used to process the input image and generate the first image. This adjustment effectively improves image clarity in low-light environments.
[0040] The calculation of the brightness mean of the first image is a key step.
[0041] For example, the calculated mean brightness of the first image is 80, and the preset threshold is 100. Since the threshold is not exceeded, the first image is directly output as the second image. This mechanism avoids unnecessary secondary processing, saving computing resources while ensuring image quality. The histogram equalization algorithm is suitable for scenarios where brightness exceeds the standard. Suppose in another scenario, the mean brightness of the first image is 120, exceeding the threshold of 100. In this case, the algorithm redistributes the grayscale values of the image pixels, enhancing contrast and generating a second image. This processing can significantly improve the rendering of image details in high-brightness environments.
[0042] Based on the pixel distribution of the second image, a convolutional neural network model is preset to extract features.
[0043] For example, the network transforms the second image into a 256-dimensional feature vector through multiple layers of convolution and pooling. This vector contains information about the image's texture and structure, providing a reliable basis for subsequent matching. The efficiency of feature extraction directly impacts the accuracy of normalization determination. The feature vector is matched against a pre-set standard feature library to determine the normalization standard.
[0044] Exemplarily, the feature library contains feature vectors of multiple standard images. By calculating the Euclidean distance, it is determined whether the feature vector of the second image is less than 0.05 away from a vector in the standard library. If the conditions are met, it is determined to be a normalized image. This matching method ensures the consistency of the image processing results. If the normalization judgment does not meet the standards, the model parameters need to be adjusted and reprocessed. For example, the model may reduce the brightness gain factor from 1.2 to 1.1, regenerate the first image and repeat the subsequent steps. This iterative process can gradually optimize the image quality until the normalization standard is met.
[0045] Obtaining a first feature set based on the preprocessed image and the pre-trained deep convolutional neural network model;
[0046] As a specific implementation, the process of obtaining the first feature set based on the preprocessed image and the pretrained deep convolutional neural network model includes:
[0047] A preset deep convolutional neural network model is used to extract the features of the target object from the normalized image, and the extracted features are subjected to dimensionality reduction processing through a preset pooling layer to obtain a compressed feature set; a principal component analysis algorithm is used to perform feature selection on the compressed feature set to obtain a key feature set; the feature variance of the key feature set is calculated and it is determined whether it is lower than a preset threshold. If so, the third feature set is optimized through a feature enhancement algorithm to obtain a first feature set; if not, the key feature set is used as the first feature set.
[0048] Specifically, by performing dimensionality reduction processing on the first feature set through a preset pooling layer, computational complexity can be effectively reduced. The pooling layer uses a maximum pooling method to reduce the spatial resolution of the feature map.
[0049] For example, after 2x2 pooling of the 1024 features in the initial feature set, a compressed feature set of 512 features is generated. This dimensionality reduction preserves key information while reducing redundancy. Pooling also enhances the robustness of features to noise, ensuring stability in subsequent processing.
[0050] The principal component analysis algorithm is used to select features from the compressed feature set to further simplify the features. The principal component analysis extracts the principal components of the features through linear transformation and retains the part that contributes the most to the variance.
[0051] For example, after principal component analysis of the 512 features in the second feature set, a key feature set of 256 features was generated, retaining 90% of the variance information. This method effectively removes low-contribution features and improves the compactness of the feature set.
[0052] For example, if the variance of a key feature set falls below a preset threshold, it needs to be optimized using a feature enhancement algorithm. Assuming the threshold is 0.8 and the variance of the key feature set is 0.6, a feature enhancement algorithm, such as gradient boosting-based feature reconstruction, is used to generate the first feature set. The variance of the enhanced feature set increases to 0.85, meeting the requirement. If the variance exceeds 0.8, the key feature set is directly output as the first feature set. This dynamic adjustment mechanism ensures that the feature set always maintains high information content.
[0053] If the final judgment result does not meet the standards, the deep convolutional neural network model parameters need to be adjusted.
[0054] For example, the convolution kernel size is adjusted from 3x3 to 5x5, features are re-extracted, and the process is repeated. This iterative optimization ensures accurate feature extraction. The entire process achieves efficient object recognition through multi-level feature processing and dynamic adjustment.
[0055] Predicting the motion of the target object based on the first feature set to obtain a first motion trajectory;
[0056] As a specific implementation, the process of obtaining the first motion trajectory includes:
[0057] Based on the first feature set, a Kalman filter algorithm is used to predict the original motion trajectory of the target object, including the position and velocity of the next frame. The original motion trajectory is segmented through a preset time window to obtain a set of trajectory segments. The mean filter algorithm is used to smooth the set of trajectory segments to eliminate noise interference and obtain the first motion trajectory.
[0058] For example, in an image-based target tracking scenario, a Kalman filter algorithm is used to process a first feature set extracted from the target object to predict its trajectory. The Kalman filter establishes a state transition model and an observation model, combining prior estimates with current observations to predict the target object's position and velocity in the next frame.
[0059] The Kalman filter reduces the prediction error and ensures the stability of trajectory prediction by iteratively updating the state covariance matrix.
[0060] The original motion trajectory is segmented using a preset time window to form a set of trajectory segments. The time window can be set to 5 frames, dividing the continuous trajectory into multiple segments, each containing position and velocity information.
[0061] For example, a car's trajectory segment set within 5 frames may include the path from the starting point (90, 140) to the end point (110, 160). This segmentation facilitates subsequent analysis of the local characteristics of the trajectory.
[0062] The mean filter algorithm is used to smooth segmented motion trajectories and eliminate noise. For example, if the coordinates of a segmented trajectory point are jittery, such as (105, 155), (106, 158), and (104, 153), the mean filter takes the average of these three points and smooths the point (105, 155.3), forming the third trajectory. This smoothing process reduces the effects of sensor noise or environmental interference.
[0063] If the deviation between the first motion trajectory and the target position in the actual frame image exceeds a preset threshold, an auxiliary data set is obtained, and a second feature set is obtained based on the first feature set and the auxiliary data set; based on the second feature set, a feature enhancement algorithm based on the attention mechanism is used to obtain a third feature set;
[0064] As a specific implementation, the process of obtaining the second feature set includes:
[0065] The first feature set and the auxiliary data set are temporally and spatially aligned to obtain a fused data set; the fused data set is subjected to dimensionality reduction processing using the principal component analysis algorithm to extract the main eigenvector; if the variance of the eigenvector is lower than a preset threshold, the eigenvector is complemented by a linear interpolation method to obtain a second feature set.
[0066] Specifically, when the deviation between the first motion trajectory and the target position in the actual frame image exceeds a preset threshold (e.g., 5 meters), the fusion module first performs spatiotemporal alignment on the visual data (e.g., the vehicle body outline captured by the camera), radar data (e.g., target distance and speed), and infrared data (e.g., nighttime thermal imaging). In one possible implementation, the visual data is sampled at 30 frames per second, the radar data at 10 times per second, and the infrared data at 5 times per second. Through timestamp alignment and coordinate system conversion, the three are ensured to be in the same spatiotemporal reference system to generate a second feature set. This process effectively integrates multi-source information and improves feature integrity.
[0067] Specifically, the principal component analysis algorithm is used to perform dimensionality reduction on the fused data set.
[0068] For example, if the fused dataset contains high-dimensional features, principal component analysis can be used to extract the first three principal components while retaining 90% of the variance. This dimensionality reduction process can reduce computational complexity while preserving key information. If the variance of the feature vector falls below a preset threshold (such as 0.1), the missing features are supplemented through linear interpolation.
[0069] For example, if the speed feature of a frame is missing due to occlusion, it can be supplemented based on the linear estimation of the speed of the previous and next frames. If the variance is higher than the threshold, the fusion data set is directly subjected to dimensionality reduction processing, and the extracted main feature vector is used as the second feature set to ensure the representativeness of the feature.
[0070] Exemplarily, a preset convolutional neural network model is used to extract the initial feature vector from the second feature set to obtain a first feature vector set. The first feature vector set is grouped by a clustering algorithm to generate a feature cluster set and obtain the feature distribution within the cluster. If the compactness of the feature distribution within the cluster is lower than a preset threshold, the feature vector is complemented by a linear interpolation method to obtain a second feature vector set; if the compactness is higher than the preset threshold, the first feature vector set is directly output as the second feature vector set. Based on the second feature vector set, the similarity between features is calculated using a preset kernel function method to obtain a feature similarity matrix. Using the feature similarity matrix, the feature vectors are integrated using a weighted fusion method to obtain a third feature vector set. If the dimension of the third feature vector set is higher than the preset threshold, the principal component analysis algorithm is used to perform dimensionality reduction processing to obtain a third feature set; if the dimension is lower than the preset threshold, the third feature vector set is directly output as the third feature set.
[0071] Performing a time series analysis on the third feature set to obtain a continuous tracking state of the target, and if the target is lost, using a target re-identification algorithm based on a region proposal network to obtain a restored target tracking state;
[0072] As a specific implementation method, the third feature set is temporally modeled through a long short-term memory network, the historical frame feature sequence is analyzed, the current state of the target is predicted, and a continuous tracking state sequence is obtained.
[0073] As a specific implementation method, if the target is lost, a target re-identification algorithm based on a region proposal network is used. The process of obtaining the recovered target tracking state includes:
[0074] If target loss is detected in the tracking state sequence, a region proposal network is used to generate candidate target regions from the current frame to obtain the regional position of the candidate target. Based on the regional position of the candidate target, the feature vectors of each region are extracted to generate a set of regional feature vectors. Using a preset re-identification model, the similarity between the set of regional feature vectors and the historical target features is calculated to obtain the regional feature vector with the highest similarity and determine the re-identified target position. Based on the re-identified target position, the tracking state sequence is updated to obtain the restored target tracking state.
[0075] Specifically, in time-series tracking scenarios based on LSTM networks, target loss is a common problem, requiring a region proposal network to generate candidate regions to resume tracking. The region proposal network scans the current frame, identifies regions that may contain the target, and generates multiple candidate boxes.
[0076] For example, assuming the current frame is a frame from a surveillance video, the region proposal network can generate 50 candidate boxes, each containing position coordinates and size information, such as the upper left corner coordinates of (100, 150) and the width and height of (80, 120). These candidate boxes cover the potential target area, ensuring that no target is missed.
[0077] Specifically, when extracting feature vectors from candidate boxes, a pre-trained convolutional neural network can be used to extract the visual features of each region and generate a set of regional feature vectors.
[0078] For example, each candidate box generates a 256-dimensional feature vector, and the set contains 50 vectors. The re-ID model then calculates the similarity of these vectors with historical object features. Historical object features can come from known objects in previous frames, such as the color and texture of a red car.
[0079] The similarity calculation may be based on the cosine distance, and the region with the highest similarity is selected as the re-identified target location.
[0080] For example, if the feature vector of a candidate box has a similarity of 0.95 with historical features, much higher than the 0.7 of other regions, it is determined to be the target location. If the match degree of the re-identified target location falls below a threshold (such as 0.8), the target's motion trajectory is estimated using an optical flow algorithm. The optical flow algorithm analyzes the pixel motion between the current frame and the previous frame to generate a motion vector.
[0081] For example, a target might move from (200, 300) in the previous frame to (220, 310) in the current frame, with a motion vector of (20, 10). Based on this vector, the re-identified target position is adjusted to generate a corrected position, such as (218, 308). The similarity between the feature vector of the corrected position and the historical feature vector is then recalculated. If the similarity improves to 0.85, the tracking state sequence is updated. The updated tracking state sequence more accurately reflects the target's continuous motion trajectory.
[0082] Preferably, the re-ID model can incorporate multimodal features, such as the shape and color of the target, to further improve the reliability of the similarity calculation. These measures together ensure stable updates of the tracking state sequence and adapt to the needs of dynamic scenes.
[0083] As a specific implementation method, if the matching degree between the restored target tracking state and the historical target features is lower than a preset threshold, the motion trajectory of the target in the current frame is estimated by the optical flow algorithm to obtain an estimated motion vector; according to the estimated motion vector, the re-identified target position is adjusted to generate a corrected target position; according to the corrected target position, the similarity between the regional feature vector and the historical target features is recalculated, and the tracking state sequence is updated.
[0084] Specifically, obtaining state sequence data from the recovered target tracking state involves extracting information such as the target's position, speed, and direction from consecutive frames.
[0085] When analyzing environmental complexity, a complexity coefficient is calculated using a pre-set complexity assessment model. This model may be based on the density and motion patterns of objects in the background. For example, in an urban road scene, pedestrians, vehicles, and traffic lights increase background complexity. The model outputs a complexity coefficient of 0.8, which is higher than the pre-set threshold of 0.6, indicating a high-complexity environment. High-complexity environments can cause tracking drift, so an adaptive threshold algorithm is used to adjust tracking parameters.
[0086] Exemplarily, the algorithm reduces the confidence threshold of the detection box from 0.7 to 0.5, increases the screening range of candidate targets, generates an optimized parameter set, and updates the state sequence.
[0087] In a possible implementation, the state smoothing process may be implemented by a weighted average method.
[0088] For example, the coordinate data of the previous three frames is combined to calculate the smoothed position of the current frame and reduce jitter. Motion features are extracted from the smoothed state to generate a continuity index. Assume the continuity index is 0.9, which is above the threshold of 0.85, indicating stable tracking. If the index is below the threshold, for example, 0.7, the motion features are optimized using a Kalman filter. The Kalman filter uses a prediction and update step to estimate the true position and generate a corrected state.
[0089] The tracking optimization algorithm based on adaptive threshold is used to optimize the recovered target tracking state to obtain the final tracking result.
[0090] As a specific implementation method, the process of optimizing the recovered target tracking state using a tracking optimization algorithm based on an adaptive threshold to obtain the final tracking result includes:
[0091] Based on the state sequence data obtained from the recovered target tracking state, the state feature vector is generated by the feature extraction method to obtain the state feature set; according to the state feature set, the environmental complexity is analyzed, and the complexity coefficient is calculated by the preset complexity evaluation model to determine the environmental complexity level; if the environmental complexity level is higher than the preset threshold, the tracking parameters are adjusted by the adaptive threshold algorithm to generate an optimized parameter set; according to the optimized parameter set, the target tracking state sequence is updated, and the state smoothing method is adopted to obtain the smoothed tracking state; the motion features are extracted from the smoothed tracking state, and the continuity index is generated by the motion feature analysis to determine whether the continuity meets the preset conditions; if the continuity index is lower than the preset conditions, the motion features are optimized by the Kalman filter algorithm to generate a corrected tracking state; according to the corrected tracking state, the state feature vector is recalculated, the tracking state sequence is updated, and the final tracking result is obtained.
[0092] The spatiotemporal features of the target are extracted from the final tracking results, and the association analysis algorithm based on graph neural network is used to obtain the interactive relationship between multiple targets, and a multi-target tracking sequence including interactive information is obtained.
[0093] For example, in the field of target tracking, obtaining spatiotemporal feature sequences from the final tracking results is the basis for building multi-target interaction models. The spatiotemporal feature sequences usually contain information such as the target's position, velocity, and timestamp.
[0094] Normalization during data preprocessing can eliminate the effects of different dimensions. Specifically, location coordinates might be in meters, and speed in meters per second. Normalization maps these features to the same range, such as a distribution with a mean of 0 and a variance of 1, to facilitate subsequent algorithm processing.
[0095] In one possible implementation, a graph neural network algorithm is used to model the interactions between multiple targets. The graph neural network treats targets as nodes and interactions as edges to generate an interaction matrix.
[0096] For example, in a pedestrian tracking scenario, matrix elements might represent the distance or speed similarity between pedestrians. If the distance between two pedestrians is less than 2 meters, the matrix element value is high, indicating strong interaction. It should be noted that if some matrix element values fall below a preset threshold, such as 0.3, this indicates weak interaction between the targets and may require grouping.
[0097] Specifically, clustering algorithms can be used to group objects.
[0098] For example, based on the K-means algorithm, pedestrians are divided into different groups based on spatial proximity, such as those near the entrance and those near the exit. After grouping, the sequence recombination method rearranges the spatiotemporal feature sequence.
[0099] For example, prioritizing feature sequences of pedestrians in the same group can reduce data fragmentation during subsequent processing. Interaction feature extraction focuses on the dynamic relationships between targets within a group, such as the relative speed differences between pedestrians. Feature fusion methods can generate multi-target interaction feature vectors through weighted averaging.
[0100] For example, the position and velocity features are fused with weights of 0.6 and 0.4 to form a comprehensive vector.
[0101] In one embodiment, a sequence generation algorithm updates a multi-target tracking sequence.
[0102] For example, based on the long short-term memory network and combined with the interactive feature vector, a tracking sequence containing interactive information is generated. If there is missing data in the sequence, the missing position can be filled by the linear interpolation algorithm.
[0103] For example, the intermediate position during occlusion can be estimated based on the coordinates of the previous and next frames. This method can generate a complete multi-target tracking sequence and ensure the continuity of tracking.
[0104] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for tracking aerial targets based on multi-frame fusion, characterized in that: The following steps are involved: Acquiring ambient light intensity data, setting image correction model parameters based on the ambient light intensity data, and preprocessing the image based on the image correction model; Obtaining a first feature set based on the preprocessed image and the pre-trained deep convolutional neural network model; Predicting the motion of the target object based on the first feature set to obtain a first motion trajectory; If the deviation between the first motion trajectory and the target position in the actual frame image exceeds a preset threshold, obtaining an auxiliary data set, and obtaining a second feature set based on the first feature set and the auxiliary data set; Based on the second feature set, a feature enhancement algorithm based on the attention mechanism is used to obtain the third feature set; Performing a time series analysis on the third feature set to obtain a continuous tracking state of the target, and if the target is lost, using a target re-identification algorithm based on a region proposal network to obtain a restored target tracking state; The tracking optimization algorithm based on adaptive threshold is used to optimize the recovered target tracking state to obtain the final tracking result.
2. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: The preprocessing process includes: Ambient light intensity data is obtained through a sensor, and dynamic configuration parameters of an image correction model are obtained based on the ambient light intensity data; based on the dynamic configuration parameters, an input image is processed to obtain a first image with adjusted brightness and contrast; if the brightness mean of the first image exceeds a preset threshold, the first image is adjusted twice using a histogram equalization algorithm to obtain a second image; based on the pixel distribution of the second image, image features are extracted using a preset convolutional neural network model to obtain a feature vector; the feature vector is matched with a preset standard feature library to determine whether the image meets the normalization standard. If not, the image correction model parameters are adjusted to reprocess the input image to obtain a normalized image.
3. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: The process of obtaining the first feature set based on the preprocessed image and the pre-trained deep convolutional neural network model includes: A preset deep convolutional neural network model is used to extract the features of the target object from the normalized image, and the extracted features are subjected to dimensionality reduction processing through a preset pooling layer to obtain a compressed feature set; a principal component analysis algorithm is used to perform feature selection on the compressed feature set to obtain a key feature set; the feature variance of the key feature set is calculated and it is determined whether it is lower than a preset threshold. If so, the third feature set is optimized through a feature enhancement algorithm to obtain a first feature set; if not, the key feature set is used as the first feature set.
4. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: The process of obtaining the first motion trajectory includes: Based on the first feature set, a Kalman filter algorithm is used to predict the original motion trajectory of the target object, including the position and velocity of the next frame. The original motion trajectory is segmented through a preset time window to obtain a set of trajectory segments. The mean filter algorithm is used to smooth the set of trajectory segments to eliminate noise interference and obtain the first motion trajectory.
5. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: The process of obtaining the second feature set includes: The first feature set and the auxiliary data set are temporally and spatially aligned to obtain a fused data set; the fused data set is subjected to dimensionality reduction processing using the principal component analysis algorithm to extract the main eigenvector; if the variance of the eigenvector is lower than a preset threshold, the eigenvector is complemented by a linear interpolation method to obtain a second feature set.
6. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: The third feature set is temporally modeled through a long short-term memory network, the historical frame feature sequence is analyzed, the current state of the target is predicted, and a continuous tracking state sequence is obtained.
7. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: If the target is lost, the target re-identification algorithm based on the region proposal network is used. The process of obtaining the recovered target tracking state includes: If target loss is detected in the tracking state sequence, a region proposal network is used to generate candidate target regions from the current frame to obtain the regional position of the candidate target. Based on the regional position of the candidate target, the feature vectors of each region are extracted to generate a set of regional feature vectors. Using a preset re-identification model, the similarity between the set of regional feature vectors and the historical target features is calculated to obtain the regional feature vector with the highest similarity and determine the re-identified target position. Based on the re-identified target position, the tracking state sequence is updated to obtain the restored target tracking state.
8. The aerial target tracking method based on multi-frame fusion according to claim 7, characterized in that: If the matching degree between the restored target tracking state and the historical target features is lower than the preset threshold, the motion trajectory of the target in the current frame is estimated by the optical flow algorithm to obtain the estimated motion vector; according to the estimated motion vector, the re-identified target position is adjusted to generate a corrected target position; according to the corrected target position, the similarity between the regional feature vector and the historical target features is recalculated, and the tracking state sequence is updated.
9. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: The tracking optimization algorithm based on adaptive threshold is used to optimize the recovered target tracking state. The process of obtaining the final tracking result includes: Based on the state sequence data obtained from the recovered target tracking state, the state feature vector is generated by the feature extraction method to obtain the state feature set; according to the state feature set, the environmental complexity is analyzed, and the complexity coefficient is calculated by the preset complexity evaluation model to determine the environmental complexity level; if the environmental complexity level is higher than the preset threshold, the tracking parameters are adjusted by the adaptive threshold algorithm to generate an optimized parameter set; according to the optimized parameter set, the target tracking state sequence is updated, and the state smoothing method is adopted to obtain the smoothed tracking state; the motion features are extracted from the smoothed tracking state, and the continuity index is generated by the motion feature analysis to determine whether the continuity meets the preset conditions; if the continuity index is lower than the preset conditions, the motion features are optimized by the Kalman filter algorithm to generate a corrected tracking state; according to the corrected tracking state, the state feature vector is recalculated, the tracking state sequence is updated, and the final tracking result is obtained.
10. The aerial target tracking method based on multi-frame fusion according to claim 1, characterized in that: It also includes: extracting the spatiotemporal features of the target from the final tracking results, using an association analysis algorithm based on a graph neural network to obtain the interactive relationship between multiple targets, and obtaining a multi-target tracking sequence including interactive information.
Citation Information
Cited By
Multi-target tracking method based on trajectory pipeline regression prediction and association
CN120953322A
Accurate identification method and system based on air dynamic target
CN121010895A