Video image target tracking method and processing device based on multi-feature fusion
Through the multi-feature fusion video image target tracking method, Kalman filtering and particle filtering algorithm combined with lightweight convolutional neural networks, the problem of tracking error accumulation in resource-confined scenarios in video image target tracking is solved, and the accuracy and real-time tracking is improved.
Patent Information
- Application Number
- CN202510612747.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, video image target tracking is accumulated due to premature switching of view angles in scenarios where resource limitations are found, resulting in poor tracking accuracy and real-time.
Using a multi-feature fusion method, combined with Kalman filtering and particle filtering algorithm, the target's color histogram, LBP texture features and optical flow motion features are extracted through optical flow method and lightweight convolutional neural network, a multi-feature description model is constructed, feature weights are dynamically adjusted, candidate targets are generated, feature similarity is calculated, and computational overhead is optimized.
It improves the accuracy and real-timeness of target tracking, effectively solves the problem of tracking error accumulation in resource-constrained scenarios, and ensures the continuity of target tracking in complex scenarios.
Smart Images

Figure CN120471955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image target tracking technology, and in particular to a video image target tracking method and processing device based on multi-feature fusion. Background Art
[0002] With the continuous development of computer vision technology, video image object tracking has been widely used in many fields, such as intelligent surveillance, autonomous driving, and robotic navigation. However, achieving efficient and reliable object tracking faces many challenges. Objects in video scenes may move at high speeds and be easily obscured by other objects. To accurately track a target, the system must be able to quickly locate and maintain a continuous lock, while also dynamically adjusting the tracking strategy based on information about surrounding obstacles to ensure continuous tracking.
[0003] In one existing technology, a single feature is mainly relied upon to accurately locate and continuously lock onto the target. During the movement of the target, it may be temporarily blocked by other objects. At the moment of the occlusion, the target's movement trend is predicted, and then the tracking strategy is dynamically planned based on the obstacle information to adjust the observation angle in advance to track the target.
[0004] In the existing technology, tracking errors accumulate due to premature switching of perspectives. In resource-constrained scenarios, the computational overhead of deep learning models is usually large, resulting in poor tracking accuracy and real-time performance. Summary of the Invention
[0005] The present invention provides a video image target tracking method and processing device based on multi-feature fusion to solve the problem of accumulation of tracking errors caused by premature switching of perspectives. In resource-constrained scenarios, the computational overhead of deep learning models is usually large, resulting in poor tracking accuracy and real-time performance.
[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a video image target tracking method based on multi-feature fusion, comprising: Get the target's initial position parameters and current frame detection parameters; Calculate according to the initial position parameters to obtain a pixel-level motion vector field; Predicting the target position based on the pixel-level motion vector field; Comparing the predicted position parameter with the current frame detection parameter, and determining that the target is not blocked when the difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold; When the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, it is determined that the target is blocked, and an occlusion processing mechanism is started to track the target.
[0007] In an optional implementation, the calculating according to the initial position parameters to obtain the pixel-level motion vector field includes: Calculating the motion vector of the target between adjacent frames using an optical flow method according to the initial position parameters to obtain a first motion vector field of the target; Inputting the first motion vector field into a Kalman filter for calculation to obtain motion state parameters of the target; wherein the Kalman filter is a recursive estimation algorithm; Recursive estimation is performed based on the motion state parameters to obtain a pixel-level motion vector field.
[0008] In an optional implementation, performing prediction based on the pixel-level motion vector field to obtain predicted position parameters of the target includes: Calculating according to the pixel-level motion vector field to obtain the current frame position coordinates of the target; Determine the target's search center parameters based on the current frame position coordinates; Searching based on the search center parameter to obtain a target search range parameter; Prediction is performed based on the search range parameters to obtain predicted position parameters of the target.
[0009] In an optional embodiment, when the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, determining that the target is occluded, and starting an occlusion processing mechanism to track the target include: Obtain the target's historical trajectory data, color histogram, LBP texture features, and optical flow motion features; Estimating the target's motion trend data based on the historical trajectory data; Probabilistically estimating the target position using a particle filter algorithm based on the motion trend data to obtain parameters of the area where the target may appear; Using a lightweight convolutional neural network to jointly compress and optimize the color histogram, the LBP texture feature, and the optical flow motion feature to obtain multi-feature compressed data; A sliding window is used to generate all candidate targets in the area according to the area parameters and multi-feature fusion data; Calculating the distance between the candidate target and the template target based on the candidate target to obtain the Euclidean distance of each candidate target; Calculating the feature similarity between the candidate target and the template target based on the candidate target to obtain the cosine similarity of each candidate target; Calculate the similarity score of each candidate target based on the Euclidean distance and the cosine similarity; Sorting is performed based on the similarity scores. If the matching degree between the candidate target with the highest similarity score and the historical target feature does not reach a preset threshold, the observation angle is adjusted based on the target motion trend estimated by the optical flow method, and the target compression feature is re-extracted and matched; If the matching degree between the candidate target with the highest similarity score and the historical target feature exceeds a preset threshold, the candidate target is determined to be the tracking target of the current frame, and the target position information is updated to track the target.
[0010] In an optional embodiment, the probability estimation of the target position using a particle filter algorithm based on the motion trend data to obtain parameters of an area where the target may appear includes: A particle filter algorithm is used to predict the state of the target based on the motion trend data to obtain a set of particle samples; wherein the particle sample represents a possible location of the target; Calculating based on the particle sample to obtain the observation likelihood probability of each particle; The particle weights are updated according to the observation likelihood probability to obtain particle weight distribution data; Resampling is performed according to the weight distribution data to obtain parameters of the area where the target may appear.
[0011] In an optional embodiment, the color histogram, the LBP texture feature, and the optical flow motion feature are jointly compressed and optimized using a lightweight convolutional neural network to obtain multi-feature compressed data, including: Simplifying the color histogram, the LBP texture feature, and the optical flow motion feature using low-rank decomposition and pruning technology to obtain feature coding; Designed based on the feature encoding, a lightweight spatial attention layer is obtained; Writing according to the spatial attention layer to obtain a lightweight temporal memory unit; A lightweight convolutional neural network is used for training according to the temporal memory unit to obtain multi-feature compressed data.
[0012] In an optional embodiment, calculating the feature similarity between the candidate target and the template target based on the candidate target to obtain the cosine similarity of each candidate target includes: The cosine similarity is calculated as follows: = in, is the cosine similarity, is the vector dot product.
[0013] In an optional embodiment, calculating the distance between the candidate target and the template target based on the candidate target to obtain the Euclidean distance of each candidate target includes: The Euclidean distance is calculated as follows:
[0014] in, is the Euclidean distance, is the vector dot product.
[0015] In an optional implementation, the calculation based on the Euclidean distance and the cosine similarity to obtain a similarity score for each candidate target includes: The similarity score is calculated as follows: in, is the similarity score, is a weight parameter, is the cosine similarity, is the Euclidean distance, is the distance threshold.
[0016] In a second aspect, the present invention provides a video image target tracking device based on multi-feature fusion, comprising: The data acquisition module is used to obtain the initial position parameters of the target and the current frame detection parameters; A data calculation module for calculating according to the initial position parameters to obtain a pixel-level motion vector field; A data prediction module for performing prediction based on the pixel-level motion vector field to obtain predicted position parameters of the target; a data comparison module, configured to compare the predicted position parameter with the current frame detection parameter, and determine that the target is not blocked when the difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold; The data judgment module is used to judge that the target is blocked when the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, and to start the blockage processing mechanism to track the target.
[0017] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements any one of the above-mentioned video image target tracking methods based on multi-feature fusion.
[0018] In a fourth aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned video image target tracking methods based on multi-feature fusion.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a video image target tracking method based on multi-feature fusion. The method comprises obtaining initial position parameters and current frame detection parameters of a target; performing calculations based on the initial position parameters to obtain a pixel-level motion vector field; performing predictions based on the pixel-level motion vector field to obtain predicted position parameters of the target; and comparing the predicted position parameters with the current frame detection parameters. When the difference between the predicted position parameters and the current frame detection parameters is less than a preset threshold, the target is determined to be unobstructed; and when the difference between the predicted position parameters and the current frame detection parameters is greater than a preset threshold, the target is determined to be obstructed, and an occlusion processing mechanism is activated to track the target. The method constructs a multi-feature description model by extracting the target's color histogram, LBP texture features, and optical flow motion features. A lightweight convolutional neural network is used to jointly compress and optimize the features, reducing computational overhead. The method combines Kalman filtering and particle filtering algorithms to predict the target's position, generates candidate targets within a possible area, and calculates feature similarity. The weights of each feature in the multi-feature fusion model are dynamically adjusted, prioritizing the most robust features, effectively improving tracking accuracy and real-time performance. The present invention can solve the problem in the prior art of accumulation of tracking errors caused by premature switching of perspectives. In resource-constrained scenarios, the computational overhead of deep learning models is usually large, resulting in poor tracking accuracy and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 1 is a flow chart of a method for video image target tracking based on multi-feature fusion provided by the first embodiment of the present invention; Figure 2 Schematic diagram of the system for video image target tracking provided by the present invention; Figure 3 It is a structural diagram of a video image target tracking device based on multi-feature fusion provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] Reference Figure 1 The first embodiment of the present invention provides a video image target tracking method based on multi-feature fusion, comprising the following steps: S11, obtain the initial position parameters of the target and the current frame detection parameters; S12, calculates according to the initial position parameters to obtain a pixel-level motion vector field; S13, predicting the target position parameters based on the pixel-level motion vector field; S14, comparing the predicted position parameter with the current frame detection parameter, and when a difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold, determining that the target is not blocked; S15: When the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, it is determined that the target is blocked, and an occlusion processing mechanism is started to track the target.
[0023] In step S11, the initial position parameters of the target and the current frame detection parameters are obtained.
[0024] The initial position parameters of an object refer to its position in the video image at the start of the tracking task. These parameters include coordinate information, size information, and initial pose information. Position information refers to the coordinates (x, y) of the center point of the object in the image, or the coordinates of the upper left and lower right corners of the object's bounding box. The bounding box is the minimum enclosing rectangle of the object in the image and is used to accurately locate the object. The width and height h of the object are calculated from the coordinates of the bounding box. The initial position parameters are obtained by the object detection algorithm in the first frame of the video. Object detection algorithms (such as deep learning-based object detection models such as YOLO, SSD, and Faster R-CNN) can identify the object's position and category and output the aforementioned parameters. The current frame detection parameters refer to the output of the object detection algorithm in the currently processed video frame. These parameters are similar to the initial position parameters, but reflect the state of the object in the current frame. The current frame detection parameters include the current frame's coordinate information, current frame size information, and current frame pose information. The current frame's coordinate information refers to the coordinates of the target's center point or bounding box in the current frame, and the current frame's size information refers to the target's width and height in the current frame, calculated in the same way as the initial size information. The current frame's pose information means that if the target's pose changes within the current frame, the target's pose information, such as the rotation angle, must be recorded. In the first frame of the video, the target's initial position parameters are obtained using the object detection algorithm. These parameters serve as the initial input to the tracking algorithm. In each subsequent frame, the object detection algorithm is rerun to obtain the target's position parameters for the current frame. These parameters are used to update the tracking algorithm's state to ensure tracking accuracy.
[0025] In step S12, calculation is performed based on the initial position parameters to obtain a pixel-level motion vector field.
[0026] In one implementation, the motion vector of the target between adjacent frames is calculated using the optical flow method based on the initial position parameters to obtain a first motion vector field of the target; the first motion vector field is input into a Kalman filter for calculation to obtain the motion state parameters of the target; wherein the Kalman filter is a recursive estimation algorithm; recursive estimation is performed based on the motion state parameters to obtain a pixel-level motion vector field.
[0027] Specifically, optical flow is a commonly used motion estimation technique that infers target motion by calculating changes in pixel intensity across an image sequence. When tracking fast-moving targets, optical flow can effectively capture the displacement of a target between consecutive frames. For example, for a high-speed vehicle, the motion vector of each pixel can be calculated across consecutive video frames, resulting in a dense vector field. This vector field reflects the vehicle's overall motion trend as well as local changes in details. The Kalman filter is a recursive estimation algorithm that effectively handles noisy measurement data. Inputting the optical flow vector field into a Kalman filter yields a smoother and more accurate motion estimate. For example, for a leaf blown by wind, the original optical flow vector may contain some random noise, but after Kalman filtering, a more stable trajectory is obtained. Calculating the target's center of mass coordinates is a key step in the tracking process. By taking a weighted average of the filtered vector field, the target's precise position in the current frame can be determined. For example, for a moving pedestrian, even if the pedestrian's limbs move in different directions during movement, the center of the pedestrian's body can still be determined by calculating the average of the overall motion vectors. Predicting the target's position in the next frame is crucial for continuous tracking. A Kalman filter can predict the target's future position based on the current motion state and historical data. For example, for a uniformly moving ball, the current position and velocity information can be used to predict the range of possible positions for the ball in the next frame. Performing a local search around the predicted position is an effective tracking strategy. This approach reduces computational effort and improves tracking efficiency. For example, for a fast-flying bird, a search window can be set near the predicted position, limiting target detection to this limited area rather than performing a global search across the entire image. Analyzing the motion vector field within the search area can further improve tracking accuracy. By comparing the motion characteristics of different parts of the search area, it is possible to distinguish the target from the background. For example, when tracking a bee moving against a complex background, the motion patterns within the search area can be analyzed to identify the bee's unique flight trajectory, distinguishing it from the stationary or slow-moving background. Finally, by continuously updating the target state in the Kalman filter, continuous tracking of the target is achieved. This process is similar to the attention mechanism in the human visual system, which enables sustained focus on a target of interest in complex visual scenes. For example, in a soccer game, it is possible to keep tracking a specific player by continuously updating the target state even if the players move quickly and are frequently obscured by other players.
[0028] In step S13, prediction is performed based on the pixel-level motion vector field to obtain predicted position parameters of the target.
[0029] In one implementation, calculation is performed based on the pixel-level motion vector field to obtain the current frame position coordinates of the target; judgment is performed based on the current frame position coordinates to obtain the search center parameters of the target; search is performed with the search center parameters as the center to obtain the search range parameters of the target; prediction is performed based on the search range parameters to obtain the predicted position parameters of the target.
[0030] The goal is to predict the target's position parameters based on the pixel-level motion vector field. Optical flow methods calculate the motion vector for each pixel by analyzing the change in pixel intensity between consecutive frames. Common methods include region-based methods (such as the Lucas-Kanade algorithm) and feature-based methods. Based on the pixel-level motion vector field, the target's center position coordinates and the sum of its bounding box coordinates are updated. For example, by calculating the average of the motion vectors of all pixels within the target region, the target's overall motion direction and velocity can be determined. The search center parameter is determined based on the target's position coordinates in the current frame, with the center of the target in the current frame serving as the search center. If the target's bounding box changes in the current frame, the search center position needs to be adjusted to ensure the accuracy of the search range. The search center position is dynamically adjusted based on the target's motion trends and bounding box changes to ensure that the search range covers all possible target locations. The search range parameter defines the specific area in the image to search for the target. The search range is determined based on the target's size (such as the width and height of the bounding box) and the search center. The search range is typically expanded appropriately, for example, by a certain ratio (e.g., 1.5 times) to improve search robustness. Boundary processing is also performed to ensure that the search range does not exceed the image boundaries, and the search range is cropped if necessary. The final step is to predict the target's position parameters in the next frame based on the search range parameters. Combining the target's motion model (such as a uniform motion model or a uniformly accelerated motion model), the target's position in the next frame is predicted based on the current frame's position coordinates and motion vector field. The current frame's position coordinates and motion vector field are input into the Kalman filter, and its recursive estimation characteristics are used to further optimize the prediction results. The Kalman filter estimates the system state through two steps: prediction and update, and can effectively handle noise and uncertainty. The predicted position parameters are output, including the predicted center position coordinates and bounding box coordinates.
[0031] In step S14, the predicted position parameter is compared with the current frame detection parameter. When the difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold, it is determined that the target is not blocked.
[0032] The goal is to compare the predicted position parameters with the detection parameters of the current frame to determine whether the target is occluded. The purpose of this comparison is to determine whether the difference between the predicted and actual detected positions is within an acceptable range. A preset threshold is set based on historical data analysis. If the calculated difference is less than the preset threshold, it indicates that the predicted and actual detected positions are very close and the target is not occluded.
[0033] In step S15, when the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, it is determined that the target is blocked, and an occlusion processing mechanism is started to track the target.
[0034] In one implementation, the historical trajectory data, color histogram, LBP texture features and optical flow motion features of the target are obtained; the motion trend data of the target is estimated based on the historical trajectory data; the probability of the target position is estimated using a particle filter algorithm based on the motion trend data to obtain the area parameters where the target may appear; the color histogram, the LBP texture features and the optical flow motion features are jointly compressed and optimized using a lightweight convolutional neural network to obtain multi-feature compressed data; the sliding window is used to generate all candidate targets in the area based on the area parameters and multi-feature fusion data; the distance between the candidate target and the template target is calculated based on the candidate target, and the target is obtained. The Euclidean distance to each candidate target is calculated; the feature similarity between the candidate target and the template target is calculated based on the candidate target to obtain the cosine similarity of each candidate target; the similarity score of each candidate target is calculated based on the Euclidean distance and the cosine similarity; the candidate targets are sorted according to the similarity score, and if the feature matching degree between the candidate target with the highest similarity score and the historical target does not reach a preset threshold, the observation angle is adjusted according to the target motion trend estimated by the optical flow method, and the target compression feature is re-extracted and matched; if the feature matching degree between the candidate target with the highest similarity score and the historical target exceeds a preset threshold, the candidate target is determined to be the tracking target of the current frame, and the target position information is updated to track the target.
[0035] In one implementation, the target's speed, direction, and future position can be calculated and estimated based on historical trajectory data, thereby obtaining the target's motion trend data. Based on this motion trend data, a particle filter algorithm is used to predict the target's state, generating a set of particle samples. Each particle sample represents a possible location for the target. Calculations are performed based on these particle samples to obtain the observation likelihood of each particle. Particle weights are updated based on these observation likelihoods to obtain particle weight distribution data. Resampling is performed based on this weight distribution data to obtain parameters for the target's possible location.
[0036] In one implementation, the color histogram, the LBP texture features, and the optical flow motion features are simplified using low-rank decomposition and pruning techniques to obtain feature coding; a lightweight spatial attention layer is designed based on the feature coding; writing is performed based on the spatial attention layer to obtain a lightweight temporal memory unit; and a lightweight convolutional neural network is trained based on the temporal memory unit to obtain multi-feature compressed data.
[0037] The cosine similarity is calculated as follows: = in, is the cosine similarity, is the vector dot product.
[0038] The Euclidean distance is calculated as follows:
[0039] in, is the Euclidean distance, is the vector dot product.
[0040] The similarity score is calculated as follows: in, is the similarity score, is a weight parameter, is the cosine similarity, is the Euclidean distance, is the distance threshold.
[0041] The current frame image and target position prediction result are obtained, and the current frame image is input into the target detection model to obtain the current frame detection result. The difference between the current frame detection result and the target position prediction result is calculated. If the difference exceeds a preset threshold, the target is determined to be occluded, triggering the occlusion processing mechanism. Based on the historical frame information and the current frame image, the Kalman filter algorithm is used to predict the target position in the next frame as the target position estimate in the case of occlusion. The features of the occluded area of the target in the current frame image are extracted, and similar feature areas are searched in the historical frames to determine the possible location of the target. The Kalman filter predicted position and the feature matching position are combined to obtain the optimal estimated position of the target in the current frame. The optimal estimated position is used as the prediction of the target position in the next frame and is used for occlusion judgment and processing in subsequent frames. The target is continuously tracked. When the target is occluded in multiple consecutive frames, the target appearance model is updated to improve the matching accuracy when it reappears after occlusion.
[0042] In the similarity score calculation formula, the distance threshold (often denoted as T) is a key parameter used to control the impact of Euclidean distance on the similarity score. It normalizes or standardizes the Euclidean distance so that the Euclidean distance value can be compared and weighted on the same scale as the cosine similarity value.
[0043] It is important to note that the target's position coordinates in previous frames are recorded, and the color distribution features of the target region (such as histograms in RGB or HSV color space) are extracted. The local binary pattern (LBP) of the target region is extracted to describe texture information. The target's motion vector in consecutive frames is calculated using optical flow to describe the target's direction and velocity. A motion model (such as a constant velocity model) is fitted using historical trajectory data to predict the target's future position. For example, the target's trajectory is fitted using the least squares method to obtain velocity and acceleration. The particle filter algorithm randomly samples particles (hypothesized target locations) and, combining motion trends with observed data, calculates a weight for each particle to estimate the most likely location of the target. For example, if the target is likely to appear in the lower right corner of the image, the particle filter algorithm generates multiple hypothesized locations and selects the most likely location based on motion trends and observed data. A lightweight convolutional neural network (such as MobileNet) is used to fuse and compress color, texture, and motion features to generate a low-dimensional feature vector. In areas where targets may appear, a sliding window is used to traverse the image, extracting features from each window as candidate targets. For example, in the lower right corner of the image, a sliding window is used to extract multiple candidate targets (such as a red car, a red traffic sign, etc.). The Euclidean distance between each candidate target's feature vector and the template target's feature vector is calculated to measure their spatial distance. The cosine similarity between each candidate target's feature vector and the template target's feature vector is calculated to measure their directional similarity. The Euclidean distance and cosine similarity are combined to calculate a comprehensive similarity score (e.g., a weighted sum). If the highest-scoring candidate target's feature match with the historical target exceeds a preset threshold, it is determined to be the target and its location information is updated. If the match does not reach the threshold, the observation angle is adjusted using the optical flow method, and the features are re-extracted and matched.
[0044] Specifically, the object detection model identifies objects in the current frame and outputs their location, size, and category. For example, in a pedestrian tracking scenario, the detection model might output the pedestrian's bounding box coordinates (x, y, width, height) and a confidence score. This detection result is compared with the previously predicted object position, calculating the Euclidean distance or intersection over union (IoU) between the two. If the distance is greater than a preset threshold (e.g., 30 pixels) or the IoU is less than a threshold (e.g., 0.5), the object is considered potentially occluded. Once an occlusion detection is triggered, the system initiates the occlusion handling mechanism. The Kalman filter algorithm uses the object's historical motion trajectory and current observations to predict the object's likely location in the next frame. The filter maintains the object's state vector, including information such as position and velocity, and continuously optimizes the state estimate through two phases: prediction and update. In the case of occlusion, the prediction step becomes particularly important due to the lack of reliable observation data. Simultaneously, the system extracts visual features of the occluded area from the current frame, such as color histograms, texture features, or deep learning features. These features are used to search for similar regions in previous frames to estimate the object's likely location. Feature matching can use a sliding window approach, comparing feature similarity pixel by pixel across historical frames to identify the best matching position. The system then fuses the Kalman filter predicted position with the feature matched position to obtain the optimal estimated target position. The fusion method can be a simple weighted average or a more complex probabilistic fusion model. For example, if the Kalman filter predicted position is (100, 200) and the feature matched position is (110, 190), with weights of 0.6 and 0.4, respectively, the fused estimated position is (104, 196). The optimal estimated position serves as the target position prediction for the next frame and is used for occlusion detection and handling in subsequent frames. This prediction-update cycle maintains tracking continuity during brief target occlusions. When a target is occluded for multiple consecutive frames (e.g., five frames), the system updates the target's appearance model. The appearance model can be a color histogram, HOG features, or a deep learning feature representation. The update process may involve online feature learning or fine-tuning of model parameters to accommodate potential changes in the target's appearance. This integrated approach effectively handles target occlusion in complex scenes. For example, when tracking a specific customer in a crowded shopping mall, the system can maintain tracking even if the target is briefly obscured by other people. Kalman filtering provides motion prediction, while feature matching supplements appearance information. The combination of these two improves the accuracy of position estimation during occlusion. Dynamic updates to the appearance model ensure that the system can adapt to changes in the target's appearance, such as when a customer changes coats or puts on a hat. This robust occlusion handling mechanism significantly improves the performance and reliability of target tracking systems in real-world applications.
[0045] To facilitate understanding of the present invention, some preferred embodiments of the present invention are further described below.
[0046] The following describes the working process of the present invention using a common scenario as an example. Figure 2 , which is Figure 1 Schematic diagram of the working scenario of the method.
[0047] In an intelligent surveillance system, the goal is to track pedestrians entering the surveillance area in real time. The system first uses an object detection algorithm to obtain the initial position parameters of the pedestrian in the first frame of the video, including the bounding box coordinates and the center point position. In subsequent frames, the system continues to detect the target to obtain the detection parameters for the current frame. Next, it uses optical flow to calculate the pixel-level motion vector field of the target between adjacent frames, analyzing pixel intensity changes to infer the target's motion direction and velocity. This motion vector field is then input into a Kalman filter, which predicts the target's position in the next frame based on the target's motion model. The Kalman filter uses recursive estimation, combining the motion model with observation data to optimize the predicted target position. The system compares the predicted position parameters with the detection parameters in the current frame. If the difference between the two is less than a preset threshold (e.g., a Euclidean distance less than 10 pixels), the target is considered unobstructed and tracking continues. If the difference is greater than the threshold, the target is considered potentially occluded and the occlusion handling mechanism is activated. This occlusion handling mechanism involves collecting the target's historical trajectory data, including position, velocity, and acceleration information from previous frames, and extracting the target's color histogram, LBP texture features, and optical flow motion features. Based on historical trajectory data, a Kalman filter or other motion model is used to estimate the target's motion trend and predict its possible direction and speed. A particle filter algorithm is then used to probabilistically estimate the target's location, generating a set of particle samples, each representing a possible target location. The observation likelihood of each particle is calculated, and its weight is updated based on these probabilities. Resampling is then used to determine the parameters of the target's possible location. The color histogram, LBP texture features, and optical flow motion features are fed into a lightweight convolutional neural network for joint compression and optimization, generating multi-feature compressed data. This network uses low-rank decomposition and pruning techniques to simplify feature encoding and incorporates spatial attention layers and temporal memory units to further improve feature extraction efficiency and accuracy. Based on the parameters of the target's possible location, a sliding window or selective search algorithm is used to generate all candidate targets within that area. For each candidate target, the Euclidean distance and cosine similarity are calculated between it and the template target. These two distances are combined to generate a similarity score for each candidate target. Candidate targets are sorted based on their similarity scores. If the candidate with the highest similarity score does not meet a threshold for matching the historical target features, the observation angle is adjusted based on the target motion trend estimated by the optical flow method, and the target compression features are re-extracted and matched. If the matching degree exceeds a preset threshold, the candidate target is determined to be the tracking target for the current frame, and the target position information is updated and tracking continues. In practical applications, such as in a busy shopping mall monitoring scenario, pedestrians may be obscured by other pedestrians or shopping carts. With the above technical solution, the monitoring system can still accurately track targets even when pedestrians are obscured.When a pedestrian was briefly obscured by a shopping cart, the system used historical trajectory data and motion trend estimation, combined with a particle filter algorithm and multi-feature fusion, to successfully predict and relocate the target, ensuring tracking continuity. The system performed exceptionally well in complex scenarios, and its occlusion-handling mechanism enabled it to quickly resume tracking even when the target was rapidly moving or obscured, ensuring the effectiveness and reliability of the monitoring system.
[0048] In summary, the present invention discloses a method for tracking video targets based on multi-feature fusion, comprising obtaining initial position parameters and current frame detection parameters of a target; performing calculations based on the initial position parameters to obtain a pixel-level motion vector field; performing predictions based on the pixel-level motion vector field to obtain predicted position parameters of the target; comparing the predicted position parameters with the current frame detection parameters; and determining that the target is not occluded when the difference between the predicted position parameters and the current frame detection parameters is less than a preset threshold; and activating an occlusion processing mechanism to track the target when the difference between the predicted position parameters and the current frame detection parameters is greater than a preset threshold. The present invention constructs a multi-feature description model by extracting the target's color histogram, LBP texture features, and optical flow motion features. A lightweight convolutional neural network is used to jointly compress and optimize the features, reducing computational overhead. The Kalman filter and particle filter algorithms are combined to predict the target's position, generate candidate targets within a possible area, and calculate feature similarity. The weights of each feature in the multi-feature fusion model are dynamically adjusted, prioritizing the most robust features, effectively improving tracking accuracy and real-time performance. The present invention can solve the problem in the prior art of accumulation of tracking errors caused by premature switching of perspectives. In resource-constrained scenarios, the computational overhead of deep learning models is usually large, resulting in poor tracking accuracy and real-time performance.
[0049] Reference Figure 3 The second embodiment of the present invention provides a video image target tracking device based on multi-feature fusion, comprising: The data acquisition module is used to obtain the initial position parameters of the target and the current frame detection parameters; A data calculation module for calculating according to the initial position parameters to obtain a pixel-level motion vector field; A data prediction module for performing prediction based on the pixel-level motion vector field to obtain predicted position parameters of the target; a data comparison module, configured to compare the predicted position parameter with the current frame detection parameter, and determine that the target is not blocked when the difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold; The data judgment module is used to judge that the target is blocked when the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, and to start the blockage processing mechanism to track the target.
[0050] It should be noted that the video image target tracking device based on multi-feature fusion provided in an embodiment of the present invention is used to execute all the process steps of the video image target tracking method based on multi-feature fusion in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.
[0051] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a computer algorithm program. When the processor executes the computer program, the steps in the above-mentioned embodiments of the method for tracking a target in a video image based on multi-feature fusion are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as the data calculation module.
[0052] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0053] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.
[0054] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.
[0055] The memory can be used to store the computer programs and / or modules. The processor implements the various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0056] If the module / unit integrated into the electronic device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0057] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.
[0058] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A video image target tracking method based on multi-feature fusion, characterized in that: Executed by a computer, including: Get the target's initial position parameters and current frame detection parameters; Calculate according to the initial position parameters to obtain a pixel-level motion vector field; Predicting the target position based on the pixel-level motion vector field; Comparing the predicted position parameter with the current frame detection parameter, and determining that the target is not blocked when the difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold; When the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, it is determined that the target is blocked, and an occlusion processing mechanism is started to track the target.
2. The video image target tracking method based on multi-feature fusion according to claim 1, characterized in that: The calculating according to the initial position parameters to obtain a pixel-level motion vector field includes: Calculating the motion vector of the target between adjacent frames using an optical flow method according to the initial position parameters to obtain a first motion vector field of the target; Inputting the first motion vector field into a Kalman filter for calculation to obtain motion state parameters of the target; wherein the Kalman filter is a recursive estimation algorithm; Recursive estimation is performed based on the motion state parameters to obtain a pixel-level motion vector field.
3. The video image target tracking method based on multi-feature fusion according to claim 1, characterized in that: The step of performing prediction based on the pixel-level motion vector field to obtain predicted position parameters of the target includes: Calculating according to the pixel-level motion vector field to obtain the current frame position coordinates of the target; Determine the target's search center parameters based on the current frame position coordinates; Searching based on the search center parameter to obtain a target search range parameter; Prediction is performed based on the search range parameters to obtain predicted position parameters of the target.
4. The video image target tracking method based on multi-feature fusion according to claim 1, characterized in that: When the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, it is determined that the target is blocked, and an occlusion processing mechanism is started to track the target, including: Obtain the target's historical trajectory data, color histogram, LBP texture features, and optical flow motion features; Estimating the target's motion trend data based on the historical trajectory data; Probabilistically estimating the target position using a particle filter algorithm based on the motion trend data to obtain parameters of the area where the target may appear; Using a lightweight convolutional neural network to jointly compress and optimize the color histogram, the LBP texture feature, and the optical flow motion feature to obtain multi-feature compressed data; A sliding window is used to generate all candidate targets in the area according to the area parameters and multi-feature fusion data; Calculating the distance between the candidate target and the template target based on the candidate target to obtain the Euclidean distance of each candidate target; Calculating the feature similarity between the candidate target and the template target based on the candidate target to obtain the cosine similarity of each candidate target; Calculate the similarity score of each candidate target based on the Euclidean distance and the cosine similarity; Sorting is performed based on the similarity scores. If the matching degree between the candidate target with the highest similarity score and the historical target feature does not reach a preset threshold, the observation angle is adjusted based on the target motion trend estimated by the optical flow method, and the target compression feature is re-extracted and matched; If the matching degree between the candidate target with the highest similarity score and the historical target feature exceeds a preset threshold, the candidate target is determined to be the tracking target of the current frame, and the target position information is updated to track the target.
5. The video image target tracking method based on multi-feature fusion according to claim 4 is characterized in that: The method of using a particle filter algorithm to estimate the probability of the target position based on the motion trend data to obtain parameters of the area where the target may appear includes: A particle filter algorithm is used to predict the state of the target based on the motion trend data to obtain a set of particle samples; wherein the particle sample represents a possible location of the target; Calculating based on the particle sample to obtain the observation likelihood probability of each particle; The particle weights are updated according to the observation likelihood probability to obtain particle weight distribution data; Resampling is performed according to the weight distribution data to obtain parameters of the area where the target may appear.
6. The video image target tracking method based on multi-feature fusion according to claim 4 is characterized in that: The method of using a lightweight convolutional neural network to jointly compress and optimize the color histogram, the LBP texture feature, and the optical flow motion feature to obtain multi-feature compressed data includes: Simplifying the color histogram, the LBP texture feature, and the optical flow motion feature using low-rank decomposition and pruning technology to obtain feature coding; Designed based on the feature encoding, a lightweight spatial attention layer is obtained; Writing according to the spatial attention layer to obtain a lightweight temporal memory unit; A lightweight convolutional neural network is used for training according to the temporal memory unit to obtain multi-feature compressed data.
7. The video image target tracking method based on multi-feature fusion according to claim 4 is characterized in that: The step of calculating the feature similarity between the candidate target and the template target based on the candidate target to obtain the cosine similarity of each candidate target includes: The cosine similarity is calculated as follows: = ; in, is the cosine similarity, is the vector dot product.
8. The video image target tracking method based on multi-feature fusion according to claim 4 is characterized in that: The calculating the distance between the candidate target and the template target based on the candidate target to obtain the Euclidean distance of each candidate target includes: The Euclidean distance is calculated as follows: ; in, is the Euclidean distance, is the vector dot product.
9. The video image target tracking method based on multi-feature fusion according to claim 4, characterized in that: The calculation based on the Euclidean distance and the cosine similarity to obtain a similarity score for each candidate target includes: The similarity score is calculated as follows: ; in, is the similarity score, is a weight parameter, is the cosine similarity, is the Euclidean distance, is the distance threshold.
10. A video image target tracking device based on multi-feature fusion, characterized in that: include: The data acquisition module is used to obtain the initial position parameters of the target and the current frame detection parameters; A data calculation module for calculating according to the initial position parameters to obtain a pixel-level motion vector field; A data prediction module for performing prediction based on the pixel-level motion vector field to obtain predicted position parameters of the target; a data comparison module, configured to compare the predicted position parameter with the current frame detection parameter, and determine that the target is not blocked when the difference between the predicted position parameter and the current frame detection parameter is less than a preset threshold; The data judgment module is used to judge that the target is blocked when the difference between the predicted position parameter and the current frame detection parameter is greater than a preset threshold, and to start the blockage processing mechanism to track the target.
Citation Information
Patent Citations
Moving object tracking method under complicated background and sheltering condition
CN103077539A
Moving target tracking method under shielding background
CN110458862A
Real-time pedestrian tracking method applied to unmanned vehicle
CN111488795A
Target tracking method, tracking device and readable storage medium
CN113763428A
Moving target anti-shielding re-tracking method based on correlation filtering
CN118537365A
Cited By
Image processing transmission method suitable for display control tracking and related product
CN120856841A
Image inter-frame displacement detection method and device, storage medium and program product
CN121458999A
Coal gangue tracking method based on hand-eye system of sorting robot
CN121616851A
Coal gangue tracking method based on sorting robot hand-eye system
CN121616851B
Ship tracking method based on AIS signal and high-resolution satellite image
CN122289377A