A multi-target tracking method, system and computer storage medium
The unified cost matrix approach with adaptive weight adjustment and improved preprocessing enhances target tracking in extreme weather and complex scenarios, achieving higher precision and efficiency.
Patent Information
- Application Number
- CN202510520258.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing target tracking methods are difficult to achieve stable multi-target matching in extreme weather and complex scenarios, especially in heavy fog and heavy rain.
Using an improved multi-objective tracking method, a unified cost matrix of IoU and ReID feature similarity is constructed, and a unified cost matrix is dynamically adjusted by combining weather conditions and target density factors, one-to-one matching is performed, and a dark channel prior defogging algorithm and camera motion compensation are introduced to enhance feature extraction capabilities, combining Kalman filter and trajectory smoothing optimization.
In extreme weather and complex scenarios, the accuracy and stability of target tracking are significantly improved, the computing efficiency is improved, the system's robustness and adaptability are enhanced, and it is suitable for monitoring, search and rescue and traffic management fields.
Smart Images

Figure CN120107317B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target tracking, and more specifically, relates to a multi-target tracking method, system, and computer storage medium in extreme weather and complex scenarios. Background Art
[0002] With the increasing complexity of the traffic network and the growing demand for intelligent management, target tracking technology plays an increasingly important role in safety supervision, efficiency improvement, and accident prevention. In recent years, thanks to the rapid development of computer vision and deep learning technologies, vision-based target tracking methods have made significant progress. However, in practical applications, especially under extreme weather conditions such as heavy fog and rainstorms, or in complex scenarios such as dense targets and fast movements, target tracking still faces many challenges.
[0003] First of all, extreme weather such as heavy fog, rainstorms, and strong light will seriously affect the image quality, resulting in blurred or partially occluded targets, which greatly increases the difficulty of accurate tracking. Secondly, the environment is complex and changeable, and the targets may move quickly or turn suddenly. Traditional tracking algorithms often have difficulty coping with these situations in a timely manner and are prone to target loss. In addition, in areas with dense targets such as intersections or busy lanes, problems such as mutual occlusion and ID switching between multiple targets also pose great challenges to continuous and stable tracking.
[0004] Currently, commonly used target tracking methods mainly include algorithms based on SORT (Simple Online and Realtime Tracking) and their improved versions, such as DeepSORT, ByteTrack, etc. These algorithms perform well under ideal conditions, but still have obvious deficiencies in extreme weather and complex scenarios, mainly manifested in limited accuracy and matching efficiency of target matching, which limits their application in extreme weather and complex scenarios. Summary of the Invention
[0005] Aiming at the above defects or improvement requirements of the prior art, the present invention provides a multi-target tracking method, system, and computer storage medium, aiming to improve the accuracy and tracking efficiency of multi-target tracking in extreme weather and complex scenarios.
[0006] To achieve the above object, the present invention provides a multi-target tracking method, including:
[0007] Preprocess the target video data collected under extreme weather or complex scenarios to obtain each preprocessed frame image;
[0008] Input each preprocessed frame image into a trained target detection model, and output each target detection box included in each frame image;
[0009] Predict M target prediction boxes in the current frame image based on the M target detection boxes included in the previous frame image; calculate the intersection over union (IoU) between the N target detection boxes and the M target prediction boxes in the current frame image to construct an M×N dimensional IoU cost matrix; at the same time, calculate the ReID feature similarity between the N target detection boxes and the M target prediction boxes to construct an M×N dimensional ReID feature similarity cost matrix; use the weighted IoU cost matrix and the ReID feature similarity cost matrix as the data association cost matrix; wherein, input the image corresponding to the target detection box or the target prediction box into the trained ReID network to output the corresponding ReID feature vector;
[0010] Use the data association cost matrix as the input of the Hungarian algorithm to perform one-to-one matching of the N target detection boxes and the M target prediction boxes in the current frame image, realizing the association of targets in two adjacent frame images; then manage each target trajectory to achieve multi-target tracking.
[0011] Further, the IoU cost matrix is:
[0012]
[0013] wherein, represents the i th prediction box in the current frame image, represents the j th detection box in the current frame image, and the IoU function calculates the intersection over union of two bounding boxes;
[0014] The ReID feature similarity cost matrix is:
[0015]
[0016] wherein, and respectively represent the ReID feature vectors corresponding to the i th prediction box and the j th detection box in the current frame image; represents the L2 norm operation;
[0017] The data association cost matrix is:
[0018]
[0019] wherein, is an adjustable feature discrimination weight coefficient.
[0020] Further, the feature discrimination weight coefficient The determination method is as follows:
[0021] Calculate the average transmittance of the current frame image, and use the average transmittance as the visibility index of the current frame image ;
[0022] Normalize the brightness index , contrast index and the visibility index of the current frame image, and perform weighted fusion to obtain the comprehensive weather condition index ;
[0023] Calculate the target density factor in the current frame image :
[0024]
[0025] where N is the number of targets detected in the current frame image, W and H are the width and height of the current frame image respectively, is the ratio of the average target area to the area of the current frame image in the current frame image;
[0026] When the comprehensive weather condition index is small, increase the value of the feature discrimination weight coefficient ; when the comprehensive weather condition index is large, decrease the value of the feature discrimination weight coefficient ;
[0027] When the target density factor is large, decrease the value of the feature discrimination weight coefficient ; when the target density factor is small, increase the value of the feature discrimination weight coefficient ;
[0028] Furthermore, the feature discrimination weight coefficient is:
[0029]
[0030] where, is the reference weight, is the weather impact adjustment coefficient, is the target density impact adjustment coefficient.
[0031] Furthermore, the preprocessing includes performing defogging preprocessing on each frame of image using the dark channel prior defogging algorithm, or / and, performing camera motion compensation processing on each frame of image;
[0032] Among them, based on the brightness and contrast of each frame of image, the parameters for retaining a certain degree of fog in the dark channel prior dehazing algorithm are adaptively adjusted , specifically including:
[0033] Convert each frame of image into a grayscale image, calculate the average grayscale value of the grayscale image, and use it as the brightness index of each frame of image ;
[0034] Take the variance of the image grayscale distribution as the contrast index of each frame of image ;
[0035] According to the brightness index of each frame of image and the contrast index dynamically adjust the value in the dark channel prior algorithm:
[0036]
[0037] Among them: is the reference value, and are the influence coefficients of brightness and contrast respectively, and are the historical maximum values of the brightness index and the contrast index respectively.
[0038] Furthermore, the trained object detection model is an improved YOLOv8s model; the improved YOLOv8s model includes:
[0039] A feature extraction module for extracting the features of each frame of image to obtain the corresponding feature map ;
[0040] The CBAM module includes a channel attention enhancement unit and a spatial attention enhancement unit connected in series in sequence; among them, the channel attention enhancement unit is used to enhance the channel attention of each feature in the feature map to obtain the feature map after channel attention enhancement :
[0041]
[0042] In the formula, represents the average pooling operation, represents the maximum pooling operation, is the sigmoid activation function, is the multi-layer perceptron;
[0043] The spatial attention enhancement unit is used to enhance the spatial attention of the features in the feature map to obtain the feature map after spatial attention enhancement :
[0044]
[0045] In the formula, represents a 7x7 convolution operation, represents and perform a vector concatenation operation;
[0046] A prediction module, configured to predict the object detection results included in each frame of image based on the enhanced feature map;
[0047] A global-local attention module is added to the backbone network of the trained ReID network; the global-local attention module is used to enhance the features, and the implementation method is:
[0048]
[0049] In the formula, is the feature map obtained by extracting features from the object image corresponding to the object detection box or the object prediction box; represents global average pooling, represents the average pooling of the th local region in the image corresponding to the object detection box or the object prediction box, is the learned weight, D is the number of divided local regions, represents concatenating the global feature map and the local feature map ;
[0050] Furthermore, a Kalman filter is used to predict M object prediction boxes included in the current frame of image based on M object detection boxes included in the previous frame of image;
[0051] The management of each object trajectory includes:
[0052] For objects that are detected for consecutive frames but no object association is achieved, initialize the trajectory formed by the objects detected for consecutive frames as a new trajectory;
[0053] For objects for which the association is successfully achieved, update the state of the Kalman filter to obtain the object prediction boxes included in the next frame of image;
[0054] For objects that are not detected for consecutive frames, mark them as deleted and delete the corresponding objects.
[0055] Furthermore, after managing each target trajectory, trajectory smoothing and ID stability optimization processing are also included;
[0056] The trajectory smoothing adopts an improved exponential moving average method, and the calculation formula is:
[0057]
[0058] In the formula, is the estimated position of each target smoothed in the current frame, represents the estimated position of each target in the current frame, is the estimated position of each target smoothed in the previous frame, is the adaptive smoothing factor:
[0059]
[0060] In the formula, and are preset parameters, is the target estimated speed;
[0061] The ID stability optimization includes:
[0062] For targets that reappear after being occluded for a long time, multi-feature fusion scores based on appearance and motion are used for ID recovery; the multi-feature fusion scores The calculation method is:
[0063]
[0064] In the formula, 、 and respectively represent the ReID feature similarity, motion consistency score, and shape similarity score between two targets, 、 and are the corresponding weight coefficients;
[0065] If it exceeds the preset value, it is considered that the two targets are the same target, and the two targets are set to the same ID.
[0066] The present invention also provides a multi-target tracking system, including a computer-readable storage medium and a processor;
[0067] The computer-readable storage medium is used to store executable instructions;
[0068] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the multi-target tracking method described in any one of the above.
[0069] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the multi-object tracking method described in any one of the above is implemented.
[0070] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0071] (1) In the present invention, considering that the commonly used two-stage matching strategy constructs a data association cost matrix by using a sequential cascaded processing method. Specifically, first, preliminary matching is performed based on IoU information, and then ReID features are used for secondary matching of unmatched targets. This method has obvious deficiencies: First, since the IoU matching and ReID feature matching are split into two independent stages, the system cannot balance the importance of spatial position information (IoU information) and appearance feature information (ReID features) simultaneously when associating targets, and it is easy to produce suboptimal matching results in extreme weather or complex scenarios, resulting in low target tracking accuracy; Second, the staged processing method inevitably leads to the need to repeatedly use the Hungarian algorithm for target matching, significantly increasing the computational overhead of the algorithm and being unfavorable for improving the real-time performance of tracking in extreme weather or complex scenarios; Third, since the ReID feature matching in the second stage only processes the targets unmatched in the first stage, the incompleteness of this information utilization may cause some potentially better matches to be ignored, further reducing the target tracking accuracy.
[0072] In view of the above problems, the improved unified cost matrix matching strategy in the present invention realizes the following remarkable advantages by constructing a complete cost matrix of the IoU relationship and ReID feature similarity relationship between all prediction boxes and detection box pairs included in the current frame image: First, by constructing an IoU cost matrix and a ReID feature similarity cost matrix of the same dimension, the single cost matrix (data association cost matrix) obtained after weighting the two simultaneously fuses spatial position information and appearance feature information, can consider these two types of key information simultaneously under the framework of global optimization, and can dynamically balance the importance of the two types of key information through the adjustment of the weight coefficient, so as to obtain the optimal matching result in different scenarios; Second, since the repeated calculations caused by staged processing are avoided, and the IoU and ReID feature similarity are calculated in parallel, the computational efficiency is significantly improved, better meeting the requirements of real-time processing in extreme weather or complex scenarios; Third, by performing unified similarity calculation and optimal solution for all possible matching pairs, the integrity of information utilization and the global optimality of the matching result are ensured, effectively improving the tracking accuracy and stability.
[0073] (2) Further, the feature discrimination weight coefficient based on the weather condition index and target density factor proposed by the present invention The value adaptive adjustment mechanism realizes the intelligent dynamic regulation of matching weights by comprehensively evaluating multi-dimensional quality indicators such as image brightness, contrast, and visibility, as well as the density distribution characteristics of targets in the scene. It effectively solves the problem of errors caused by directly using automatic calculation of weight coefficients in extreme weather scenes and target-dense scenes. When encountering bad weather conditions, it can automatically increase the weight of position information to maintain stable tracking performance. In target-dense scenes, it copes with frequent occlusions and intersections by enhancing the importance of ReID features, thus maintaining good robustness and reliability in various scenarios and significantly improving the adaptability and tracking performance of the method in the actual application environment.
[0074] (3)Preferably, the present invention provides a specific feature discrimination weight coefficient adjustment formula, and based on it, the feature discrimination weight coefficient is adjusted, so that the method of the present invention can maintain good robustness and reliability in various scenarios, and the calculation method is relatively simple.
[0075] (4)Preferably, during the preprocessing of each frame of image, the present invention uses an improved dark channel prior dehazing algorithm to perform dehazing preprocessing on each frame of image. By introducing an adaptive parameter adjustment mechanism, the value is dynamically adjusted according to the overall brightness and contrast of the image to adapt to different degrees of foggy weather conditions, thereby improving the quality of the preprocessed image. By performing camera motion compensation processing on each frame of image, the influence of camera motion on target tracking can be eliminated, further improving the quality of each frame of image.
[0076] (5)Preferably, an improved YOLOv8s model is used to obtain the target detection results included in each frame of image. By adding a CBAM module to the backbone network, the ability to extract key features can be enhanced; and by introducing a global-local attention module into the ReID network, the ability to extract target detail features is enhanced.
[0077] (6)Furthermore, during the process of smoothing the target trajectory, the present invention dynamically adjusts the value according to the target motion state, which can maintain a high response speed when the target moves quickly, and provide a smoother trajectory when the target moves slowly; during the ID recovery process after a long-term occlusion, by weighting, comprehensively considering three similarity scores, namely the ReID feature similarity, motion consistency score, and shape similarity score between two targets, the calculation is more reliable.
[0078] All in all, the method of the present invention is particularly suitable for continuous tracking of targets in bad weather conditions such as heavy fog and heavy rain, as well as complex scenarios such as target density and fast movement. Description of the Drawings
[0079] Figure 1 This is the flowchart of the multi - target tracking method in the embodiment of the present invention.
[0080] Figure 2 This is the flowchart of the parallel multi - target matching strategy in the embodiment of the present invention.
[0081] Figure 3 This is the schematic diagram of the improved YOLOv8s model structure in the embodiment of the present invention.
[0082] Figure 4 This is the detailed structure diagram of the CBAM attention mechanism in the embodiment of the present invention.
[0083] Figure 5 This is the structure diagram of the improved ReID framework provided by the embodiment of the present invention. Detailed Embodiment
[0084] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0085] Embodiment 1
[0086] As Figure 1 、 Figure 2 shown, an embodiment of the present invention provides a multi - target tracking method, which is applicable to extreme weather and complex scenarios. The method mainly includes:
[0087] S1. Pre - process the target video data collected under extreme weather or complex scenarios to obtain each pre - processed frame image;
[0088] S2. Input each pre - processed frame image into a trained target detection model to output the target detection results included in each frame image. Among them, the target detection results include the target detection boxes (target positions), sizes and category information of each target;
[0089] S3. Predict M target prediction boxes in the current frame image based on the M target detection boxes included in the previous frame image; calculate the intersection over union (IoU) between the N target detection boxes included in the current frame image and the M target prediction boxes to construct an M×N-dimensional IoU cost matrix; at the same time, calculate the ReID feature similarity between the N target detection boxes and the M target prediction boxes to construct an M×N-dimensional ReID feature similarity cost matrix; after weighting the IoU cost matrix and the ReID feature similarity cost matrix, use it as the data association cost matrix; wherein, input the image corresponding to the target detection box or the target prediction box into the trained ReID network to output the corresponding ReID feature vector;
[0090] S4. Use the data association cost matrix as the input of the Hungarian algorithm to perform one-to-one matching on the N target detection boxes and the M target prediction boxes in the current frame image to realize the association of targets in adjacent two frame images; then manage each target trajectory to realize multi-target tracking.
[0091] As a preferred implementation, the preprocessing in S1 includes performing defogging preprocessing on each frame of image, or / and, performing camera motion compensation processing. In the embodiments of the present invention, a high-resolution camera device is used to collect target video data in extreme weather or complex scenarios at a frame rate of not less than 30 frames per second.
[0092] In the embodiments of the present invention, an improved dark channel prior defogging algorithm is used to perform defogging preprocessing on each frame of image, which specifically includes:
[0093]
[0094] where, represents the coordinate of any pixel point in each frame of image, represents the clear image after defogging, represents the original foggy image, A represents the atmospheric light value, represents the atmospheric transmittance, is a preset transmittance threshold. In the embodiments of the present invention, the estimation formula of the transmittance is:
[0095]
[0096] where, is a parameter for retaining a certain degree of fog; represents the local window centered on the pixel coordinate , c represents the color channel ; represents the image corresponding to the color channel c, represents the atmospheric light value corresponding to the color channel.
[0097] In the embodiments of the present invention, an adaptive parameter adjustment mechanism is introduced to dynamically adjust the value according to the overall brightness and contrast of the image to adapt to different degrees of foggy weather conditions, thereby improving the quality of the preprocessed image. The specific method includes:
[0098] Convert each frame of the image into a grayscale image, and calculate the average grayscale value of the grayscale image, which is used as the brightness index of each frame of the image :
[0099]
[0100] where W and H are the width and height of each frame of the image respectively, is the grayscale value at pixel point ;
[0101] Take the variance based on the grayscale distribution of the image as the contrast index of each frame of the image :
[0102]
[0103] where, is the average grayscale of the whole image.
[0104] Dynamically adjust the value in the dark channel prior algorithm according to the brightness index and the contrast index of each frame of the image:
[0105]
[0106] where, is the reference value (usually taken as 0.95), and are the influence coefficients of brightness and contrast respectively, and are the historical maximum values of each index.
[0107] To ensure the stability of the system, set upper and lower limit constraints on the value:
[0108]
[0109] where, and are the upper and lower limit values of the value respectively, and are taken according to experience.
[0110] As a preferred implementation, perform camera motion compensation processing on each frame of the image, specifically including:
[0111] Extract the feature points of every two adjacent frames of images; in the embodiments of the present invention, the FAST corner detection algorithm is used to extract the feature points of every two adjacent frames of images;
[0112] Use the RANSAC algorithm (Random Sample Consensus algorithm) to perform feature point matching on the feature points of every two adjacent frames of images, so as to delete the dissimilar feature points in every two adjacent frames of images and obtain the similar feature points in every two adjacent frames of images;
[0113] Based on the similar feature points, use the affine transformation model to estimate the camera motion matrix H;
[0114] Compensate the pixel points in each frame of image with the estimated camera motion matrix H to eliminate the influence of camera motion on target tracking. In the embodiments of the present invention, the compensation formula is:
[0115]
[0116] In the formula, and respectively represent the pixel coordinates in each frame of image before and after compensation, and H is a 3×3 affine transformation matrix, that is, the estimated camera motion matrix.
[0117] As a preferred implementation, as Figure 3 shown, in S2, the trained target detection model is an improved YOLOv8s model. This improved YOLOv8s model includes: a feature extraction module, a CBAM (Convolutional Block Attention Module) module, and a prediction module.
[0118] The feature extraction module is used to extract the features of each frame of image to obtain the corresponding feature map F.
[0119] The CBAM module includes a channel attention enhancement unit and a spatial attention enhancement unit connected in series in sequence, as Figure 4 shown; the channel attention enhancement unit performs channel attention enhancement on each feature in the feature map F to obtain the feature map after channel attention enhancement:
[0120]
[0121] Among them, represents the average pooling operation, represents the maximum pooling operation, is the sigmoid activation function, is the multi-layer perceptron.
[0122] The spatial attention enhancement unit is used to perform spatial attention enhancement on the feature map Spatial attention enhancement is performed on the features in it to obtain the feature map after spatial attention enhancement :
[0123]
[0124] Among them, represents a 7x7 convolution operation, represents the operation of and for vector concatenation operation.
[0125] The prediction module is used to predict the object detection results included in each frame of image based on the enhanced feature map; among them, the object detection results include the bounding boxes (object positions), sizes, and category information of each object.
[0126] In the embodiments of the present invention, the model training adopts a transfer learning strategy, based on the pre-trained YOLOv8s model, and uses the collected object dataset for fine-tuning. During the training process, techniques such as multi-scale training and mosaic data augmentation are introduced to improve the generalization ability of the model and the detection performance for small objects.
[0127] In the embodiments of the present invention, by adding a CBAM module to the backbone network, the ability to extract key features is enhanced.
[0128] As a preferred implementation, in S3, based on the M object detection boxes included in the previous frame of image, the Kalman filter is used to predict the M object prediction boxes included in the current frame of image. In the embodiments of the present invention, the state vector of the Kalman filter is defined as , including the center coordinates of the object (i.e., the center coordinates of the detection box), the area ratio s of the object position in the current frame of image, the aspect ratio r of the detection box, and its corresponding speed information (i.e., , respectively representing the change rates of the abscissa a, ordinate b, and area ratio s of the object center). The state equation and the observation equation are as follows:
[0129] State equation: ;
[0130] Observation equation: ;
[0131] In the formula, P is the state transition matrix, Q is the observation matrix, and are the process noise and the observation noise respectively, and both are assumed to follow a zero-mean Gaussian distribution; represents the state of the current frame (corresponding to the current time k ), represents the state corresponding to the previous frame, represents in the current frame Observation value
[0132] In the embodiment of the present invention, the intersection over union (IoU) between N target detection boxes and M target prediction boxes included in the current frame image is calculated to construct an M×N-dimensional IoU cost matrix; wherein, the constructed IoU cost matrix is :[[]]
[0133]
[0134] Wherein represents the i th prediction box in the current frame image, represents the j th detection box in the current frame image, and the IoU function calculates the intersection over union of two bounding boxes, represents the i th prediction box and the j th detection box in the current frame image, indicating the degree of difference in spatial position between them.
[0135] In the embodiment of the present invention, the ReID feature similarity between the N target detection boxes and the M target prediction boxes is calculated to construct an M×N-dimensional ReID feature similarity cost matrix; wherein, the constructed ReID feature similarity cost matrix :[[]]
[0136]
[0137] Wherein and respectively represent the ReID feature vectors (i.e., appearance feature vectors) corresponding to the i th prediction box and the j th detection box in the current frame image. The denominator term normalizes the ReID feature vectors to ensure that the similarity values are in the range of [0,1]. represents taking the L2 norm (Euclidean norm) of the ReID feature vector, represents the i th prediction box and the j th detection box in the current frame image, indicating the degree of difference in appearance features between them.
[0138] As a preferred implementation, as shown in Figure 5 , an improved ReID network is used to extract ReID features from the images corresponding to each target detection box or target prediction box. In the embodiment of the present invention, the improved ReID network includes: a feature extraction module, a global-local attention module (Global-Local Attention module), and a ReID feature vector extraction module.
[0139] The feature extraction module is used to extract features from the target images cropped from the current frame image based on the target detection boxes or target prediction boxes, and obtain the corresponding feature maps. ;
[0140] The global-local attention module combines global and local information to enhance the ability to extract target detail features. The specific implementation is as follows:
[0141]
[0142] In the formula, represents global average pooling, represents the average pooling of the th local region in the image corresponding to the target detection box or target prediction box, is the learned weight, is the number of divided local regions, represents concatenating the global feature map and the local feature map .
[0143] The ReID feature vector extraction module is used to obtain the ReID feature vectors of the images corresponding to each target detection box or target prediction box based on the features output by the global-local attention module.
[0144] In the embodiments of the present invention, by introducing a global-local attention module into the ReID network, the ability to extract target detail features is enhanced.
[0145] As a preferred implementation manner, based on the above two cost matrices and with the same dimension, the final data association cost matrix is constructed by a weighted method:
[0146]
[0147] Among them, is an adjustable feature discrimination weight coefficient, which is used to dynamically adjust the relative importance of spatial position information and appearance feature information in the matching process according to the specific scene characteristics and application requirements.
[0148] As a further design of the present invention, considering that under extreme weather conditions, due to the degradation of image quality, the feature extraction is unstable. If the automatically calculated feature discrimination weight coefficient Errors may occur, that is, the weights are prone to over - bias towards a certain feature, reducing the robustness of target tracking. To avoid over - relying on the automatically calculated feature discrimination weight coefficient when fusing IoU features and ReID features , the present invention proposes a feature discrimination weight coefficient adaptive adjustment mechanism based on weather conditions and target density, which dynamically adjusts the weight ratio of IoU information and ReID features in target association by evaluating the current image quality and target density, specifically including:
[0149] Based on the transmittance estimated by the dark channel prior algorithm , calculate the average transmittance of the current frame image and use this average transmittance as the visibility index of the current frame image :
[0150]
[0151] Normalize the brightness index , contrast index and visibility index of the current frame image and perform weighted fusion as the comprehensive weather condition index :
[0152]
[0153] Among them, , , are the weighting coefficients and satisfy ; , , are the historical maximum values of each index respectively; , the closer the value is to 1, the better the weather condition.
[0154] Calculate the target density factor in the current frame image :
[0155]
[0156] Among them, N is the number of targets detected in the current frame image, is the ratio of the average target area to the area of the current frame image in the current frame image. For the N targets detected in the current frame image, each target i is represented by its bounding box: , where is the upper - left coordinate of the bounding box, is the width of the bounding box, is the height of the bounding box, then The calculation formula is:
[0157]
[0158] Wherein, is the bounding box area of the i th target in the current frame image, represents the area ratio of a single target occupying the current frame image. The larger is, the larger the average target size is. When is close to 1, it means that the target occupies most of the image area. When is close to 0, it means that the target is relatively small in the image. The target density factor in the current frame image reflects the target density situation in the scene.
[0159] Based on the comprehensive weather condition index and the target density factor dynamically adjust the feature discrimination weight coefficient :
[0160] When the weather condition is poor (such as heavy fog, heavy rain, etc.), that is, when the comprehensive weather condition index is small, the appearance features of the target will be seriously affected, resulting in a decrease in the quality of ReID feature extraction and a reduction in reliability. At this time, the weight of the ReID feature should be reduced, and the weight of the IoU position information should be increased, that is, increase the value of the feature discrimination weight coefficient . When the weather condition is good, that is, when the comprehensive weather condition index is large, the appearance features of the target are clear, the ReID feature extraction quality is high, and the reliability is good. At this time, more reliance can be placed on the ReID feature for tracking, that is, reduce the value of the feature discrimination weight coefficient to increase the weight ratio of the ReID feature. When the target density is high, that is, when the target density factor is large, the probability of multi-target occlusion and intersection increases, and the reliability of the IoU information decreases (because the position overlap of multiple targets leads to inaccurate IoU calculation). At this time, more reliance should be placed on the ReID feature to distinguish different targets, that is, reduce the value of the feature discrimination weight coefficient . When the target density is low, that is, when the target density factor is small, the targets are independent of each other, the position relationship is clear, and the IoU information is more reliable. More reliance can be placed on the position information for tracking. At this time, increase the value of the feature discrimination weight coefficient .
[0161] As a preferred implementation manner, based on the above adjustment strategy, a specific feature discrimination weight coefficient adjustment method is provided in the embodiment of the present invention:
[0162]
[0163] Among them, is the reference weight, and it can take the value of 0.5. is the weather impact adjustment coefficient. is the target density impact adjustment coefficient, which can be adjusted according to the actual situation.
[0164] Using this adjustment formula for the feature discrimination weight coefficient adjustment can maintain good robustness and reliability in various scenarios, and the calculation method is relatively simple.
[0165] To ensure stability, upper and lower limit constraints are set for the value, and the final value is:
[0166]
[0167] Among them, is the minimum weight to ensure IoU information. is to ensure the minimum contribution of ReID appearance feature information; in the embodiments of the present invention, , .
[0168] As a preferred implementation method, the management of each target trajectory includes:
[0169] Trajectory initialization: For targets that are detected in consecutive frames but the target association is not achieved, the trajectory formed by the targets detected in consecutive frames is initialized as a new trajectory.
[0170] Trajectory update: For targets that have successfully achieved association, update the Kalman filter state to obtain the predicted bounding box of the target included in the next frame of the image.
[0171] Trajectory deletion: For targets that are not detected in consecutive frames, mark them as deleted and delete the corresponding targets.
[0172] As a further design of the present invention, after step S4, post-processing optimization is further included, specifically including: trajectory smoothing and ID stability optimization; among them, trajectory smoothing adopts the adaptive exponential moving average method, and ID stability optimization includes the ID recovery mechanism after long-term occlusion.
[0173] In the embodiments of the present invention, trajectory smoothing adopts an improved exponential moving average method, and its calculation formula is:
[0174]
[0175] Wherein, is the estimated position of each target smoothed in the current frame, represents the estimated position of each target in the current frame, is the estimated position of each target smoothed in the previous frame, is the adaptive smoothing factor. The present invention introduces an adaptive mechanism to dynamically adjust value according to the target motion state:
[0176]
[0177] Wherein, and are preset parameters, is the target estimated speed. Dynamically adjusting value can maintain a high response speed when the target moves quickly, and provide a smoother trajectory when the target moves slowly.
[0178] In the embodiment of the present invention, the ID recovery after long-term occlusion mainly includes:
[0179] For the target that reappears after being occluded for a long time, a multi-feature fusion strategy based on appearance and motion is adopted for ID recovery. The fusion score is calculated as follows:
[0180]
[0181] Wherein, , and respectively represent the ReID feature similarity, motion consistency score and shape similarity score between two targets, , and are the corresponding weight coefficients. If it exceeds the preset value, it is considered that the two targets are the same target and the IDs are the same. In the ID recovery process after long-term occlusion, by weighting, comprehensively considering the three similarities of the ReID feature similarity, motion consistency score and shape similarity score between two targets, the calculated credibility is higher.
[0182] Through the above technical solutions, the present invention can effectively address numerous challenges faced by target tracking under extreme weather conditions, including but not limited to: 1. By improving the dark channel prior dehazing algorithm, the image quality under extreme weather is enhanced, laying a foundation for subsequent processing. 2. Camera jitter problem: Camera motion compensation is introduced, effectively eliminating the impact of camera motion on target tracking and improving the stability of the system. 3. The improved YOLOv8s model enhances the detection ability of targets under complex backgrounds by introducing an attention mechanism, especially the detection performance for small targets and partially occluded targets. 4. An improved ReID network is adopted, introducing a global-local attention module, which improves the model's ability to distinguish similar targets and provides a reliable feature representation for multi-target tracking. 5. Through the parallel weighted IoU and ReID matching strategies and post-processing optimization, by fusing spatial position information and appearance feature information in a single cost matrix, the system can consider these two types of key information simultaneously within a globally optimized framework and dynamically balance their importance through adjustable weight coefficients, thereby obtaining optimal matching results in different scenarios; secondly, by avoiding the repeated calculations caused by phased processing and calculating the IoU and ReID feature similarities in parallel, the computational efficiency of the algorithm is significantly improved, better meeting the requirements of real-time processing; 6. A value adaptive adjustment mechanism based on the weather condition index and target density factor is designed, significantly enhancing the adaptability and tracking performance of the system in the actual application environment; 7. An adaptive trajectory smoothing algorithm is introduced, providing a smoother target trajectory while ensuring fast response; 8. During the ID recovery process after long-term occlusion, by weighting, comprehensively considering the three similarity scores of the ReID feature similarity, motion consistency score, and shape similarity score between two targets, the calculated credibility is higher. The method of the present invention is not only applicable to target tracking under extreme weather conditions but can also be extended and applied to target tracking scenarios in other complex environments, such as surveillance, search and rescue, traffic management, etc. The implementation of this method can significantly improve the efficiency and accuracy of safety supervision, providing reliable technical support for intelligent transportation and safety.
[0183] In addition, the method of the present invention has strong scalability and adaptability. By adjusting model parameters and optimization strategies, it can quickly adapt to different application scenarios and environmental conditions. For example, the parameters of the dehazing algorithm can be adjusted according to specific environmental characteristics and weather conditions; the network structure of the target detection model can be optimized according to the type and size of the target; the parameters of the multi-target tracking algorithm can be adjusted according to the specific requirements of the tracking task, etc.
[0184]
[0185] The multi-object tracking method provided by the present invention achieves high-performance object tracking in complex environments. This method not only improves the accuracy and stability of tracking, but also enhances the robustness and adaptability of the system, making an important contribution to the technological progress in the fields of security and intelligent transportation.
[0186] Embodiment 2
[0187] An embodiment of the present invention provides a multi-object tracking system, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the multi-object tracking method in Embodiment 1 above are implemented.
[0188] Among them, the so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. The processor executes the corresponding functions by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.
[0189] The related technical solutions are the same as above and will not be elaborated here.
[0190] Embodiment 3
[0191] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-object tracking method in Embodiment 1 above are implemented.
[0192] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0193] The related technical solutions are the same as above and will not be elaborated here.
[0194] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-object tracking method, characterized in that, Including: Preprocess the target video data collected under extreme weather or complex scenarios to obtain each preprocessed frame image; Input each preprocessed frame image into a trained object detection model and output each object detection box included in the frame image; Predict M object prediction boxes included in the current frame image based on the M object detection boxes included in the previous frame image; Calculate the intersection over union (IoU) between the N object detection boxes and the M object prediction boxes included in the current frame image to construct an M×N-dimensional IoU cost matrix; at the same time, calculate the ReID feature similarity between the N object detection boxes and the M object prediction boxes to construct an M×N-dimensional ReID feature similarity cost matrix; Use the weighted IoU cost matrix and the ReID feature similarity cost matrix as the data association cost matrix; wherein, input the image corresponding to the object detection box or object prediction box into a trained ReID network to output the corresponding ReID feature vector; Use the data association cost matrix as the input of the Hungarian algorithm to perform one-to-one matching on the N object detection boxes and the M object prediction boxes in the current frame image to realize the association of objects in adjacent two frame images; then manage each object trajectory to realize multi-object tracking; The IoU cost matrix is as follows: Among them, represents the i-th predicted bounding box in the current frame image, represents the j-th detected bounding box in the current frame image, and the IoU function calculates the intersection over union of the two bounding boxes; The ReID feature similarity cost matrix is as follows: Among them, and respectively represent the ReID feature vectors corresponding to the $i$-th predicted bounding box and the $j$-th detected bounding box in the current frame image; represents the L2 norm operation; The data association cost matrix is as follows: Among them, is an adjustable feature discrimination weight coefficient; The weight coefficient of feature distinctiveness is determined as follows: Calculate the average transmittance of the current frame image and use the average transmittance as the visibility index of the current frame image ; The brightness index of the current frame image , the contrast index and the visibility index are normalized and then weighted and fused to serve as the comprehensive weather condition index ; Calculate the target density factor in the current frame image : Among them, is the number of detected targets in the current frame image, and are the width and height of the current frame image respectively, is the ratio of the average target area in the current frame image to the area of the current frame image; When the comprehensive weather condition index is small, increase the value of the feature discrimination weight coefficient ; when the comprehensive weather condition index is large, decrease the value of the feature discrimination weight coefficient ; When the target density factor is large, reduce the value of the feature discrimination weight coefficient ; when the target density factor is small, increase the value of the feature discrimination weight coefficient .
2. The multi-object tracking method according to claim 1, wherein The feature discrimination weight coefficient is as follows: Among them, is the reference weight, is the weather impact adjustment coefficient, is the target density impact adjustment coefficient.
3. The multi-object tracking method according to claim 1 or 2, characterized in that The preprocessing includes using the dark channel prior dehazing algorithm to perform dehazing preprocessing on each frame image, or / and, performing camera motion compensation processing on each frame image; Among them, the parameters for retaining a certain degree of fog in the dark channel prior dehazing algorithm are adaptively adjusted based on the brightness and contrast of each frame of image , specifically including: Convert each frame of the image into a grayscale image, calculate the average grayscale value of the grayscale image, and use it as the brightness index of each frame of the image ; Take the variance of the grayscale distribution of the image as the contrast index for each frame of the image ; According to the brightness index of each frame of image and the contrast index dynamically adjust the value in the dark channel prior algorithm: Wherein: is the reference value, and are the influence coefficients of brightness and contrast respectively, and are the historical maximum values of the brightness index and the contrast index respectively.
4. The multi-target tracking method according to claim 1, characterized in that The trained object detection model is an improved YOLOv8s model; the improved YOLOv8s model includes: A feature extraction module, configured to extract features of each frame of image to obtain a corresponding feature map ; The CBAM module includes a channel attention enhancement unit and a spatial attention enhancement unit connected in series in sequence; wherein, the channel attention enhancement unit is used to perform channel attention enhancement on each feature in the feature map to obtain a feature map with enhanced channel attention : In the formula, represents the average pooling operation, represents the max pooling operation, is the sigmoid activation function, is the multi-layer perceptron; The spatial attention enhancement unit is used to enhance the spatial attention of the features in the feature map to obtain a feature map with enhanced spatial attention : In the formula, represents a 7x7 convolution operation, represents and for vector concatenation operation; A prediction module, configured to predict the object detection results included in each frame of image based on the enhanced feature map ; A global-local attention module is added to the backbone network of the trained ReID network; the global-local attention module is used to enhance features, and the implementation method is: In the formula, is the feature map obtained by extracting features from the target image corresponding to the target detection box or target prediction box; represents global average pooling, represents the average pooling of the th local region in the image corresponding to the target detection box or target prediction box, is the learned weight, is the number of divided local regions, represents concatenating the global feature map and the local feature map .
5. The multi-object tracking method according to claim 1, wherein Use a Kalman filter to predict M object prediction boxes included in the current frame image based on the M object detection boxes included in the previous frame image; The management of each object trajectory includes: For consecutive frames that are detected but no target with target association is achieved, the trajectory formed by the targets detected in consecutive frames is initialized as a new trajectory; For the objects that are successfully associated, update the Kalman filter state to obtain the object prediction boxes included in the next frame image; For consecutive targets for which no frames are detected, mark them as deleted and delete the corresponding targets.
6. The multi-object tracking method according to claim 1, wherein After managing each object trajectory, it also includes trajectory smoothing and ID stability optimization processing; The trajectory smoothing uses an improved exponential moving average method, and the calculation formula is: In the formula, is the smoothed estimated position of each target in the current frame, represents the estimated position of each target in the current frame, is the smoothed estimated position of each target in the previous frame, is the adaptive smoothing factor: In the formula, and are preset parameters, is the target estimated speed; The ID stability optimization includes: For a target that reappears after being occluded for a long time, an ID recovery is performed using a multi-feature fusion score based on appearance and motion; the multi-feature fusion score is calculated as follows: Wherein, , and respectively represent the ReID feature similarity, motion consistency score, and shape similarity score between two targets, , and are the corresponding weight coefficients; If it exceeds the preset value, the two targets are considered to be the same target, and the two targets are set to the same ID.
7. A multi-target tracking system, characterized in that Including a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the multi-object tracking method according to any one of claims 1-6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it realizes the multi-object tracking method according to any one of claims 1-6.
Citation Information
Patent Citations
Pedestrian re-identification method based on global-local feature dynamic alignment
CN113408492A
Target tracking method and device for multimedia data
CN116912508A
Multi-target tracking method based on fusion information association and camera motion compensation
CN117036397A
Airport aircraft target tracking method and system based on low earth orbit satellite combined monitoring
CN118552863A