Multi-target tracking method and system and computer storage medium
By constructing the IoU and ReID feature similarity cost matrix in extreme weather and complex scenarios, and combining Hungarian algorithms to match the target, the problem of multi-target tracking accuracy and efficiency in the existing technology in extreme weather and complex scenarios is solved, and high-precision and high-efficiency multi-target tracking is achieved, and the robustness and adaptability of the system are improved.
Patent Information
- Application Number
- CN202510520258.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In extreme weather and complex scenarios, existing target tracking methods are difficult to achieve high-precision and efficient multi-target tracking, especially in extreme weather conditions such as heavy fog and heavy rain or in complex scenarios with dense targets and rapid movement.
By preprocessing the target video data in extreme weather or complex scenarios, using the trained object detection model and ReID network, the IoU cost matrix and ReID feature similarity cost matrix are constructed, combined with the Hungarian algorithm to match one-to-one, realize the correlation of targets in two adjacent frames of images, and manage the trajectories of each target to achieve multi-objective tracking.
It significantly improves the multi-objective tracking accuracy and tracking efficiency in extreme weather and complex scenarios. By dynamically adjusting the feature distinction weight coefficient, it adapts to the importance of information in different scenarios, and improves the robustness and adaptability of the system.
Smart Images

Figure CN120107317A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target tracking, and more specifically, relates to a multi-target tracking method, system and computer storage medium in extreme weather and complex scenes. Background Art
[0002] With the increasing complexity of traffic networks and the growing demand for intelligent management, target tracking technology plays an increasingly important role in safety supervision, efficiency improvement and accident prevention. In recent years, thanks to the rapid development of computer vision and deep learning technology, vision-based target tracking methods have made significant progress. However, in practical applications, target tracking still faces many challenges, especially in extreme weather conditions such as heavy fog and heavy rain or in complex scenes such as dense targets and fast movement.
[0003] First, extreme weather conditions such as heavy fog, heavy rain, and strong light can seriously affect image quality, causing the target to be blurred or partially blocked, which greatly increases the difficulty of accurate tracking. Second, the environment is complex and changeable, and the target may move quickly or turn suddenly. Traditional tracking algorithms often have difficulty responding to these situations in a timely manner, which can easily cause the target to be lost. In addition, in areas with dense targets such as intersections or busy lanes, problems such as mutual occlusion between multiple targets and ID switching also pose great challenges to continuous and stable tracking.
[0004] At present, the commonly used target tracking methods mainly include algorithms based on SORT (Simple Online and Realtime Tracking) and its improved versions, such as DeepSORT, ByteTrack, etc. These algorithms perform well under ideal conditions, but still have obvious shortcomings in extreme weather and complex scenes, mainly in the limited accuracy and matching efficiency of target matching, which limits their application in extreme weather and complex scenes. Summary of the invention
[0005] In view of the above defects or improvement needs of the prior art, the present invention provides a multi-target tracking method, system and computer storage medium, which aims to improve the accuracy and tracking efficiency of multi-target tracking in extreme weather and complex scenarios.
[0006] To achieve the above object, the present invention provides a multi-target tracking method, comprising: Preprocess the target video data collected in extreme weather or complex scenes to obtain each frame of the preprocessed image; Input each preprocessed frame image into the trained target detection model, and output each target detection frame contained in each frame image; Based on the M target detection frames contained in the previous frame image, the M target prediction frames contained in the current frame image are predicted; the intersection over union (IoU) between the N target detection frames contained in the current frame image and the M target prediction frames is calculated to construct an M×N dimensional IoU cost matrix; at the same time, the ReID feature similarity between the N target detection frames and the M target prediction frames is calculated to construct an M×N dimensional ReID feature similarity cost matrix; the IoU cost matrix and the ReID feature similarity cost matrix are weighted as a data association cost matrix; wherein, the image corresponding to the target detection frame or the target prediction frame is input into the trained ReID network, and the corresponding ReID feature vector is output; The data association cost matrix is used as the input of the Hungarian algorithm to perform one-to-one matching between N target detection frames and M target prediction frames in the current frame image to achieve the association of targets in two adjacent frame images; then the trajectory of each target is managed to achieve multi-target tracking.
[0007] Furthermore, the IoU cost matrix for: in, Indicates the first i prediction boxes, Indicates the first j The IoU function calculates the intersection-over-union ratio of two bounding boxes. The ReID feature similarity cost matrix for: in, and Respectively represent the first i prediction box and j ReID feature vector corresponding to the detection box; Indicates taking the L2 norm operation; The data association cost matrix for: in, is an adjustable feature discrimination weight coefficient.
[0008] Furthermore, the feature discrimination weight coefficient The method of determining is: Calculate the average transmittance of the current frame image and use the average transmittance as the visibility index of the current frame image ; The brightness index of the current frame image , contrast index and the visibility metrics After normalization, weighted fusion is performed as a comprehensive index of weather conditions ; Calculate the target density factor in the current frame image : Where N is the number of targets detected in the current frame image, W and H are the width and height of the current frame image respectively. is the ratio of the average target area in the current frame image to the area of the current frame image; When the weather condition comprehensive index When it is smaller, increase the feature discrimination weight coefficient The value of weather condition comprehensive index When it is larger, reduce the feature discrimination weight coefficient The value of When the target density factor When it is larger, reduce the feature discrimination weight coefficient The value of; when the target density factor When it is smaller, increase the feature discrimination weight coefficient The value of .
[0009] Furthermore, the feature discrimination weight coefficient for: in, is the base weight, is the weather adjustment coefficient, It is the target density impact adjustment coefficient.
[0010] Furthermore, the preprocessing includes performing defogging preprocessing on each frame of image using a dark channel priori defogging algorithm, or / and performing camera motion compensation processing on each frame of image; The parameters of the dark channel priori defogging algorithm for retaining a certain degree of fog are adaptively adjusted based on the brightness and contrast of each frame of the image. , specifically including: Convert each frame of image into a grayscale image and calculate the average grayscale value of the grayscale image as the brightness index of each frame of image ; The variance of the image grayscale distribution is used as the contrast index of each frame image ; According to the brightness index of each frame image and contrast index Dynamically adjust the dark channel prior algorithm value: in: is the base value, and are the influence coefficients of brightness and contrast, respectively. and They are the historical maximum values of brightness and contrast indicators respectively.
[0011] Furthermore, the trained target detection model is an improved YOLOv8s model; the improved YOLOv8s model includes: Feature extraction module, used to extract the features of each frame image and obtain the corresponding feature map ; The CBAM module comprises a channel attention enhancement unit and a spatial attention enhancement unit connected in series; wherein the channel attention enhancement unit is used to perform channel attention enhancement on each feature in the feature map to obtain a feature map after channel attention enhancement. : In the formula, represents the average pooling operation, represents the maximum pooling operation, is the sigmoid activation function, is a multi-layer perceptron; The spatial attention enhancement unit is used to The features in the image are spatially enhanced to obtain the feature map after spatial attention enhancement. : In the formula, represents a 7x7 convolution operation, Express and Perform vector concatenation operations; A prediction module, used to predict the target detection results contained in each frame of the image based on the enhanced feature map; A global-local attention module is added to the backbone network of the trained ReID network; the global-local attention module is used to enhance the features, and the implementation method is as follows: In the formula, A feature map obtained by extracting features from a target image corresponding to a target detection frame or a target prediction frame; represents global average pooling, Indicates the first Average pooling of local regions, is the learned weight, D is the number of divided local areas, Represents the global feature map and local feature maps Make stitching.
[0012] Furthermore, a Kalman filter is used to predict the M target prediction frames contained in the current frame image based on the M target detection frames contained in the previous frame image; The management of each target trajectory includes: For continuous The frame is detected but the target association is not achieved, and the The trajectory formed by the detected target in the frame is initialized as a new trajectory; For the target that is successfully associated, the Kalman filter state is updated to obtain the target prediction box contained in the next frame image; For continuous If a frame is not detected, it is marked as deleted and the corresponding target is deleted.
[0013] Furthermore, after managing each target trajectory, it also includes trajectory smoothing and ID stability optimization processing; The trajectory smoothing adopts the improved exponential moving average method, and the calculation formula is: In the formula, is the estimated position of each target after smoothing in the current frame, represents the estimated position of each target in the current frame, is the estimated position of each target after smoothing in the previous frame, is the adaptive smoothing factor: In the formula, and is the preset parameter, Estimate the speed for the target; The ID stability optimization includes: For targets that reappear after being blocked for a long time, a multi-feature fusion score based on appearance and motion is used to restore the ID; the multi-feature fusion score The calculation method is: In the formula, , and Represent the ReID feature similarity, motion consistency score and shape similarity score between two targets respectively. , and is the corresponding weight coefficient; If the number of targets exceeds the preset value, the two targets are considered to be the same target and the same ID is set for the two targets.
[0014] The present invention also provides a multi-target tracking system, comprising a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute any of the multi-target tracking methods described above.
[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the multi-target tracking method as described in any one of the above items is implemented.
[0016] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects: (1) In the present invention, it is considered that the commonly used two-stage matching strategy adopts a cascade processing method to construct a data association cost matrix. Specifically, a preliminary matching is first performed based on the IoU information, and then the ReID feature is used to perform a secondary matching on the unmatched target. This method has obvious shortcomings: First, since the IoU matching and ReID feature matching are separated into two independent stages, the system cannot simultaneously weigh the importance of spatial location information (IoU information) and appearance feature information (ReID feature) when associating targets, which easily produces suboptimal matching results in extreme weather or complex scenes, resulting in low target tracking accuracy; second, the staged processing method inevitably leads to the need to repeatedly use the Hungarian algorithm for target matching, which significantly increases the computational overhead of the algorithm and is not conducive to improving the real-time tracking performance in extreme weather or complex scenes; third, since the ReID feature matching in the second stage only processes the targets that were not matched in the first stage, the incompleteness of this information utilization may cause some potential better matches to be ignored, further reducing the accuracy of target tracking.
[0017] Based on the consideration of the above problems, the improved unified cost matrix matching strategy in the present invention achieves the following significant advantages by constructing a complete cost matrix of the IoU relationship and the ReID feature similarity relationship between all the prediction boxes and detection box pairs contained in the current frame image: First, by constructing the IoU cost matrix and the ReID feature similarity cost matrix of the same dimension, the single cost matrix (data association cost matrix) obtained after weighting the two integrates both the spatial position information and the appearance feature information, and can simultaneously consider these two types of key information under the framework of global optimization, and can dynamically balance the importance of the two types of key information by adjusting the weight coefficient, so as to obtain the best matching results in different scenarios; second, by avoiding the repeated calculation caused by staged processing, and calculating the IoU and ReID feature similarity in parallel, the calculation efficiency is significantly improved, and the requirements of real-time processing in extreme weather or complex scenes are better met; third, by performing unified similarity calculation and optimization solution for all possible matching pairs, the integrity of information utilization and the global optimality of matching results are ensured, and the tracking accuracy and stability are effectively improved.
[0018] (2) Furthermore, the feature discrimination weight coefficient based on the weather condition index and the target density factor proposed in the present invention is The adaptive value adjustment mechanism realizes intelligent dynamic regulation of matching weights by comprehensively evaluating multi-dimensional quality indicators such as image brightness, contrast and visibility, as well as the density distribution characteristics of targets in the scene. It effectively solves the problem of errors caused by directly using automatic calculation of weight coefficients in extreme weather scenes and scenes with dense targets. It can automatically increase the weight of position information to maintain stable tracking performance when encountering severe weather conditions. In scenes with dense targets, it can cope with frequent occlusions and intersections by enhancing the importance of ReID features. This ensures good robustness and reliability in various scenarios, significantly improving the adaptability and tracking performance of the method in actual application environments.
[0019] (3) As a preferred embodiment, the present invention provides a specific feature discrimination weight coefficient Adjustment formula, based on which the feature discrimination weight coefficient is calculated The method of the present invention can maintain good robustness and reliability in various scenarios, and the calculation method is relatively simple.
[0020] (4) Preferably, in the process of preprocessing each frame of the image, the present invention adopts an improved dark channel prior defogging algorithm to preprocess each frame of the image, and introduces an adaptive parameter adjustment mechanism to dynamically adjust the overall brightness and contrast of the image. The value of can be adjusted to adapt to different degrees of foggy weather conditions, thereby improving the quality of the pre-processed image. By performing camera motion compensation processing on each frame of image, the influence of camera motion on target tracking can be eliminated, further improving the quality of each frame of image.
[0021] (5) As a preferred method, an improved YOLOv8s model is used to obtain the target detection results contained in each frame of the image. By adding a CBAM module to the backbone network, the ability to extract key features can be enhanced; and by introducing a global-local attention module in the ReID network, the ability to extract target detail features is enhanced.
[0022] (6) Furthermore, in the process of smoothing the trajectory of the target, the present invention dynamically adjusts the target according to the target motion state. The value can maintain a high response speed when the target moves quickly, and provide a smoother trajectory when the target moves slowly. In the ID recovery process after long-term occlusion, the three similarity scores of ReID feature similarity, motion consistency score and shape similarity score between two targets are comprehensively considered in a weighted manner, and the calculation credibility is higher.
[0023] In summary, the method of the present invention is particularly suitable for continuous target tracking in adverse weather conditions such as heavy fog and rainstorm, as well as in complex scenes such as densely packed targets and fast movements. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Flow chart of the multi-target tracking method in an embodiment of the present invention.
[0025] Figure 2 The figure is a flow chart of a parallel multi-target matching strategy in an embodiment of the present invention.
[0026] Figure 3 Schematic diagram of the improved YOLOv8s model structure in an embodiment of the present invention.
[0027] Figure 4 This is a detailed structural diagram of the CBAM attention mechanism in an embodiment of the present invention.
[0028] Figure 5 A structural diagram of the improved ReID framework provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0030] Example 1 like Figure 1 , Figure 2 As shown, an embodiment of the present invention provides a multi-target tracking method suitable for extreme weather and complex scenes. The method mainly includes: S1. Preprocessing target video data collected in extreme weather or complex scenes to obtain each frame of image after preprocessing; S2, input each frame of the preprocessed image into the trained target detection model, and output the target detection result contained in each frame of the image; wherein the target detection result includes each target detection frame (target position), size and category information; S3, predicting the M target prediction frames contained in the current frame image based on the M target detection frames contained in the previous frame image; calculating the intersection over union (IoU) between the N target detection frames contained in the current frame image and the M target prediction frames to construct an M×N dimensional IoU cost matrix; calculating the ReID feature similarity between the N target detection frames and the M target prediction frames at the same time to construct an M×N dimensional ReID feature similarity cost matrix; weighting the IoU cost matrix and the ReID feature similarity cost matrix as the data association cost matrix; wherein, inputting the image corresponding to the target detection frame or the target prediction frame into the trained ReID network, and outputting the corresponding ReID feature vector; S4. The data association cost matrix is used as the input of the Hungarian algorithm to perform one-to-one matching between the N target detection boxes and the M target prediction boxes in the current frame image to achieve the association between the targets in two adjacent frame images; then the target trajectories are managed to achieve multi-target tracking.
[0031] As a preferred implementation, the preprocessing in S1 includes performing defogging preprocessing on each frame of image, or / and performing camera motion compensation processing. In the embodiment of the present invention, a high-resolution camera device is used to collect target video data in extreme weather or complex scenes at a frame rate of not less than 30 frames per second.
[0032] In the embodiment of the present invention, an improved dark channel priori defogging algorithm is used to perform defogging preprocessing on each frame of image, specifically including: In the formula, Represents the coordinates of any pixel point in each frame of the image. represents the clear image after dehazing, represents the original foggy image, A represents the atmospheric light value, represents the transmittance of the atmosphere, is a preset transmittance threshold. In the embodiment of the present invention, the transmittance The estimation formula is: In the formula, To retain the parameters of a certain degree of fog; In pixel coordinates is the local window centered on, c represents the color channel ; Represents the image corresponding to color channel c, Indicates the atmospheric light value of the corresponding color channel.
[0033] In the embodiment of the present invention, an adaptive parameter adjustment mechanism is introduced to dynamically adjust the overall brightness and contrast of the image. The value of can be adjusted to adapt to different degrees of foggy weather conditions, thereby improving the quality of pre-processed images. The specific methods include: Convert each frame of image into a grayscale image and calculate the average grayscale value of the grayscale image as the brightness index of each frame of image : Among them, W and H are the width and height of each frame image respectively. Pixel The gray value at ; The variance based on the image grayscale distribution is used as the contrast index of each frame image : in, is the average grayscale of the entire image.
[0034] According to the brightness index of each frame image and contrast index Dynamically adjust the dark channel prior algorithm value: in, is the reference value (usually 0.95), and are the influence coefficients of brightness and contrast, respectively. and It is the historical maximum value of each indicator.
[0035] To ensure system stability, Set upper and lower limits for the value: in, and They are The upper and lower limits of the value are determined based on experience.
[0036] As a preferred implementation method, camera motion compensation processing is performed on each frame of image, specifically including: Extracting feature points of every two adjacent frames of images; In the embodiment of the present invention, the FAST corner point detection algorithm is used to extract feature points of every two adjacent frames of images; The RANSAC algorithm (random sample consensus algorithm) is used to match the feature points of each two adjacent frames of images, so as to delete the dissimilar feature points in each two adjacent frames of images and obtain the similar feature points in each two adjacent frames of images; Based on similar feature points, the camera motion matrix H is estimated using the affine transformation model; The estimated camera motion matrix H is used to compensate the pixels in each frame to eliminate the influence of the camera motion on the target tracking. In the embodiment of the present invention, the compensation formula is: In the formula, and They represent the pixel coordinates in each frame before and after compensation, respectively. H is a 3*3 affine transformation matrix, i.e., the estimated camera motion matrix.
[0037] As a preferred implementation, Figure 3 As shown, in S2, the trained target detection model is an improved YOLOv8s model. The improved YOLOv8s model includes: a feature extraction module, a CBAM (Convolutional Block Attention Module) module and a prediction module.
[0038] The feature extraction module is used to extract the features of each frame of the image and obtain the corresponding feature map F.
[0039] The CBAM module includes a channel attention enhancement unit and a spatial attention enhancement unit connected in series, such as Figure 4 As shown; the channel attention enhancement unit performs channel attention enhancement on each feature in the feature map F to obtain the feature map after channel attention enhancement : in, represents the average pooling operation, represents the maximum pooling operation, is the sigmoid activation function, A multi-layer perceptron.
[0040] The spatial attention enhancement unit is used to enhance the feature map The features in the image are spatially enhanced to obtain the feature map after spatial attention enhancement. : in, represents a 7x7 convolution operation, Express and Perform vector concatenation operation.
[0041] The prediction module is used to predict the target detection result contained in each frame of the image based on the enhanced feature map; wherein the target detection result includes each target detection box (target position), size and category information.
[0042] In the embodiment of the present invention, the model training adopts a transfer learning strategy, based on the pre-trained YOLOv8s model, and fine-tuned using the collected target data set. During the training process, multi-scale training, mosaic data enhancement and other technologies are introduced to improve the generalization ability of the model and the detection performance of small targets.
[0043] In the embodiment of the present invention, the CBAM module is added to the backbone network to enhance the ability to extract key features.
[0044] As a preferred implementation, in S3, based on the M target detection frames contained in the previous frame image, a Kalman filter is used to predict the M target prediction frames contained in the current frame image. In the embodiment of the present invention, the state vector of the Kalman filter is defined as , contains the center coordinates of the target (i.e. the center coordinates of the detection frame), the area ratio s of the target position to the current frame image, the aspect ratio r of the detection frame and its corresponding speed information (i.e. , respectively represent the rate of change of the target center abscissa a, ordinate b, and area ratio s). The state equation and observation equation are as follows: Equation of state: ; Observation equation: ; In the formula, P is the state transfer matrix, Q is the observation matrix, and are process noise and observation noise, respectively, both assumed to obey zero-mean Gaussian distribution; Indicates the current frame (corresponding to the current time k ) status, Indicates the state corresponding to the previous frame, Indicates the current frame Observed value of .
[0045] In the embodiment of the present invention, the intersection over union (IoU) between the N target detection frames and the M target prediction frames contained in the current frame image is calculated to construct an M×N dimensional IoU cost matrix; wherein the constructed IoU cost matrix is : in, Indicates the first i prediction boxes, Indicates the first j The IoU function calculates the intersection-over-union ratio of two bounding boxes. Indicates the first i prediction box and j The degree of difference in the spatial positions between the detection boxes.
[0046] In the embodiment of the present invention, the ReID feature similarities between the N target detection frames and the M target prediction frames are calculated to construct an M×N dimensional ReID feature similarity cost matrix; wherein the constructed ReID feature similarity cost matrix : in, and Respectively represent the first i prediction box and j The ReID feature vector (also known as the appearance feature vector) corresponding to the detection box is normalized by the denominator to ensure that the similarity value is within the range of [0,1]. Indicates the L2 norm (Euclidean norm) of the ReID feature vector. Indicates the first i prediction box and j The degree of difference in appearance features between the detection boxes.
[0047] As a preferred implementation, Figure 5 As shown, an improved ReID network is used to extract ReID features from images corresponding to each target detection frame or target prediction frame. In an embodiment of the present invention, the improved ReID network includes: a feature extraction module, a global-local attention module (Global-Local Attention module) and a ReID feature vector extraction module.
[0048] The feature extraction module is used to extract features of the target image cropped from the current frame image based on the target detection frame or target prediction frame to obtain the corresponding feature map ; The global-local attention module combines global and local information to enhance the ability to extract target detail features. The specific implementation is as follows: In the formula, represents global average pooling, Indicates the first Average pooling of local regions, is the learned weight, is the number of local regions divided, Represents the global feature map and local feature maps Make stitching.
[0049] The ReID feature vector extraction module is used to extract features based on the output of the global-local attention module. The ReID feature vector of the image corresponding to each target detection box or target prediction box is obtained.
[0050] In the embodiment of the present invention, by introducing a global-local attention module in the ReID network, the ability to extract target detail features is enhanced.
[0051] As a preferred implementation method, based on the above two cost matrices of the same dimension and , construct the final data association cost matrix by weighting : in, It is an adjustable feature discrimination weight coefficient, which is used to dynamically adjust the relative importance of spatial position information and appearance feature information in the matching process according to specific scene characteristics and application requirements.
[0052] As a further design of the present invention, considering that under extreme weather conditions, the feature extraction is unstable due to the degradation of image quality, if the automatically calculated feature discrimination weight coefficient is directly used when fusing the IoU feature and the ReID feature, This may cause errors, that is, the weight may be overly biased towards a certain feature, reducing the robustness of target tracking. In order to avoid over-reliance on the automatically calculated feature discrimination weight coefficient when fusing IoU features and ReID features, The present invention proposes a feature discrimination weight coefficient based on weather conditions and target density. The adaptive adjustment mechanism dynamically adjusts the weight ratio of IoU information and ReID features in target association by evaluating the current image quality and target density, including: Transmittance estimated based on dark channel prior algorithm , calculate the average transmittance of the current frame image, and use the average transmittance as the visibility index of the current frame image : The brightness index of the current frame image , contrast index and viewability metrics After normalization, weighted fusion is performed as a comprehensive index of weather conditions : in, , , is the weighting coefficient and satisfies ; , , are the historical maximum values of each indicator; , the closer the value is to 1, the better the weather conditions are.
[0053] Calculate the target density factor in the current frame image : Where N is the number of targets detected in the current frame image. is the ratio of the average target area in the current frame image to the area of the current frame image. For the N targets detected in the current frame image, each target i Represented by its bounding box: ,in, is the coordinate of the upper left corner of the bounding box, is the width of the bounding box, is the height of the bounding box, then The calculation formula is: in, The current frame image i The bounding box area of the object, Indicates the area ratio of a single target in the current frame image. The larger the value, the larger the average target size. When it is close to 1, it means that the target occupies most of the image area. When it is close to 0, it means that the target is relatively small relative to the image, and the target density factor in the current frame image is Reflects the target density in the scene.
[0054] Based on comprehensive index of weather conditions and the target density factor Dynamically adjust feature discrimination weight coefficient : When the weather conditions are bad (such as heavy fog, heavy rain, etc.), the comprehensive weather index When it is small, the appearance characteristics of the target will be seriously affected, resulting in a decrease in the quality of ReID feature extraction and a decrease in reliability. At this time, the weight of the ReID feature should be reduced and the weight of the IoU position information should be increased, that is, the feature discrimination weight coefficient should be increased. When the weather conditions are good, the comprehensive weather index When it is larger, the appearance features of the target are clear, the ReID feature extraction quality is high, and the reliability is good. At this time, more reliance can be placed on the ReID feature for tracking, that is, the feature discrimination weight coefficient is reduced. The value of is used to increase the weight ratio of the ReID feature. When the target density is high, the target density factor When it is larger, the probability of multiple targets blocking and crossing each other increases, and the reliability of IoU information decreases (because the positions of multiple targets overlap, resulting in inaccurate IoU calculation). At this time, more reliance should be placed on ReID features to distinguish different targets, that is, to reduce the feature discrimination weight coefficient. When the target density is low, the target density factor When it is smaller, the targets are independent of each other, the position relationship is clear, the IoU information is more reliable, and tracking can be more dependent on position information. In this case, the feature discrimination weight coefficient is increased. The value of .
[0055] As a preferred implementation method, based on the above adjustment strategy, a specific feature discrimination weight coefficient is provided in the embodiment of the present invention: Adjustment method: in, is the base weight, which can be 0.5. is the weather adjustment coefficient, The target density affects the adjustment coefficient, which can be adjusted according to the actual situation.
[0056] Use this adjustment formula to calculate the feature discrimination weight coefficient Adjustment can maintain good robustness and reliability in various scenarios, and the calculation method is relatively simple.
[0057] To ensure stability, The value sets the upper and lower limits, and the final value for: in, To ensure the minimum weight of IoU information, To ensure the minimum contribution of ReID appearance feature information; in the embodiment of the present invention, , .
[0058] As a preferred implementation method, each target trajectory is managed, including: Trajectory initialization: For continuous The frame is detected but the target association is not achieved, and the The trajectory formed by the detected target in the frame is initialized as a new trajectory; Trajectory update: For successfully associated targets, update the Kalman filter state to obtain the target prediction box contained in the next frame image; Track deletion: For continuous If a frame is not detected, it is marked as deleted and the corresponding target is deleted.
[0059] As a further design of the present invention, after step S4, post-processing optimization is also included, specifically including: trajectory smoothing and ID stability optimization; wherein, trajectory smoothing adopts an adaptive exponential moving average method, and ID stability optimization includes an ID recovery mechanism after long-term occlusion.
[0060] In the embodiment of the present invention, the trajectory smoothing adopts the improved exponential moving average method, and its calculation formula is: In the formula, is the estimated position of each target after smoothing in the current frame, represents the estimated position of each target in the current frame, is the estimated position of each target after smoothing in the previous frame, is the adaptive smoothing factor. The present invention introduces an adaptive mechanism to dynamically adjust the value: In the formula, and is the preset parameter, Estimate the speed of the target. Dynamically adjust according to the target's motion state A higher value can maintain a high response speed when the target moves quickly, and provide a smoother trajectory when the target moves slowly.
[0061] In the embodiment of the present invention, the ID recovery after long-term obstruction mainly includes: For targets that reappear after being occluded for a long time, a multi-feature fusion strategy based on appearance and motion is used to restore the ID. The fusion score is calculated as follows: In the formula, , and Represent the ReID feature similarity, motion consistency score and shape similarity score between two targets respectively. , and is the corresponding weight coefficient. If the value exceeds the preset value, the two targets are considered to be the same target and have the same ID. In the ID recovery process after long-term occlusion, the three similarities between the two targets, namely the ReID feature similarity, motion consistency score and shape similarity score, are comprehensively considered in a weighted manner to obtain a higher calculation reliability.
[0062] Through the above technical solutions, the present invention can effectively solve many challenges faced by target tracking under extreme weather conditions, including but not limited to: 1. Through the improved dark channel prior dehazing algorithm, the image quality under extreme weather conditions is improved, laying the foundation for subsequent processing. 2. Camera shake problem: The introduction of camera motion compensation effectively eliminates the impact of camera motion on target tracking and improves the stability of the system. 3. The improved YOLOv8s model enhances the detection ability of targets in complex backgrounds by introducing an attention mechanism, especially the detection performance of small targets and partially occluded targets. 4. The improved ReID network is adopted and the global-local attention module is introduced to improve the model's ability to distinguish similar targets and provide reliable feature representation for multi-target tracking. 5. Through parallel weighted IoU and ReID matching strategies and post-processing optimization, by integrating spatial location information and appearance feature information in a single cost matrix, the system can simultaneously consider these two types of key information under the framework of global optimization, and dynamically balance their importance through adjustable weight coefficients, so as to obtain the best matching results in different scenarios; secondly, by avoiding repeated calculations caused by staged processing, and calculating IoU and ReID feature similarities in parallel, the algorithm's computational efficiency is significantly improved, better meeting the requirements of real-time processing; 6. A weather condition index and target density factor based on the weather condition index and target density factor are designed. The adaptive value adjustment mechanism significantly improves the system's adaptability and tracking performance in actual application environments; 7. The introduction of an adaptive trajectory smoothing algorithm provides a smoother target trajectory while ensuring a fast response; 8. In the ID recovery process after a long period of occlusion, the three similarity scores between the two targets, namely the ReID feature similarity, motion consistency score and shape similarity score, are comprehensively considered in a weighted manner, and the calculation credibility is higher.
[0063] The method of the present invention is not only suitable for target tracking under extreme weather conditions, but can also be extended to target tracking scenarios in other complex environments, such as monitoring, search and rescue, traffic management, etc. The implementation of this method can significantly improve the efficiency and accuracy of safety supervision and provide reliable technical support for intelligent transportation and safety.
[0064] In addition, the method of the present invention has strong scalability and adaptability. By adjusting the model parameters and optimization strategies, it can quickly adapt to different application scenarios and environmental conditions. For example, the parameters of the defogging algorithm can be adjusted according to the specific environmental characteristics and weather conditions; the network structure of the target detection model can be optimized according to the type and size of the target; the parameters of the multi-target tracking algorithm can be adjusted according to the specific requirements of the tracking task, etc.
[0065] The multi-target tracking method provided by the present invention realizes high-performance target tracking in complex environments. The method not only improves the accuracy and stability of tracking, but also enhances the robustness and adaptability of the system, making an important contribution to the technological progress in the fields of safety and intelligent transportation.
[0066] Example 2 An embodiment of the present invention provides a multi-target tracking system, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the multi-target tracking method in the above-mentioned embodiment 1 when executing the computer program.
[0067] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory may be used to store computer programs and / or modules, and the processor executes corresponding functions by running or executing computer programs and / or modules stored in the memory and calling data stored in the memory.
[0068] The relevant technical solutions are the same as above and will not be described again here.
[0069] Example 3 An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-target tracking method in the above-mentioned embodiment 1 are implemented.
[0070] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0071] The relevant technical solutions are the same as above and will not be described again here.
[0072] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-target tracking method, characterized in that: include: Preprocess the target video data collected in extreme weather or complex scenes to obtain each frame of the preprocessed image; Input each preprocessed frame image into the trained target detection model, and output each target detection frame contained in each frame image; Predict the M target prediction frames contained in the current frame image based on the M target detection frames contained in the previous frame image; Calculate the intersection over union (IoU) between the N target detection frames and the M target prediction frames contained in the current frame image to construct an M×N dimensional IoU cost matrix; and simultaneously calculate the ReID feature similarity between the N target detection frames and the M target prediction frames to construct an M×N dimensional ReID feature similarity cost matrix; The IoU cost matrix and the ReID feature similarity cost matrix are weighted as a data association cost matrix; wherein the image corresponding to the target detection box or the target prediction box is input into the trained ReID network, and the corresponding ReID feature vector is output; The data association cost matrix is used as the input of the Hungarian algorithm to perform one-to-one matching between N target detection frames and M target prediction frames in the current frame image to achieve the association of targets in two adjacent frame images; then the trajectory of each target is managed to achieve multi-target tracking.
2. The multi-target tracking method according to claim 1, characterized in that: The IoU cost matrix for: in, Indicates the first i prediction boxes, Indicates the first j The IoU function calculates the intersection-over-union ratio of two bounding boxes. The ReID feature similarity cost matrix for: in, and Respectively represent the first i prediction box and j ReID feature vector corresponding to the detection box; Indicates taking the L2 norm operation; The data association cost matrix for: in, is an adjustable feature discrimination weight coefficient.
3. The multi-target tracking method according to claim 2, characterized in that: The feature discrimination weight coefficient The method of determining is: Calculate the average transmittance of the current frame image and use the average transmittance as the visibility index of the current frame image ; The brightness index of the current frame image , contrast index and the visibility metrics After normalization, weighted fusion is performed as a comprehensive index of weather conditions ; Calculate the target density factor in the current frame image : Where N is the number of targets detected in the current frame image, W and H are the width and height of the current frame image respectively. is the ratio of the average target area in the current frame image to the area of the current frame image; When the weather condition comprehensive index When it is smaller, increase the feature discrimination weight coefficient The value of weather condition comprehensive index When it is larger, reduce the feature discrimination weight coefficient The value of When the target density factor When it is larger, reduce the feature discrimination weight coefficient The value of; when the target density factor When it is smaller, increase the feature discrimination weight coefficient The value of .
4. The multi-target tracking method according to claim 3, characterized in that: The feature discrimination weight coefficient for: in, is the base weight, is the weather adjustment coefficient, It is the target density impact adjustment coefficient.
5. The multi-target tracking method according to any one of claims 1 to 4, characterized in that: The preprocessing includes performing defogging preprocessing on each frame of image using a dark channel priori defogging algorithm, or / and performing camera motion compensation processing on each frame of image; The parameters of the dark channel priori defogging algorithm for retaining a certain degree of fog are adaptively adjusted based on the brightness and contrast of each frame of the image. , specifically including: Convert each frame of image into a grayscale image and calculate the average grayscale value of the grayscale image as the brightness index of each frame of image ; The variance of the image grayscale distribution is used as the contrast index of each frame image ; According to the brightness index of each frame image and contrast index Dynamically adjust the dark channel prior algorithm value: in: is the base value, and are the influence coefficients of brightness and contrast, respectively. and They are the historical maximum values of brightness and contrast indicators respectively.
6. The multi-target tracking method according to claim 1, characterized in that: The trained target detection model is an improved YOLOv8s model; the improved YOLOv8s model includes: Feature extraction module, used to extract the features of each frame image and obtain the corresponding feature map ; The CBAM module comprises a channel attention enhancement unit and a spatial attention enhancement unit connected in series; wherein the channel attention enhancement unit is used to perform channel attention enhancement on each feature in the feature map to obtain a feature map after channel attention enhancement. : In the formula, represents the average pooling operation, represents the maximum pooling operation, is the sigmoid activation function, is a multi-layer perceptron; The spatial attention enhancement unit is used to The features in the image are spatially enhanced to obtain the feature map after spatial attention enhancement. : In the formula, represents a 7x7 convolution operation, Express and Perform vector concatenation operations; A prediction module, used to predict the target detection results contained in each frame of the image based on the enhanced feature map; A global-local attention module is added to the backbone network of the trained ReID network; the global-local attention module is used to enhance the features, and the implementation method is as follows: In the formula, A feature map obtained by extracting features from a target image corresponding to a target detection frame or a target prediction frame; represents global average pooling, Indicates the first Average pooling of local regions, is the learned weight, D is the number of divided local areas, Represents the global feature map and local feature maps Make stitching.
7. The multi-target tracking method according to claim 1, characterized in that: Using the Kalman filter, the M target prediction frames contained in the current frame image are predicted based on the M target detection frames contained in the previous frame image; The management of each target trajectory includes: For continuous The frame is detected but the target association is not achieved, and the The trajectory formed by the detected target in the frame is initialized as a new trajectory; For the target that is successfully associated, the Kalman filter state is updated to obtain the target prediction box contained in the next frame image; For continuous If a frame is not detected, it is marked as deleted and the corresponding target is deleted.
8. The multi-target tracking method according to claim 1, characterized in that: After managing each target trajectory, it also includes trajectory smoothing and ID stability optimization processing; The trajectory smoothing adopts the improved exponential moving average method, and the calculation formula is: In the formula, is the estimated position of each target after smoothing in the current frame, represents the estimated position of each target in the current frame, is the estimated position of each target after smoothing in the previous frame, is the adaptive smoothing factor: In the formula, and is the preset parameter, Estimate the speed for the target; The ID stability optimization includes: For targets that reappear after being blocked for a long time, a multi-feature fusion score based on appearance and motion is used to restore the ID; the multi-feature fusion score The calculation method is: In the formula, , and Represent the ReID feature similarity, motion consistency score and shape similarity score between two targets respectively. , and is the corresponding weight coefficient; If the number of targets exceeds the preset value, the two targets are considered to be the same target and the same ID is set for the two targets.
9. A multi-target tracking system, characterized in that: comprising a computer readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the multi-target tracking method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multi-target tracking method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Pedestrian re-identification method based on global-local feature dynamic alignment
CN113408492A
Unmanned aerial vehicle task planning method and system based on comprehensive empowerment and storage medium
CN115755953A
Target tracking method and device for multimedia data
CN116912508A
Multi-target tracking method based on fusion information association and camera motion compensation
CN117036397A
Self-adaptive environment fog penetration enhancement method and enhancement system
CN118469865A
Cited By
Multi-target tracking detection method, device and equipment and readable storage medium
CN120726100A
Strong convection weather tracking and early warning method based on radar data and deep learning
CN120762034A
Target tracking and early warning method, system and equipment
CN121437914A
Target tracking and warning method, system, and device
CN121437914B