Multi-target tracking method based on three-dimensional depth information decoupling
By introducing three-dimensional deep information decoupling technology and improved clustering algorithms into the multi-objective tracking algorithm, the problem of target positioning and identity maintenance in complex scenarios is solved, and efficient and robust multi-objective tracking effect is achieved.
Patent Information
- Application Number
- CN202411981440.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In complex scenarios, multi-objective tracking algorithms are difficult to achieve accurate and robust target positioning and identity maintenance under dense occlusion and interference from similar objects.
A multi-objective tracking method based on three-dimensional deep information decoupling is adopted. By constructing a multi-objective tracking model including attention mechanism feature extraction network, camera pose estimation network, depth estimation prediction head and object detection prediction head, combined with the improved Mean-shift clustering algorithm and deep information decoupling strategy, data association and tracking of dense occlusion and non-intensive occlusion targets are achieved.
It effectively solves the problem of inaccurate IOU matching caused by dense occlusion, realizes multi-objective tracking in complex scenarios, and improves the accuracy of target positioning and the robustness of identity maintenance.
Smart Images

Figure CN119941789A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision tasks, and in particular relates to a multi-target tracking method based on three-dimensional depth information decoupling. Background Art
[0002] Multiple object tracking (MOT) is a key computer vision task that aims to track the motion trajectories of multiple targets (such as pedestrians) in a video sequence in real time. With the rapid development of deep learning and artificial intelligence technology, MOT has been widely used in various fields. For example, in intelligent monitoring systems, multiple pedestrian tracking technology is widely used in public safety monitoring. By deploying cameras at key locations in the city and using MOT technology to track the movement trajectory of the crowd in real time, security personnel can promptly detect abnormal behaviors, such as gang fights and suspicious packages left behind, thereby improving the level of urban safety management. In the field of autonomous driving, autonomous vehicles need to identify and track multiple dynamic targets such as pedestrians and bicycles in complex road environments to avoid collisions and ensure the safety of passengers and pedestrians. Through deep learning algorithms, the MOT system can accurately locate and predict the movement trajectory of pedestrians in real-time traffic scenes, so as to make reasonable driving decisions. Intelligent traffic management is another important application field of multiple pedestrian tracking technology. In a crowded traffic environment, MOT technology can help traffic management departments monitor road conditions in real time, identify and track pedestrian movements, optimize traffic light control strategies, alleviate traffic congestion, and improve traffic efficiency. In addition, this task is also the basis of advanced computer vision tasks such as posture estimation, behavior recognition, behavior analysis, and video analysis. In short, the application background of multi-target tracking technology in the field of artificial intelligence is extensive and far-reaching. With the continuous advancement of computer vision and deep learning technology, MOT technology will further promote the intelligent development of all walks of life and bring more convenience and safety to society.
[0003] The main task of multi-target tracking is to output the motion trajectories of all targets from a given video and maintain the identity information (Identity, ID) of each target. The tracking target can be a pedestrian, a vehicle or other objects. The current mainstream tracking method splits the MOT task into three subtasks: target detection, feature extraction, and data association. This idea has been well developed. However, due to challenges such as occlusion and interference from similar objects in the actual tracking process, maintaining robust tracking is still a current research difficulty. In order to meet the requirements of accurate, robust, and real-time tracking of multiple targets in complex scenes, further research and improvement of the MOT algorithm is needed. However, robust tracking in complex scenes is still a current research difficulty, which is mainly reflected in the following three aspects: ① Frequent occlusion during tracking makes it difficult to accurately locate the target ② Different targets may have a high appearance similarity, which increases the difficulty of maintaining the target ID ③ Interaction between targets may cause tracking box drift.
[0004] Existing multi-target tracking algorithms usually rely on the joint action of the detection module, appearance feature module and position prediction module. The existing basic detection algorithm can extract the deep feature information of the target to complete accurate target detection and obtain accurate detection results. However, for the appearance feature extraction module, on the one hand, the feature differentiation of the pedestrian target itself is not high, and on the other hand, in the case of dense occlusion, the appearance feature information extracted by the detection frame is inaccurate due to occlusion. Therefore, frequent ID switches will occur in the case of occlusion and target overlap. Summary of the invention
[0005] To solve the above technical problems, the present invention provides a multi-target tracking method based on three-dimensional depth information decoupling, which solves the problem of inaccurate IOU (Intersection over Union) matching caused by dense occlusion and realizes multi-target tracking in complex scenarios with multiple pedestrian targets.
[0006] The technical solution adopted by the present invention is: a multi-target tracking method based on three-dimensional depth information decoupling, the specific steps are as follows:
[0007] S1. Build a multi-target tracking model based on three-dimensional depth information decoupling;
[0008] The multi-target tracking model includes: 1 feature extraction network based on attention mechanism, 1 camera posture estimation network, 3 depth estimation prediction heads, and 3 target detection prediction heads.
[0009] S2, training the multi-target tracking model constructed in step S1;
[0010] S3, based on the multi-target tracking model trained in step S2, input the image at the current time t, and obtain the target detection data of the image at the time t and the three-dimensional depth information at the time t;
[0011] Among them, each detected target is represented by a bounding box without ID information. The box without ID is the detection box; the target detection prediction head outputs multiple pedestrian target boxes for each input image.
[0012] S4, based on step S3, clustering multiple detection frames at time t and trajectory frames at time t-1 is performed using an improved Mean-shift clustering algorithm to identify dense occlusions, and each class is a dense occlusion;
[0013] Among them, when the number of targets in each clustering result exceeds 5, it is regarded as dense occlusion; for "dense occlusion" with less than 5 targets, they are regarded as targets with less severe occlusion or even no occlusion, that is, non-dense occlusion targets.
[0014] S5, based on step S4, realizing data association between non-dense occlusion targets and dense occlusion targets;
[0015] S6. Process the unmatched detection frame targets and trajectory frame targets in step S5, track the targets frame by frame to obtain target trajectories, and implement multi-target tracking based on depth information decoupling.
[0016] Furthermore, the step S2 is specifically as follows:
[0017] S21, extracting deep feature information through a feature extraction network based on the attention mechanism;
[0018] The feature extraction network is a U-shaped backbone network, including: a Conv-stem convolution module, an average pooling layer, a stage 2 convolution layer, a stage 3 convolution layer, a stage 4 convolution layer, and three upsampling and convolution modules.
[0019] Among them, the convolution kernel size of the stage 1-4 convolution layers is 3×3, and the step size is 2; the stage 2 convolution layer and the stage 3 convolution layer both include: 1 downsampling layer, 3 continuous expansion convolution modules CDC, and 1 local-global feature interaction module LGFI; the stage 4 convolution layer includes: 1 downsampling layer, 6 continuous expansion convolution modules CDC, and 1 local-global feature interaction module LGFI. The upsampling and convolution module includes: 1 upsampling layer and 1 convolution layer.
[0020] The feature extraction network is used to extract deep feature information, that is, the encoder of like-UNet is used to extract deep feature information, which is divided into four stages, as follows:
[0021] 1) Phase 1:
[0022] The original image of size H×W×3 is used as the input feature and input into the Conv-stem convolution module for downsampling to obtain a size of H / 2×W / 2×C1 tmp1 The output feature map of .
[0023] Among them, H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image; the Conv-stem module consists of three convolutional layers. The first layer uses a 3×3 convolution kernel with a step size of 2 to achieve convolution to achieve the effect of downsampling, while the next two layers use a 3×3 convolution kernel with a step size of 1 to achieve convolution to achieve the effect of local feature extraction.
[0024] 2) Second stage:
[0025] The original image is processed by the average pooling layer to obtain a feature map of size H / 2×W / 2×3, which is different from the feature map of size H / 2×W / 2×C1 obtained in the first stage. tmp1 The feature maps of the image are concatenated to obtain a feature map of size H / 2×W / 2×C1 as the input feature, and then input into the convolution layer with a convolution kernel size of 3×3 and a step size of 2 for downsampling to obtain a feature map of size H / 4×W / 4×C2 tmp1 The downsampled features are then passed through three consecutive dilated convolution modules CDC and one local-global feature interaction module LGFI to obtain a size of H / 4×W / 4×C2 tmp2 The output feature map of .
[0026] Among them, the window size of the average pooling layer in this stage is set to H / 2×H / 2, and the step size is set to H / 2, that is, the original image with a size of H×W×3 passes through this average pooling layer, and the output size is H / 2×W / 2×3.
[0027] 3) The third stage:
[0028] The size obtained in the second stage is H / 4×W / 4×C2 tmp1 The size of the initial down-sampling feature map and the final output of the stage is H / 4×W / 4×C2 tmp2 The feature map of the original image after the average pooling layer is concatenated to obtain a feature map of size H / 4×W / 4×C2 as the input feature, which is then input into a convolution layer with a convolution kernel size of 3×3 and a stride of 2 for downsampling to obtain a feature map of size H / 8×W / 8×C3. tmp1 The downsampled feature map is then passed through three CDC modules and one LGFI module to obtain a size of H / 8×W / 8×C3 tmp2 The output feature map of .
[0029] Among them, the average pooling window size in this stage is set to H / 4×H / 4, the step size is set to H / 4, and the output size of the average pooling layer is H / 4×W / 4×3.
[0030] 4) The fourth stage:
[0031] The size obtained in the third stage is H / 8×W / 8×C3 tmp1 The size of the initial down-sampling feature map and the final output of the stage is H / 8×W / 8×C3 tmp2 The feature map of the original image and the feature map of size H / 8×W / 8×3 after the average pooling layer are concatenated to obtain a feature map of size H / 8×W / 8×C3 as the input feature, which is then input into a convolution layer with a convolution kernel size of 3×3 and a stride of 2 for downsampling, and the sampling result is input into 6 CDC modules and 1 LGFI module to obtain an output feature map of size H / 16×W / 16×C4.
[0032] Among them, the average pooling window size in this stage is set to H / 8×H / 8, the step size is set to H / 8, and the output size of the average pooling layer is H / 8×W / 8×3.
[0033] S22, fusing the multi-scale feature information extracted in step S21 through the depth estimation prediction head to obtain depth estimation information of different scales, and then fusing and reconstructing the multi-scale depth estimation information through the camera pose estimation completed by the camera pose estimation network Pose Net to obtain the final three-dimensional depth information, thereby realizing the reconstruction of the three-dimensional depth feature information;
[0034] S221, by fusing the features of different stages obtained in step S21, and then obtaining depth estimation information of three different scales through a depth estimation prediction head, that is, multi-scale depth estimation information;
[0035] (1) The features of size H / 16×W / 16×C4 obtained in the fourth stage are input into the first upsampling and convolution module for upsampling and residual connection. The size of the third stage is H / 8×W / 8×C3 tmp2 The features are then passed through a convolutional layer to obtain a size of H / 8×W / 8×C3 ~ The fused features are input into the depth estimation prediction head, and finally the depth estimation information of size H / 4×W / 4×1 is obtained.
[0036] Among them, the depth estimation prediction head includes: a convolution layer Conv, an upsampling layer and an activation layer Sigmoid.
[0037] (2) The size is H / 8×W / 8×C3 ~The fusion features are input into the second upsampling and convolution module for upsampling and residual connection. The size of the second stage is H / 4×W / 4×C2 tmp2 The features are then passed through a convolutional layer to obtain a size of H / 4×W / 4×C2 ~ The fused features are input into another depth estimation prediction head with the same structure, and finally the depth estimation information of size H / 2×W / 2×1 is obtained.
[0038] (3) The size is H / 4×W / 4×C2 ~ The fused features are input into the third upsampling and convolution module for upsampling, and then directly pass through the convolution layer to obtain the feature H / 2×W / 2×C1 ~ , and input it into the depth estimation prediction head of the same structure, and finally obtain the depth estimation information of size H×W×1.
[0039] S222, obtaining camera pose estimation information of two adjacent frames, that is, completing camera pose estimation through a camera pose estimation network Pose Net;
[0040] For monocular camera training, the camera pose estimation network is formed by ResNet18, with a pair of color images or six channels as input, and a four-layer convolutional pose decoder is used to estimate the corresponding 6-DOF relative pose between two adjacent frames. At the same time, horizontal inversion and random brightness, contrast, saturation and hue jitter with a range of ±0.2, ±0.2, ±0.2 and ±0.1 are performed.
[0041] Among them, the six degrees of freedom are the translational freedom and rotational freedom of the x-axis, y-axis, and z-axis in the three-dimensional coordinate system. Through these six degrees of freedom, the camera's posture is represented by a 4x4 homogeneous transformation matrix T, that is, a 3x3 rotation matrix R and a 3x1 translation vector z.
[0042] S223, correcting the original rough three-dimensional depth information obtained in step S21, that is, based on step S222, fusing and reconstructing the multi-scale depth estimation information obtained in step S221 to obtain the final three-dimensional depth information, and reconstructing the three-dimensional depth feature information through projection change and linear interpolation;
[0043] The first is the projection transformation, which projects each pixel in the depth map into the world coordinate system, and then projects it to the image plane of the target perspective according to the camera posture transformation matrix.
[0044] For each pixel point (u, v), create a pixel coordinate grid, convert the pixel coordinates to normalized coordinates, and then multiply the normalized coordinates by the depth estimation information d of the corresponding pixel point to obtain the 3D point in the world coordinate system, and then use the camera's attitude transformation matrix T to project the 3D point from the current view to the target view, and use the camera intrinsic parameter matrix K of the target view to project the 3D point back to the image plane, and normalize the projected coordinates to the image plane, that is, from three-dimensional coordinates (x, y, z) to two-dimensional coordinates (x / z, y / z, 1).
[0045] Then bilinear interpolation is performed to map the pixel value of the projection point to the target image on the image plane of the target perspective.
[0046] Normalize the projection coordinates to the range [-1,1] of the image plane, and use the bilinear interpolation method to interpolate the pixel values corresponding to the normalized projection coordinates from the image of the target perspective. Finally, the depth estimation information corrected by the camera posture estimation is obtained, that is, the final three-dimensional depth estimation map. The calculation expression is as follows:
[0047] Interpolation pixel value = (1-α)(1-β)P 00 +α(1-β)P 10 +(1-α)βP 01 +αβP 11
[0048] Among them, α and β represent the horizontal and vertical fractional parts of the normalized projection coordinates, respectively, and P ij Represents the four most recent pixel values.
[0049] The depth estimation prediction head learning objective modeling is to minimize the target image I t And the synthetic target image The image reconstruction loss between and an edge-aware smoothness loss constrained on the predicted depth map
[0050] The depth estimation prediction head loss function expression is as follows:
[0051]
[0052] in, represents the image reconstruction loss, I t represents the target image, represents the synthetic target image, m represents the fixed threshold of 0.85, SSIM(*) represents the structural similarity index between pixels, and ∥*∥ represents the L1 similarity characteristic between pixels, which is used to characterize the similarity of the mapping relationship at the pixel level. It represents the minimum photometric loss of processing out-of-view pixels and occluded objects in the source image, I sIndicates the previous or next frame of the source or target image. represents the weighted edge-aware smoothness loss, which aims to ensure that the predicted depth map remains smooth in the edge area while maintaining a certain continuity in the non-edge area. represents the average normalized depth. s Represents the weight information used for weighted edge-aware smoothing loss. Represents the sum of loss functions that include multiple scale features.
[0053] S23, based on the features of different scales obtained by fusion of the second, third and fourth stages in step S21, that is, the size obtained by fusion of the input upsampling and convolution module in step S23 is H×W×C1 ~ 、H / 2×W / 2×C2 ~ 、H / 4×W / 4×C3 ~ The fused features of each size are sent to an object detection prediction head to implement the classification subtask and the box regression subtask;
[0054] The target detection prediction head includes: a classification prediction head and a box regression prediction head.
[0055] Among them, the classification prediction head includes: 2 convolution layers with C convolution kernels of size 3×3, one ReLU activation layer, one convolution layer with 2A convolution kernels of size 3×3, and Sigmoid activation function; the box regression prediction head includes: 2 convolution layers with C convolution kernels of size 3×3, one ReLU activation layer, one convolution layer with 4A convolution kernels of size 3×3, and Sigmoid activation function; the classification prediction head is used to predict the probability of the category to which A anchor boxes belong at each pixel position, and the target categories include: background targets and foreground pedestrian targets. The box regression prediction head is used to predict the offset of (center point coordinates, center point y coordinates, width w, height h1) of each anchor box.
[0056] In the classification prediction head, two convolution layers with C convolution kernels of size 3×3 are first used, and then the output features are input into the ReLU activation layer, and then the output is passed through a convolution layer with 2A convolution kernels of size 3×3. Finally, the output is passed through the Sigmoid activation function to output 2A binary prediction results corresponding to each spatial position, that is, whether it belongs to a background target or a foreground pedestrian target.
[0057] Among them, A is set to 9, and the size of C is determined according to the number of channels of the actual input features.
[0058] Similarly, the box regression prediction head outputs 4A linear results at each spatial position. For the A anchor boxes at each spatial position, these 4 outputs represent the relative offset between the predicted anchor box and the true box.
[0059] The loss function of the target detection prediction head is That is, using focal Loss for cross entropy And L1 Loss for regression tasks To complete the training of the prediction head, the expression is as follows:
[0060]
[0061] Among them, p t Represents the predicted probability. When the sample is a positive sample, p t =p, when it is a negative sample, it is p t =1-p, p represents the direct output probability of the prediction head, α t y represents the balance factor, which is used to balance the impact of positive and negative samples. γ is used to adjust the difficulty of easy samples. When γ>0, the weight of simple samples is reduced and the attention to difficult samples is increased. i represents the true value, and Represents the predicted value, and N represents the number of samples. cls , They represent the weight information of two loss functions, which are set to 0.45 and 0.55 respectively.
[0062] Furthermore, the step S4 is specifically as follows:
[0063] For the detection frame obtained from the image at the current time t, OCR Mean-shift clustering is performed to classify the dense occlusion and non-dense occlusion in the detection frame. At the same time, OCR Mean-shift clustering is continued for the trajectory frame of the image at time t-1 to classify the dense occlusion and non-dense occlusion in the trajectory frame.
[0064] Among them, the trajectory information at the t-1th moment is known data, that is, each target in the image at the t-1th moment is represented by a bounding box and a unique ID number, the box with an ID is the trajectory box, and the three-dimensional depth information of the image at the t-1th moment is known.
[0065] For all target bounding boxes detected at each moment, the center point coordinates are taken as the sample points to be clustered. The calculation expression is as follows:
[0066]
[0067] w(x i )=λ1W1+λ2W2
[0068]
[0069] Among them, W1 and W2 represent calculation factors, which are used to measure the characteristics of whether the target itself is occluded and the connection between targets in the occlusion group; x i represents a sample point to be clustered, Represents the Euclidean distance sample point x i The nearest sample point, Represents the category probability directly output by the classification prediction head, Represents x i Sample points and The IOU value between the sample points and the corresponding bbox, express The area of the bbox corresponding to the sample point, w(x i ) represents the weight of the sample point after improvement, λ1 and λ2 represent the weight factors of the two factors, which are fixed to λ1 = 0.7 and λ2 = 0.3, and K represents the Gaussian kernel function, h is the bandwidth, which controls the range of the kernel function, and M h (x) is used to complete mean-shift clustering and update the drift point x position.
[0070] Furthermore, the step S5 is specifically as follows:
[0071] S51, realizing data association for non-densely occluded targets;
[0072] Based on the non-dense occluded targets obtained in step S4, the basic association method of SORT is directly adopted. The IOU similarity is calculated for the non-dense occluded targets of the detection box at time t and the trajectory box at time t-1. The Hungarian algorithm is used to associate the non-dense occluded targets based on the similarity matrix, and the associated IOU threshold is set to 0.5. When the IOU similarity is less than 0.5, it is considered that no association will occur.
[0073] S52, realizing data association for densely occluded targets;
[0074] Based on the multiple dense occlusions obtained in step S4, the centroids of the dense occlusions of the detection frame and the trajectory frame are calculated, and the dense occlusions are matched through the Euclidean distance of the multiple centroids of the two frames. That is, a dense occlusion of the detection frame and a dense occlusion of the trajectory frame are associated with the data within the occlusion. And the association strategy is based on the three-dimensional depth information at time t obtained in step S3.
[0075] Using the 3D depth information of the target in the dense occlusion, the depth level is evenly divided in the dense occlusion. That is, according to the depth information of each target in the dense occlusion, the deepest depth and the shallowest depth are found. The deepest depth and the shallowest depth are evenly divided into 4 levels of association areas. And for the two dense occlusions of the detection box at time t and the trajectory box at time t-1, only the targets of the same level can achieve data association by calculating IOU.
[0076] Then, according to the hierarchical association strategy, the targets are associated layer by layer. The unassociated targets will be transferred to the next layer for the same IOU association until four cascade associations are completed and an identity ID is assigned to the matched targets.
[0077] Furthermore, the step S6 is specifically as follows:
[0078] For the targets that have not been associated in step S5, the last data association is performed to establish a global IOU similarity matrix for the targets that have not been associated in the detection frame at time t and the trajectory frame at time t-1. Then, unlike the cascade matching in step S6, this association is performed once through the Hungarian algorithm to obtain the matching result, and the IOU calculation threshold is still set to 0.5. The associated detection frame targets are assigned the identity ID of the associate, and the unassociated detection frame IDs are assigned new identity IDs. For the unassociated trajectory frame targets, these IDs are discarded.
[0079] Finally, the target trajectory is obtained by tracking the target frame by frame, realizing multi-target tracking based on depth information decoupling.
[0080] Beneficial effects of the present invention: The method of the present invention first constructs a multi-target tracking model based on three-dimensional depth information decoupling and trains it, then clusters the two-dimensional target detection box to complete the detection of dense occlusion, and then obtains the three-dimensional depth information of the target through a self-supervised depth estimation method, and uses the depth information to decouple the dense occlusion to complete multi-target tracking based on spatial information. The method of the present invention does not need to solve the problem of inaccurate extraction of target appearance feature information caused by dense occlusion. It proposes that MOT based on depth information only uses IOU matching to achieve association, solves the occlusion problem through the advantages of depth information, does not require appearance features, and proposes an algorithm based on three-dimensional depth estimation. Based on the difference in depth information, hierarchical decoupling of densely occluded targets is achieved, and data association is performed at different levels. Targets at different levels cannot affect each other, and a robust multi-target tracking effect is obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 This is a flow chart of a multi-target tracking method based on three-dimensional depth information decoupling of the present invention.
[0082] Figure 2 This is a flow chart based on target detection and deep structure feature fusion in an embodiment of the present invention.
[0083] Figure 3 This is a diagram showing the effect of depth information extraction in an embodiment of the present invention.
[0084] Figure 4 Schematic diagram of clustering of the OCR Mean-shift algorithm in an embodiment of the present invention.
[0085] Figure 5 This is a clustering effect diagram of the OCR Mean-shift algorithm in an embodiment of the present invention.
[0086] Figure 6 This is a pseudo code algorithm flow chart of the OCR Mean-Shift algorithm in an embodiment of the present invention.
[0087] Figure 7 This is a diagram of the overall multi-target tracking effect in an embodiment of the present invention. DETAILED DESCRIPTION
[0088] The method of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0089] like Figure 1 As shown, a flow chart of a multi-target tracking method based on three-dimensional depth information decoupling of the present invention, the specific steps are as follows:
[0090] S1. Build a multi-target tracking model based on three-dimensional depth information decoupling;
[0091] The multi-target tracking model includes: 1 feature extraction network based on attention mechanism, 1 camera posture estimation network, 3 depth estimation prediction heads, and 3 target detection prediction heads.
[0092] S2, training the multi-target tracking model constructed in step S1;
[0093] S3, based on the multi-target tracking model trained in step S2, input the image at the current time t, and obtain the target detection data of the image at the time t and the three-dimensional depth information at the time t;
[0094] Each detected target is represented by a bounding box without ID information. The box without ID is the detection box. The target detection prediction head outputs multiple pedestrian target boxes for each input image. Three-dimensional depth information can help achieve decoupling.
[0095] S4, based on step S3, clustering multiple detection frames at time t and trajectory frames at time t-1 is performed using an improved Mean-shift clustering algorithm (OCR Mean-shift algorithm) to identify dense occlusions, and each class is a dense occlusion;
[0096] Among them, when the number of targets in each clustering result exceeds 5, it is regarded as dense occlusion; for "dense occlusion" with less than 5 targets, they are regarded as targets with less severe occlusion or even no occlusion, that is, non-dense occlusion targets.
[0097] S5, based on step S4, realizing data association between non-dense occlusion targets and dense occlusion targets;
[0098] S6. Process the unmatched detection frame targets and trajectory frame targets in step S5, track the targets frame by frame to obtain target trajectories, and implement multi-target tracking based on depth information decoupling.
[0099] like Figure 2 As shown, the process based on target detection and deep structure feature fusion, that is, the multi-target tracking model constructed in the training step S1.
[0100] In this embodiment, step S2 is specifically as follows:
[0101] S21, extracting deep feature information through a feature extraction network based on the attention mechanism;
[0102] The feature extraction network is a U-shaped backbone network, including: a Conv-stem convolution module, an average pooling layer, a stage 2 convolution layer, a stage 3 convolution layer, a stage 4 convolution layer, and three upsampling and convolution modules.
[0103] Among them, the convolution kernel size of the stage 1-4 convolution layers is 3×3, and the step size is 2; the stage 2 convolution layer and the stage 3 convolution layer both include: 1 downsampling layer, 3 continuous expansion convolution modules CDC, and 1 local-global feature interaction module LGFI; the stage 4 convolution layer includes: 1 downsampling layer, 6 continuous expansion convolution modules CDC, and 1 local-global feature interaction module LGFI. The upsampling and convolution module includes: 1 upsampling layer and 1 convolution layer.
[0104] The feature extraction network is used to extract deep feature information, that is, the encoder of like-UNet is used to extract deep feature information, which is divided into four stages, as follows:
[0105] 1) Phase 1:
[0106] The original image of size H×W×3 is used as the input feature and input into the Conv-stem convolution module for downsampling to obtain a size of H / 2×W / 2×C1 tmp1 The output feature map of .
[0107] Among them, H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image; the Conv-stem module consists of three convolutional layers. The first layer uses a 3×3 convolution kernel with a step size of 2 to achieve convolution to achieve the effect of downsampling, while the next two layers use a 3×3 convolution kernel with a step size of 1 to achieve convolution to achieve the effect of local feature extraction.
[0108] 2) Second stage:
[0109] The original image is processed by the average pooling layer to obtain a feature map of size H / 2×W / 2×3, which is different from the feature map of size H / 2×W / 2×C1 obtained in the first stage. tmp1 The feature maps of the image are concatenated to obtain a feature map of size H / 2×W / 2×C1 as the input feature, and then input into the convolution layer with a convolution kernel size of 3×3 and a step size of 2 for downsampling to obtain a feature map of size H / 4×W / 4×C2 tmp1 The downsampled features are then passed through three consecutive dilated convolution modules CDC and one local-global feature interaction module LGFI to obtain a size of H / 4×W / 4×C2 tmp2 The output feature map of .
[0110] Among them, the window size of the average pooling layer in this stage is set to H / 2×H / 2, and the step size is set to H / 2, that is, the original image with a size of H×W×3 passes through this average pooling layer, and the output size is H / 2×W / 2×3.
[0111] The reason for concatenating the output features of the first stage with the input image features after average pooling is that this can reduce the spatial information loss caused by the reduction of feature size, which will also be used in the feature processing of the subsequent stage. The CDC and LGFI modules refer to the implementation of Lite-Mono. The former uses dilated convolution to extract multi-scale local features, while the latter uses the cross-covariance attention mechanism to focus on the feature changes on the channel. These two modules can be used to learn rich hierarchical feature information.
[0112] 3) The third stage:
[0113] The size obtained in the second stage is H / 4×W / 4×C2 tmp1 The size of the initial down-sampling feature map and the final output of the stage is H / 4×W / 4×C2 tmp2 The feature map of the original image after the average pooling layer is concatenated to obtain a feature map of size H / 4×W / 4×C2 as the input feature, which is then input into a convolution layer with a convolution kernel size of 3×3 and a stride of 2 for downsampling to obtain a feature map of size H / 8×W / 8×C3. tmp1 The downsampled feature map is then passed through three CDC modules and one LGFI module to obtain a size of H / 8×W / 8×C3 tmp2 The output feature map of .
[0114] Among them, the average pooling window size in this stage is set to H / 4×H / 4, the step size is set to H / 4, and the output size of the average pooling layer is H / 4×W / 4×3.
[0115] 4) The fourth stage:
[0116] The size obtained in the third stage is H / 8×W / 8×C3 tmp1 The size of the initial down-sampling feature map and the final output of the stage is H / 8×W / 8×C3 tmp2 The feature map of the original image and the feature map of size H / 8×W / 8×3 after the average pooling layer are concatenated to obtain a feature map of size H / 8×W / 8×C3 as the input feature, which is then input into a convolution layer with a convolution kernel size of 3×3 and a stride of 2 for downsampling, and the sampling result is input into 6 CDC modules and 1 LGFI module to obtain an output feature map of size H / 16×W / 16×C4.
[0117] Among them, the average pooling window size in this stage is set to H / 8×H / 8, the step size is set to H / 8, and the output size of the average pooling layer is H / 8×W / 8×3.
[0118] In the above description, the input of the third and fourth stages needs to concatenate the features of the output of the downsampling layer of the previous stage. This design is similar to the residual connection proposed in ResNet and can better model the correlation of features across stages.
[0119] Then, bilinear upsampling is used to increase the spatial dimension of the feature maps of the second, third, and fourth stages, and convolutional layers are used to connect the features of the three stages of the encoder, that is, a depth estimation prediction head and an object detection prediction head are followed in each upsampling and convolution module. The depth estimation prediction head will output inverse depth map features at 1, 1 / 2, and 1 / 4 of the original image resolution, respectively, and this part of the features will be combined with the camera pose estimation information to obtain the final depth estimation result.
[0120] S22, fusing the multi-scale feature information extracted in step S21 through the depth estimation prediction head to obtain depth estimation information of different scales, and then fusing and reconstructing the multi-scale depth estimation information through the camera pose estimation completed by the camera pose estimation network Pose Net to obtain the final three-dimensional depth information (depth estimation result), thereby realizing the reconstruction of the three-dimensional depth feature information;
[0121] S221, by fusing the features of different stages obtained in step S21, and then obtaining depth estimation information of three different scales through a depth estimation prediction head, that is, multi-scale depth estimation information;
[0122] (1) The features of size H / 16×W / 16×C4 obtained in the fourth stage are input into the first upsampling and convolution module for upsampling and residual connection. The size of the third stage is H / 8×W / 8×C3 tmp2 The features are then passed through a convolutional layer to obtain a size of H / 8×W / 8×C3 ~The fused features are input into the depth estimation prediction head, and finally the depth estimation information of size H / 4×W / 4×1 is obtained.
[0123] Among them, the depth estimation prediction head includes: a convolution layer Conv, an upsampling layer and an activation layer Sigmoid.
[0124] (2) The size is H / 8×W / 8×C3 ~ The fusion features are input into the second upsampling and convolution module for upsampling and residual connection. The size of the second stage is H / 4×W / 4×C2 tmp2 The features are then passed through a convolutional layer to obtain a size of H / 4×W / 4×C2 ~ The fused features are input into another depth estimation prediction head with the same structure, and finally the depth estimation information of size H / 2×W / 2×1 is obtained.
[0125] (3) The size is H / 4×W / 4×C2 ~ The fused features are input into the third upsampling and convolution module for upsampling, and then directly pass through the convolution layer to obtain the feature H / 2×W / 2×C1 ~ , and input it into the depth estimation prediction head of the same structure, and finally obtain the depth estimation information of size H×W×1.
[0126] S222, obtaining camera pose estimation information of two adjacent frames, that is, completing camera pose estimation through a camera pose estimation network Pose Net;
[0127] For monocular camera training, the camera pose estimation network is formed by ResNet18, with a pair of color images or six channels as input, and a four-layer convolutional pose decoder is used to estimate the corresponding 6 degrees of freedom relative pose between two adjacent frames. At the same time, horizontal inversion and random brightness, contrast, saturation and hue jitter with a range of ±0.2, ±0.2, ±0.2 and ±0.1 are trained to ensure the robustness of camera pose estimation.
[0128] The ResNet18 network is pre-trained; the six degrees of freedom are the translational freedom and rotational freedom of the x-axis, y-axis, and z-axis in the three-dimensional coordinate system. Through these six degrees of freedom, the camera's posture can be represented by a 4x4 homogeneous transformation matrix T, that is, a 3x3 rotation matrix R and a 3x1 translation vector z. This camera posture matrix can represent the posture changes of the camera between different frames.
[0129] S223, correcting the original rough three-dimensional depth information obtained in step S21, that is, based on step S222, fusing and reconstructing the multi-scale depth estimation information obtained in step S221 to obtain the final three-dimensional depth information, and reconstructing the three-dimensional depth feature information through projection change and linear interpolation;
[0130] The first is the projection transformation, which projects each pixel in the depth map into the world coordinate system, and then projects it to the image plane of the target perspective according to the camera posture transformation matrix.
[0131] For each pixel point (u, v), create a pixel coordinate grid, convert the pixel coordinates to normalized coordinates, and then multiply the normalized coordinates by the depth estimation information d of the corresponding pixel point to obtain the 3D point in the world coordinate system, and then use the camera's attitude transformation matrix T to project the 3D point from the current view to the target view, and use the camera's intrinsic parameter matrix K (known) of the target view to project the 3D point back to the image plane, and normalize the projected coordinates to the image plane, that is, from three-dimensional coordinates (x, y, z) to two-dimensional coordinates (x / z, y / z, 1).
[0132] Then bilinear interpolation is performed to map the pixel value of the projection point to the target image on the image plane of the target perspective.
[0133] Normalize the projection coordinates to the range [-1,1] of the image plane, and use the bilinear interpolation method to interpolate the pixel values corresponding to the normalized projection coordinates from the image of the target perspective. Finally, the depth estimation information corrected by the camera posture estimation is obtained, that is, the final three-dimensional depth estimation map. The calculation expression is as follows:
[0134] Interpolation pixel value = (1-α)(1-β)P 00 +α(1-β)P 10 +(1-α)βP 01 +αβP 11
[0135] Among them, α and β represent the horizontal and vertical fractional parts of the normalized projection coordinates, respectively, and P ij Represents the four most recent pixel values.
[0136] The uncorrected depth estimation information used as the initial input in the above description has different resolutions. Therefore, for each resolution of the depth estimation information, it is necessary to first upsample to the same resolution as the original image, and then go through the described steps to reconstruct the final 3D depth estimation map of each scale. During inference, only the final 3D depth estimation corresponding to the depth estimation information with a resolution of H×W×1 is used as the final 3D depth estimation result of the input original image. The 3D depth estimation of the other two resolutions can effectively constrain the depth maps of each scale to work towards the same goal only when training the network, that is, to reconstruct the high-resolution input target image as accurately as possible, so as to achieve the effect of improving the convergence speed of the network.
[0137] The depth estimation prediction head learning objective modeling is to minimize the target image I t And the synthetic target image The image reconstruction loss between and an edge-aware smoothness loss constrained on the predicted depth map
[0138] The depth estimation prediction head loss function expression is as follows:
[0139]
[0140]
[0141] in, represents the image reconstruction loss, I t represents the target image, represents the synthetic target image, m represents the fixed threshold of 0.85, SSIM(*) represents the structural similarity index between pixels, and ∥*∥ represents the L1 similarity characteristic between pixels, which is used to characterize the similarity of the mapping relationship at the pixel level. It represents the minimum photometric loss of processing out-of-view pixels and occluded objects in the source image, I s Indicates the previous or next frame of the source or target image. represents the weighted edge-aware smoothness loss, which aims to ensure that the predicted depth map remains smooth in the edge area while maintaining a certain continuity in the non-edge area. represents the average normalized depth. s Represents the weight information used for weighted edge-aware smoothing loss, which is set to 0.001. represents the sum of the loss functions containing multiple scale features. The depth information extraction effect diagram finally obtained in this embodiment is as follows Figure 3 shown.
[0142] S23, based on the features of different scales obtained by fusion of the second, third and fourth stages in step S21, that is, the size obtained by fusion of the input upsampling and convolution module in step S23 is H×W×C1 ~ 、H / 2×W / 2×C2 ~ 、H / 4×W / 4×C3 ~ The fused features of each size are sent to an object detection prediction head to implement the classification subtask and the box regression subtask;
[0143] The target detection prediction head includes: a classification prediction head and a box regression prediction head.
[0144] Among them, the classification prediction head includes: 2 convolution layers with C convolution kernels of size 3×3, one ReLU activation layer, one convolution layer with 2A convolution kernels of size 3×3, and Sigmoid activation function; the box regression prediction head includes: 2 convolution layers with C convolution kernels of size 3×3, one ReLU activation layer, one convolution layer with 4A convolution kernels of size 3×3, and Sigmoid activation function; the classification prediction head is used to predict the probability of the category to which A anchor boxes belong at each pixel position, and the target categories include: background targets and foreground pedestrian targets. The box regression prediction head is used to predict the offset of (center point coordinates, center point y coordinates, width w, height h1) of each anchor box.
[0145] In the classification prediction head, two convolution layers with C convolution kernels of size 3×3 are first used, and then the output features are input into the ReLU activation layer, and then the output is passed through a convolution layer with 2A convolution kernels of size 3×3. Finally, the output is passed through the Sigmoid activation function to output 2A binary prediction results corresponding to each spatial position, that is, whether it belongs to a background target or a foreground pedestrian target.
[0146] Among them, A is set to 9, and the size of C is determined according to the number of channels of the actual input features.
[0147] Similarly, the box regression prediction head outputs 4A linear results at each spatial position. For the A anchor boxes at each spatial position, these 4 outputs represent the relative offset between the predicted anchor box and the true box (ground truth real data).
[0148] The box regression prediction head uses a class-independent bounding box regressor that uses fewer parameters but performs equally well. Although the object classification prediction head and the box regression prediction head share similar structures, they use independent parameters.
[0149] The loss function of the target detection prediction head is That is, using focal Loss for cross entropy And L1 Loss for regression tasks To complete the training of the prediction head, the specific expression is as follows:
[0150]
[0151] Among them, p t Represents the predicted probability. When the sample is a positive sample, p t =p, when it is a negative sample, it is p t =1-p, p represents the direct output probability of the prediction head, α t y represents the balance factor, which is used to balance the impact of positive and negative samples. γ is used to adjust the difficulty of easy samples. When γ>0, the weight of simple samples is reduced and the attention to difficult samples is increased. i represents the true value, and Represents the predicted value, and N represents the number of samples. cls , They represent the weight information of two loss functions, which are set to 0.45 and 0.55 respectively.
[0152] In this embodiment, step S4 is specifically as follows:
[0153] A new weight calculation method is introduced to implement a mean-shift clustering algorithm for densely occluded targets, the OCR mean-shift algorithm. The weight coefficient reflects the weight distribution of global sample points, but for an occluded target, the main factor affecting its offset calculation is the target that occludes it, so this global weight coefficient often ignores the more important sample weights.
[0154] For the detection frame obtained from the image at the current time t, OCR Mean-shift clustering is performed to classify the dense occlusion and non-dense occlusion in the detection frame. At the same time, OCR Mean-shift clustering is continued for the trajectory frame of the image at time t-1 to classify the dense occlusion and non-dense occlusion in the trajectory frame.
[0155] Among them, the trajectory information at the t-1th moment is known data, that is, each target in the image at the t-1th moment is represented by a bounding box and a unique ID number, the box with an ID is the trajectory box, and the three-dimensional depth information of the image at the t-1th moment is known.
[0156] The clustering diagram and clustering effect diagram of the OCR Mean-shift algorithm are as follows: Figure 4 , Figure 5 For all target bounding boxes detected at each moment, the center point coordinates are taken as the sample points to be clustered and the calculation expression is as follows:
[0157]
[0158] w(x i )=λ1W1+λ2W2
[0159]
[0160] Among them, W1 and W2 represent calculation factors, which are used to measure the characteristics of whether the target itself is occluded and the connection between targets in the occlusion group; x i Represents a sample point to be clustered (the coordinates of the center point of each box), Represents the Euclidean distance sample point x i The nearest sample point, Represents the category probability directly output by the classification prediction head, Represents x i Sample points and The IOU value between the sample points and the corresponding bbox, express The area of the sample point corresponding to the bbox, w(x i ) represents the weight of the sample point after improvement, λ1 and λ2 represent the weight factors of the two factors, which are fixed to λ1=0.7 and λ2=0.3, while K represents the Gaussian kernel function, h represents the bandwidth, which controls the range of the kernel function, and M h (x) is used to complete mean-shift clustering and update the drift point x position.
[0161] W1 mainly measures the detection effect of the target itself. Generally speaking, the detection difficulty of occluded targets is relatively high, so the classification confidence of the target box is often low. Although it may be caused by the detection challenges such as the size of the target itself and lighting changes, occlusion is also a very important reason, so this type of target is still regarded as a candidate target of the dense occlusion group, and W1 is used to reflect the value of the weight coefficient. W2 mainly measures the weight influence between the target and the target that occludes it. The reason is the difference in the size of the target in the picture. The coordinates of the sample point cannot fully reflect the occlusion situation. The area covered by the target box and the overlap with other targets are the factors that directly reflect the occlusion situation.
[0162] The improvement of OCR Mean-shift algorithm is reflected in M h (x) For the weight calculation of sample points, the overall algorithm flow is still implemented according to the mean-shift algorithm flow. The specific pseudo-code algorithm flow is as follows: Figure 6 shown.
[0163] In this embodiment, step S5 is specifically as follows:
[0164] S51, realizing data association for non-densely occluded targets;
[0165] Based on the non-dense occluded targets obtained in step S4, the basic association method of SORT is directly used. The IOU similarity is calculated for the non-dense occluded targets of the detection box at time t and the trajectory box at time t-1. The association of non-dense occluded targets is achieved using the Hungarian algorithm based on the similarity matrix. If the association of this part of the targets is relatively accurate, the associated IOU threshold is set to 0.5. When the IOU similarity is less than 0.5, it is considered that no association will occur.
[0166] Common multi-target tracking methods will achieve the identity division of the target by introducing an appearance feature extraction module. On the one hand, the extraction of such appearance features is relatively difficult, because it is necessary to additionally design a loss function to train the appearance feature extraction network, and combine multiple Re-ID training sets to extract sufficiently rich appearance features. On the other hand, this appearance feature is easily affected by other targets in the face of dense occlusion. The appearance features between targets are likely to overlap and be chaotic due to occlusion. The MOT based on depth information in this embodiment only uses IOU matching to achieve association. It uses the advantages of depth information to solve the occlusion problem without the need for appearance features.
[0167] The existing method does not consider the processing of dense occlusion, but in this embodiment, it will be processed. Because dense occlusion does not appear all the time, nor is it global, it is often short-term and local, so the mean-shift algorithm that calculates the weight of the detection box sample points based on the detection box confidence and the same-frame IOU can well realize the recognition of dense occlusion and facilitate the processing of dense occlusion.
[0168] S52, realizing data association for densely occluded targets;
[0169] For the processing of dense occlusion, it is evenly divided through depth information, so that the mutually occluded targets are decoupled at different levels. By adopting step-by-step data association, targets at different levels will not interfere with each other, thereby improving the tracking effect of densely occluded targets and reducing IDswitch.
[0170] Based on the multiple dense occlusions obtained in step S4, the centroids of the dense occlusions of the detection frame and the trajectory frame are calculated, and the dense occlusions are matched through the Euclidean distance of the multiple centroids of the two frames. That is, a dense occlusion of the detection frame and a dense occlusion of the trajectory frame are associated with the data within the occlusion. And the association strategy is based on the three-dimensional depth information at time t obtained in step S3.
[0171] The three-dimensional depth information of the target in the dense occlusion is used to achieve uniform depth hierarchy within the dense occlusion. That is, according to the depth information of each target in the dense occlusion (the three-dimensional depth corresponding to the pixel at the bottom center of the target detection frame), the deepest depth and the shallowest depth are found. Because in dense occlusion, the target is often not completely random like the global target distribution, but is relatively uniform, it is evenly divided into 4 levels of associated areas between the deepest depth and the shallowest depth. For the two dense occlusions of the detection frame at time t and the trajectory frame at time t-1, only targets of the same level can achieve data association by calculating IOU.
[0172] Then, according to the hierarchical association strategy, the targets are associated layer by layer. The unassociated targets will be transferred to the next layer for the same IOU association until four cascade associations are completed and an identity ID is assigned to the matched targets.
[0173] In this embodiment, step S6 is specifically as follows:
[0174] For the targets that have not been associated in steps S5 and S6, the last data association is performed to establish a global IOU similarity matrix for the targets that have not been associated in the detection frame at time t and the trajectory frame at time t-1. Then, unlike the cascade matching in step S6, this association is performed once through the Hungarian algorithm to obtain the matching result, and the IOU calculation threshold is still set to 0.5. The associated detection frame targets are assigned the identity ID of the associate, and the unassociated detection frame IDs are assigned new identity IDs. For the unassociated trajectory frame targets, this part of the ID is discarded.
[0175] Finally, the target trajectory is obtained by tracking the target frame by frame, realizing multi-target tracking based on depth information decoupling. The final tracking effect diagram of this embodiment is shown in FIG. Figure 7 shown.
[0176] In summary, the method of the present invention does not need to solve the problem of inaccurate extraction of target appearance feature information caused by dense occlusion. It proposes that MOT based on depth information only uses IOU matching to achieve association, and solves the occlusion problem by taking advantage of depth information. It does not require appearance features. It proposes an algorithm based on three-dimensional depth estimation (a multi-level data association strategy based on three-dimensional depth information), which realizes hierarchical decoupling of densely occluded targets based on the difference in depth information, performs data association in different levels, and targets at different levels cannot affect each other, thereby achieving efficient multi-target tracking and obtaining a robust multi-target tracking effect.
[0177] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A multi-target tracking method based on three-dimensional depth information decoupling, the specific steps are as follows: S1. Build a multi-target tracking model based on three-dimensional depth information decoupling; The multi-target tracking model includes: 1 feature extraction network based on attention mechanism, 1 camera pose estimation network, 3 depth estimation prediction heads, 3 object detection prediction heads; S2, training the multi-target tracking model constructed in step S1; S3, based on the multi-target tracking model trained in step S2, input the image at the current time t, and obtain the target detection data of the image at the time t and the three-dimensional depth information at the time t; Among them, each detected target is represented by a bounding box without ID information. The box without ID is the detection box; the target detection prediction head outputs multiple pedestrian target boxes for each input image; S4, based on step S3, clustering multiple detection frames at time t and trajectory frames at time t-1 is performed using an improved Mean-shift clustering algorithm to identify dense occlusions, and each class is a dense occlusion; When the number of targets in each clustering result exceeds 5, it is regarded as dense occlusion; for "dense occlusion" with less than 5 targets, it is regarded as targets with less severe occlusion or even no occlusion, that is, non-dense occlusion targets; S5, based on step S4, realizing data association between non-dense occlusion targets and dense occlusion targets; S6. Process the unmatched detection frame targets and trajectory frame targets in step S5, track the targets frame by frame to obtain target trajectories, and implement multi-target tracking based on depth information decoupling.
2. The multi-target tracking method based on three-dimensional depth information decoupling according to claim 1, characterized in that: The step S2 is specifically as follows: S21, extracting deep feature information through a feature extraction network based on the attention mechanism; The feature extraction network is a U-shaped backbone network, including: a Conv-stem convolution module, an average pooling layer, a stage 2 convolution layer, a stage 3 convolution layer, a stage 4 convolution layer, and three upsampling and convolution modules; Among them, the convolution kernel size of the stage 1-4 convolution layer is 3×3, and the step size is 2; the stage 2 convolution layer and the stage 3 convolution layer both include: 1 downsampling layer, 3 continuous expansion convolution modules CDC, and 1 local-global feature interaction module LGFI; the stage 4 convolution layer includes: 1 downsampling layer, 6 continuous expansion convolution modules CDC, and 1 local-global feature interaction module LGFI; the upsampling and convolution module includes: 1 upsampling layer and 1 convolution layer; The feature extraction network is used to extract deep feature information, that is, the encoder of like-UNet is used to extract deep feature information, which is divided into four stages, as follows: 1) Phase 1: The original image of size H×W×3 is used as the input feature and input into the Conv-stem convolution module for downsampling to obtain a size of H / 2×W / 2×C1 tmp1 The output feature map of Among them, H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image; the Conv-stem module consists of three convolutional layers. The first layer uses a 3×3 convolution kernel with a step size of 2 to achieve the effect of downsampling, while the next two layers use a 3×3 convolution kernel with a step size of 1 to achieve convolution, so as to achieve the effect of local feature extraction; 2) Second stage: The original image is processed by the average pooling layer to obtain a feature map of size H / 2×W / 2×3, which is the same as the feature map of size H / 2×W / 2×C1 obtained in the first stage. tmp1 The feature maps of the image are concatenated to obtain a feature map of size H / 2×W / 2×C1 as the input feature, and then input into the convolution layer with a convolution kernel size of 3×3 and a step size of 2 for downsampling to obtain a feature map of size H / 4×W / 4×C2 tmp1 The downsampled features are then passed through three consecutive dilated convolution modules CDC and one local-global feature interaction module LGFI to obtain a size of H / 4×W / 4×C2 tmp2 The output feature map of Among them, the window size of the average pooling layer in this stage is set to H / 2×H / 2, and the step size is set to H / 2, that is, the original image with a size of H×W×3 passes through this average pooling layer, and the output size is H / 2×W / 2×3; 3) The third stage: The size obtained in the second stage is H / 4×W / 4×C2 tmp1 The size of the initial down-sampling feature map and the final output of the stage is H / 4×W / 4×C2 tmp2 The feature map of the original image after the average pooling layer is concatenated to obtain a feature map of size H / 4×W / 4×C2 as the input feature, which is then input into a convolution layer with a convolution kernel size of 3×3 and a stride of 2 for downsampling to obtain a feature map of size H / 8×W / 8×C3. tmp1 The downsampled feature map is then passed through three CDC modules and one LGFI module to obtain a size of H / 8×W / 8×C3 tmp2 The output feature map of Among them, the average pooling window size in this stage is set to H / 4×H / 4, the step size is set to H / 4, and the output size of the average pooling layer is H / 4×W / 4×3; 4) The fourth stage: The size obtained in the third stage is H / 8×W / 8×C3 tmp1 The size of the initial down-sampling feature map and the final output of the stage is H / 8×W / 8×C3 tmp2 The feature map of the original image and the feature map of size H / 8×W / 8×3 after the average pooling layer are concatenated to obtain a feature map of size H / 8×W / 8×C3 as the input feature, which is then input into a convolution layer with a convolution kernel size of 3×3 and a stride of 2 for downsampling. The sampling result is input into 6 CDC modules and 1 LGFI module to obtain an output feature map of size H / 16×W / 16×C4; Among them, the average pooling window size in this stage is set to H / 8×H / 8, the step size is set to H / 8, and the output size of the average pooling layer is H / 8×W / 8×3; S22, fusing the multi-scale feature information extracted in step S21 through the depth estimation prediction head to obtain depth estimation information of different scales, and then fusing and reconstructing the multi-scale depth estimation information through the camera pose estimation completed by the camera pose estimation network Pose Net to obtain the final three-dimensional depth information, thereby realizing the reconstruction of the three-dimensional depth feature information; S221, by fusing the features of different stages obtained in step S21, and then obtaining depth estimation information of three different scales through a depth estimation prediction head, that is, multi-scale depth estimation information; (1) The features of size H / 16×W / 16×C4 obtained in the fourth stage are input into the first upsampling and convolution module for upsampling and residual connection. The size of the third stage is H / 8×W / 8×C3 tmp2 The features are then passed through a convolutional layer to obtain a size of H / 8×W / 8×C3 ~ The fused features are input into the depth estimation prediction head, and finally the depth estimation information of size H / 4×W / 4×1 is obtained; Among them, the depth estimation prediction head includes: a convolution layer Conv, an upsampling layer and an activation layer Sigmoid; (2) The size is H / 8×W / 8×C3 ~ The fusion features are input into the second upsampling and convolution module for upsampling and residual connection. The size of the second stage is H / 4×W / 4×C2 tmp2 The features are then passed through a convolutional layer to obtain a size of H / 4×W / 4×C2 ~ The fused features are input into another depth estimation prediction head with the same structure, and finally the depth estimation information of size H / 2×W / 2×1 is obtained; (3) The size is H / 4×W / 4×C2 ~ The fused features are input into the third upsampling and convolution module for upsampling, and then directly pass through the convolution layer to obtain the feature H / 2×W / 2×C1 ~ , and input it into the depth estimation prediction head of the same structure, and finally obtain the depth estimation information of size H×W×1; S222, obtaining camera pose estimation information of two adjacent frames, that is, completing camera pose estimation through a camera pose estimation network Pose Net; For monocular camera training, the camera pose estimation network is formed by ResNet18, with a pair of color images or six channels as input, and a four-layer convolutional pose decoder is used to estimate the corresponding 6-DOF relative pose between two adjacent frames; horizontal inversion and random brightness, contrast, saturation and hue jitter with a range of ±0.2, ±0.2, ±0.2 and ±0.1 are also trained; The six degrees of freedom are the translational degrees of freedom and the rotational degrees of freedom of the x-axis, y-axis, and z-axis in the three-dimensional coordinate system. Through these six degrees of freedom, the camera posture is represented by a 4x4 homogeneous transformation matrix T, that is, a 3x3 rotation matrix R and a 3x1 translation vector z; S223, correcting the original rough three-dimensional depth information obtained in step S21, that is, based on step S222, fusing and reconstructing the multi-scale depth estimation information obtained in step S221 to obtain the final three-dimensional depth information, and reconstructing the three-dimensional depth feature information through projection change and linear interpolation; The first is the projection transformation, which projects each pixel in the depth map into the world coordinate system, and then projects it to the image plane of the target perspective according to the camera posture transformation matrix; For each pixel point (u, v), create a pixel coordinate grid, convert the pixel coordinates to normalized coordinates, then multiply the normalized coordinates by the depth estimation information d of the corresponding pixel point to obtain the 3D point in the world coordinate system, and then use the camera's attitude transformation matrix T to project the 3D point from the current view to the target view, and use the camera's intrinsic parameter matrix K of the target view to project the 3D point back to the image plane, and normalize the projected coordinates to the image plane, that is, from three-dimensional coordinates (x, y, z) to two-dimensional coordinates (x / z, y / z, 1); Then, bilinear interpolation is performed to map the pixel value of the projection point to the target image on the image plane of the target perspective; Normalize the projection coordinates to the range [-1,1] of the image plane, and use the bilinear interpolation method to interpolate the pixel values corresponding to the normalized projection coordinates from the image of the target perspective. Finally, the depth estimation information corrected by the camera posture estimation is obtained, that is, the final three-dimensional depth estimation map. The calculation expression is as follows: Interpolated pixel value = (1 - α)(1 - β)P 00 + α(1 - β)P 10 + (1 - α)βP 01 + αβP 11 Among them, α and β represent the horizontal and vertical fractional parts of the normalized projection coordinates, respectively, and P ij Represents the four most recent pixel values; The depth estimation prediction head learning objective modeling is to minimize the target image I t And the synthetic target image The image reconstruction loss between and an edge-aware smoothness loss constrained on the predicted depth map The depth estimation prediction head loss function expression is as follows: in, represents the image reconstruction loss, I t represents the target image, represents the synthetic target image, m represents the fixed threshold of 0.85, SSIM(*) represents the structural similarity index between pixels, and ∥*∥ represents the L1 similarity characteristic between pixels, which is used to characterize the similarity of the mapping relationship at the pixel level; It represents the minimum photometric loss of processing out-of-view pixels and occluded objects in the source image, I s Indicates the previous or next frame of the source target image; represents the weighted edge-aware smoothness loss, which aims to ensure that the predicted depth map remains smooth in the edge area while maintaining a certain continuity in the non-edge area. represents the average normalized depth; λ s Represents the weight information used for weighted edge-aware smoothing loss; Represents the sum of loss functions containing multiple scale features; S23, based on the features of different scales obtained by fusion of the second, third and fourth stages in step S21, that is, the size obtained by fusion of the input upsampling and convolution module in step S23 is H×W×C1 ~ 、H / 2×W / 2×C2 ~ 、H / 4×W / 4×C3 ~ The fused features of each size are sent to an object detection prediction head to implement the classification subtask and the box regression subtask; The target detection prediction head includes: a classification prediction head and a box regression prediction head; Among them, the classification prediction head includes: 2 convolution layers with C convolution kernels of size 3×3, one ReLU activation layer, one convolution layer with 2A convolution kernels of size 3×3, and Sigmoid activation function; the box regression prediction head includes: 2 convolution layers with C convolution kernels of size 3×3, one ReLU activation layer, one convolution layer with 4A convolution kernels of size 3×3, and Sigmoid activation function; the classification prediction head is used to predict the probability of the category to which A anchor boxes belong at each pixel position, and the target categories include: background targets and foreground pedestrian targets; and the box regression prediction head is used to predict the offset of (center point coordinates, center point y coordinates, width w, height h1) of each anchor box; In the classification prediction head, two convolutional layers with C convolutional kernels of size 3×3 are first used, and then their output features are input into the ReLU activation layer, and then the output is passed through a convolutional layer with 2A convolutional kernels of size 3×3; finally, the output is passed through the Sigmoid activation function to output 2A binary prediction results corresponding to each spatial position, that is, whether it belongs to a background target or a foreground pedestrian target; Among them, A is set to 9, and the size of C is determined according to the number of channels of the actual input features; Similarly, the box regression prediction head outputs 4A linear results at each spatial position; for the A anchor boxes at each spatial position, these 4 outputs represent the relative offset between the predicted anchor box and the true box; The loss function of the target detection prediction head is That is, using focal Loss for cross entropy And L1 Loss for regression tasks To complete the training of the prediction head, the expression is as follows: Among them, p t Represents the predicted probability. When the sample is a positive sample, p t =p, when it is a negative sample, it is p t =1-p, p represents the direct output probability of the prediction head, α t represents the balance factor, which is used to balance the impact of positive and negative samples. γ is used to adjust the difficulty of samples. When γ>0, the weight of simple samples is reduced and the attention to difficult samples is increased. i represents the true value, and represents the predicted value, N represents the number of samples; cls , They represent the weight information of two loss functions, which are set to 0.45 and 0.55 respectively.
3. The multi-target tracking method based on three-dimensional depth information decoupling according to claim 1, characterized in that: The step S4 is specifically as follows: Perform OCR Mean-shift clustering on the detection frame obtained from the image at the current time t, and classify the dense occlusion and non-dense occlusion in the detection frame. At the same time, continue to perform OCR Mean-shift clustering on the trajectory frame of the image at time t-1, and classify the dense occlusion and non-dense occlusion in the trajectory frame; Among them, the trajectory information at the t-1th moment is known data, that is, each target in the image at the t-1th moment is represented by a bounding box and a unique ID number, the box with an ID is the trajectory box, and the three-dimensional depth information of the image at the t-1th moment is known; For all target bounding boxes detected at each moment, the center point coordinates are taken as the sample points to be clustered. The calculation expression is as follows: w(x i )=λ1W1+λ2W2 Among them, W1 and W2 represent calculation factors, which are used to measure the characteristics of whether the target itself is occluded and the connection between targets in the occlusion group; x i represents a sample point to be clustered, Represents the Euclidean distance sample point x i The nearest sample point, Represents the category probability directly output by the classification prediction head, Represents x i Sample points and The IOU value between the sample points and the corresponding bbox, express The area of the sample point corresponding to the bbox, w(x i ) represents the weight of the sample point after improvement, λ1 and λ2 represent the weight factors of the two factors, which are fixed to λ1=0.7 and λ2=0.3, while K represents the Gaussian kernel function, h represents the bandwidth, which controls the range of the kernel function, and M h (x) is used to complete mean-shift clustering and update the drift point x position.
4. The multi-target tracking method based on three-dimensional depth information decoupling according to claim 1, characterized in that: The step S5 is specifically as follows: S51, realizing data association for non-densely occluded targets; Based on the non-dense occluded targets obtained in step S4, the basic association method of SORT is directly used to calculate the IOU similarity of the detection box at time t and the trajectory box at time t-1. The Hungarian algorithm is used to associate the non-dense occluded targets based on the similarity matrix, and the associated IOU threshold is set to 0.
5. When the IOU similarity is less than 0.5, it is considered that no association occurs. S52, realizing data association for densely occluded targets; Based on the multiple dense occlusions obtained in step S4, the dense occlusion centroids of the detection frame and the trajectory frame are calculated, and the dense occlusions are matched through the Euclidean distance of the multiple centroids of the two frames; that is, a dense occlusion of the detection frame and a dense occlusion of the trajectory frame are associated with the data within the occlusion; and the association strategy is based on the three-dimensional depth information at time t obtained in step S3; Using the three-dimensional depth information of the target in the dense occlusion, the depth level is evenly divided in the dense occlusion; that is, according to the depth information of each target in the dense occlusion, the deepest depth and the shallowest depth are found; the deepest depth and the shallowest depth are evenly divided into 4 levels of associated areas; and for the two dense occlusions of the detection frame at time t and the trajectory frame at time t-1, only the targets of the same level can achieve data association by calculating IOU; Then, according to the hierarchical association strategy, the targets are associated layer by layer. The unassociated targets will be transferred to the next layer for the same IOU association until four cascade associations are completed and an identity ID is assigned to the matched targets.
5. The multi-target tracking method based on three-dimensional depth information decoupling according to claim 1, characterized in that: The step S6 is specifically as follows: For the targets that have not been associated in step S5, the last data association is performed to establish a global IOU similarity matrix for the targets that have not been associated in the detection frame at time t and the trajectory frame at time t-1. Then, unlike the cascade matching in step S6, this association is obtained by the Hungarian algorithm at one time, and the IOU calculation threshold is still set to 0.5; the associated detection frame targets are assigned the identity ID of the associate, and the unassociated detection frame IDs are assigned new identity IDs. For the unassociated trajectory frame targets, these IDs are discarded; Finally, the target trajectory is obtained by tracking the target frame by frame, realizing multi-target tracking based on depth information decoupling.
Citation Information
Patent Citations
Multi-target tracking method based on coarse-to-fine shielding processing
CN113763427A
Detection and tracking integrated algorithm research based on attention mechanism and scale fusion
CN117557810A
Multi-target tracking method, system and device based on depth correlation and storage medium
CN118334098A
Visual Multi-Object Tracking based on Multi-Bernoulli Filter with YOLOv3 Detection
US20200265591A1
Cited By
Multi-target automatic tracking method and system based on unmanned intelligent turntable
CN120428215A
Foreign matter picking method and system based on cooperation of foreign matter identification and conveyor belt
CN121869729A