A Multi-Target Pedestrian Tracking Method Based on Feature Association and Feature Enhancement

By introducing feature association modules and feature enhancement technology in the multi-objective tracking method, combining Detection and ReID branches, and using the DeepSort algorithm for target tracking, the problem of difficult speed and accuracy in the existing technology is solved, and efficient and accurate multi-objective tracking for pedestrians is achieved.

CN115631213BActive Publication Date: 2025-06-27SHENYANG LIGONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210140727.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-06-27
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

The existing multi-objective tracking methods are difficult to balance between speed and accuracy, resulting in low detection accuracy and serious ID switching.

Method used

The pedestrian multi-objective tracking method based on feature association and feature enhancement is adopted. The feature map is self-attention and mutual attention enhancement through the feature association module. Combined with the Detection branch and the ReID branch with feature enhancement, the target tracking is used using the DeepSort algorithm.

Benefits of technology

It effectively improves detection accuracy, reduces the number of ID switching, and can detect and track pedestrians more accurately, suitable for inference at video rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631213B_ABST
    Figure CN115631213B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-object pedestrian tracking method based on feature association and feature enhancement, which relates to the technical field of multi-object tracking. This method uses DLA34 as the target detection backbone network, adds a feature association module based on self-attention and mutual attention, combines the Detection branch with the ReID branch enhanced in space and channels to obtain the target detection results, and matches the detected targets with the trajectories through the DeepSort algorithm, aiming to effectively detect and distinguish pedestrians appearing in the image, improve the detection accuracy, and reduce the number of ID switches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-object tracking, and particularly to a pedestrian multi-object tracking method based on feature association and feature enhancement. Background Art

[0002] With the gradual development of modern social science and technology, monitoring devices such as cameras are gradually playing a crucial role in many places, and pedestrian multi-object tracking from video surveillance has become a hot topic of current research. Multi-object tracking identifies and tracks targets through technologies such as image processing, computer vision-related algorithms, and machine learning. The technical implementation means can be simply explained as determining whether the input picture or video frame contains a target. If there is a target, the position information and identity identifier are marked. Multi-object tracking is a technical prerequisite for behavior analysis and abnormal posture detection, and is widely used in intelligent transportation, people counting, public safety, intelligent monitoring, autonomous driving, etc., and has very important application value and practical significance.

[0003] Currently, the algorithms in the field of deep learning object tracking can be divided into two categories: Two-Step methods and One-Shot methods. The difference lies in the methods used for target position detection and feature extraction. The Two-Step method first passes the picture through a detector model to obtain the detection boxes corresponding to the targets in the image, and then uses another feature extraction model to extract the corresponding feature vectors for each detection box in the detection results. This method makes the tracking effect relatively good. Although with the development of object detection algorithms and ReID (Person re-identification) in recent years, the Two-Step method has also shown obvious performance improvement in object tracking, the Two-Step method uses two independent models and needs to be trained separately, resulting in a very slow speed and making it difficult to perform inference at video rate. The One-Shot method integrates detection and feature extraction into one model, which makes the detection speed relatively fast, but their performance is not as good as that of the Two-Step method.

[0004] To balance speed and accuracy, One-Shot methods have begun to attract more and more attention. For example, Track-RCNN adds a ReID head after Mask-RCNN to obtain a detection box and a feature vector for each target. JDE uses the method of adding a ReID head to the YOLOv3 framework to achieve an inference speed close to the video rate. FairMot found that the anchor-based detection method of JDE is not suitable for multi-object tracking tasks, so it uses an anchor-free object detector to reduce ambiguous anchor boxes. We found that the above One-Shot methods directly utilize the feature maps output by the backbone network and perform object detection and ReID tasks simultaneously. Among them, the object detection task is responsible for determining the location where the target category appears in the image and the classification of the corresponding target, while the ReID task is responsible for extracting the feature vector corresponding to each target and matching it with all the target feature vectors that have appeared before to obtain the motion trajectory of the target. The object detection task requires a large distance between the feature vectors of different classes of targets and a small distance between the feature vectors of the same class of targets. On the contrary, for the ReID task to better distinguish the IDs of each target, it requires a large feature distance between different targets of the same class. However, the object detection task and the ReID task use the same input feature map, which results in ambiguity in the representation ability of the feature map, simultaneously suppressing the effects of the object detection and ReID tasks, leading to low detection accuracy and serious ID switching in most multi-object tracking methods. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention designs a multi-object pedestrian tracking method based on feature association and feature enhancement;

[0006] A multi-object pedestrian tracking method based on feature association and feature enhancement, comprising the following steps:

[0007] Step 1: Obtain the original surveillance video containing the detection targets, and perform frame cutting on the surveillance video to cut it into a sequence of several frame images;

[0008] Step 2: Use OpenCV to read the frame images of the original surveillance video, read the first frame image, and if the reading is successful, execute steps 3-7; if the reading of the frame image fails, it means that each frame of the video has been read and tracked, and execute step 8;

[0009] Step 3: Scale the frame image obtained in step 2 to the fixed size required by the backbone network of the neural network using OpenCV to obtain the scaled image L;

[0010] Step 4: Input the image L obtained in step 3 into the DLA34 backbone network of the neural network to extract its feature map;

[0011] Step 5: Use the feature association module to enhance the self-attention and mutual attention of the feature map extracted in Step 4 to obtain the detection association feature map and the ReID association feature map;

[0012] Step 5.1: Define the output feature map of the backbone network as F, where F ∈ R C*H / 4*W / 4 , R represents the set of real numbers, H is the height of image L, W represents the width of image L, and C is the number of feature map channels; Generate the strong feature map F1 by passing F through the max pooling operation, and generate the weak feature map F2 by passing F through the average pooling operation, where {F1, F2} ∈ R C*H ' *W '; where, H' and W' are the height and width of the feature maps F1 and F2; Perform convolution operations on the strong feature map F1 and the weak feature map F2 respectively to generate the strong convolution feature map T1 and the weak convolution feature map T2; Deform the strong convolution feature map T1 and the weak convolution feature map T2 to a fixed size to obtain the strong standard feature map M1 and the weak standard feature map M2, {M1, M2} ∈ R C*N ', where N' = H' * W'; Normalize the cross product of M1 and its transpose matrix to obtain the strong self-attention map SA1, and normalize the cross product of M2 and its transpose matrix to obtain the weak self-attention map SA2, as shown in the following formula (1):

[0013]

[0014] where, C represents the number of feature map channels, i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; k represents the label of the strong and weak feature maps for this method, when k = 1, represents the element value at the i-th row and j-th column of the strong self-attention map, represents the i-th element value of the strong standard feature map, represents the j-th element value of the strong standard feature map; when k = 2, represents the element value at the i-th row and j-th column of the weak self-attention map, represents the i-th element value of the weak standard feature map, represents the j-th element value of the weak standard feature map, exp is the exponential function with the natural constant e as the base, ∑ is the accumulation operation, and · represents the matrix multiplication operation;

[0015] Step 5.2: Cross-multiply the transpose matrices of the strong standard feature map M1 and the weak standard feature map M2, and normalize the cross-multiplication result to obtain the strong mutual-attention map SC1, as shown in the following formula (2):

[0016]

[0017] where, C represents the number of feature map channels, i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; represents the value of the i-th element of the strong standard feature map, represents the value of the j-th element of the weak standard feature map;

[0018] Step 5.3: Perform a cross product on the transposed matrices of the strong standard feature map M1 and the weak standard feature map M2, and the cross product result is transposed and normalized to obtain the weak mutual attention map SC2, as shown in the following formula (3);

[0019]

[0020] where C represents the number of channels of the feature map, and i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; represents the value of the i-th element of the weak standard feature map, represents the value of the j-th element of the strong standard feature map;

[0021] Step 5.4: Add the strong self-attention map SA1 obtained in Step 5.1 and the strong mutual attention map SC1 in Step 5.2 to obtain the fused strong attention map SME1, and add SA2 and SC2 to obtain the weak attention map SME2; to prevent the feature map after attention enhancement from losing the original information, the feature map F is multiplied, added, normalized, and deformed with the strong attention map SME1 to obtain the detection correlation feature map FM1, and the deformed feature map F is multiplied, added, normalized, and deformed with the weak attention map SME2 to obtain the ReID correlation feature map FM2;

[0022] Step 6: Input the detection correlation feature map FM1 obtained in Step 5 into the Detection branch to obtain the center point coordinates of the appearance of each detection target in the image L and the width and height of the detection box; the image L is obtained by scaling in Step 3; input the ReID correlation feature map FM2 obtained in Step 5 into the ReID branch with spatial attention and channel attention enhancement to obtain the feature vector corresponding to each detection target;

[0023] Step 6.1: Use the Detection branch to perform 3 different parallel convolution operations on the detection correlation feature map FM1 obtained in Step 5 to obtain the heat map, width and height, and center point offset of the detection target; according to the heat map, the center point coordinates of the appearance of the detection target in the image L can be obtained, and the center point coordinates are corrected according to the center point offset, and the corrected result is used as the final center point coordinates. Combining the width and height corresponding to this center point coordinates, the width and height of the prediction box of the detection target can be obtained;

[0024] Step 6.2: Adopt a ReID branch with enhanced spatial attention and channel attention. First, use convolution and pooling operations on the ReID correlation feature map FM2 obtained in Step 5 to obtain the spatial hierarchical weight SW. Multiply the spatial hierarchical weight SW with the ReID correlation feature map FM2 to obtain the spatially attention-enhanced feature map SF. Then, use convolution and pooling operations on the spatially attention-enhanced feature map SF to obtain the channel hierarchical weight CW. Multiply the channel hierarchical weight CW with the above-obtained spatially attention-enhanced feature map SF to obtain the channel attention-enhanced feature map SCF. SCF is the final feature vector corresponding to each detection target.

[0025] Step 7: Package the center point coordinates of the detection targets, the width and height of the detection boxes, and the feature vectors corresponding to each detection target obtained in Step 6 into target detection information. Use the DeepSort method to perform target tracking on all detection targets appearing in the current image L to obtain the trajectory ID and coordinate position corresponding to each detection target in the current image L, that is, the tracking result. Then, jump to Step 2 to obtain the next frame of the image. After successful acquisition, execute Steps 3 - 7. If the acquisition fails, it means that all frames of the image have been acquired, and directly execute Step 8.

[0026] Step 8: Create a txt file and write the tracking results of all frames of the image obtained in Step 7 into the txt file to obtain the multi-object pedestrian tracking result.

[0027] Advantageous technical effects of the present invention:

[0028] The present invention uses DLA34 as the target detection backbone network and adds a feature correlation module. By combining the Detection branch and the ReID branch with feature enhancement, the target detection result is obtained. The detection target is matched with the trajectory through the DeepSort algorithm, aiming to effectively detect and distinguish pedestrians appearing in the image, improve the detection accuracy, and reduce the number of ID switches.

[0029] It can make full use of the image information in the current video, and obtain pedestrian feature information and pedestrian trajectory information to the greatest extent from a practical perspective, with stronger general applicability.

[0030] Express the pedestrian tracking behavior according to the data itself, which is more real and closer to the changes in the pedestrian trajectory, conforms to the motion law, and has a fast data processing speed, further improving the detection and processing efficiency.

[0031] If there is an overlap of pedestrians during pedestrian tracking and detection, it will cause missing frames of overlapping pedestrians, so that effective matching cannot be performed on the current frame. By setting the threshold of the number of missing frames, the overlapping pedestrians exceeding the threshold of the number of missing frames are discarded. On the contrary, when the threshold of the number of missing frames is not exceeded and the overlapping pedestrians walk out of the occlusion area and appear in the camera's field of view, they will be rematched to the original trajectory, which will not affect the overall result. Description of the Drawings

[0032] Figure 1 It is a schematic flowchart of the multi-object pedestrian tracking method of this embodiment;

[0033] Figure 2 It is a schematic diagram of the feature association module of the multi-object pedestrian tracking method of this embodiment;

[0034] Figure 3 It is a schematic diagram of the ReID branch of the multi-object pedestrian tracking method of this embodiment. Specific Embodiments

[0035] The present invention will be further described below in conjunction with the drawings and embodiments;

[0036] Figure 1 It is a specific flowchart of the multi-object pedestrian tracking method of this embodiment: First, obtain the video data to be recognized. Read the frame image sequence of the video in turn, scale each frame of the picture to a fixed size to obtain picture L, input picture L into the backbone network, and the backbone network outputs the center points, offsets, widths and heights of the detected frames corresponding to all pedestrians in picture L, and the feature values of each target. A set of prediction frames can be uniquely determined according to the center point, offset, width and height of the detection frame. After encapsulating the prediction frame of each target and the feature value corresponding to each target into a detection object, the DeepSort algorithm is used for target tracking, and the existing trajectories are matched using the Kalman filter, appearance feature distance, and IOU (Intersection over Union, overlap degree) matching algorithm. For the trajectories that fail to match successfully, it is considered that the target state is lost in the current frame. For the targets that fail to match successfully, a new trajectory is generated for them and a trajectory ID is assigned. The above operations are repeated for all frame pictures. Finally, the tracking results of all frames are written into a txt file. For some frames, due to problems such as occlusion and overlap, the trajectories of the undetected targets are manually set with a missing frame number threshold. The threshold of this method is 30. If the number of consecutive missing frames exceeds the threshold, the trajectory will be discarded to avoid occupying computing space by the disappeared trajectory. For the targets that reappear within the missing frame number threshold, they are determined to be retrieved and their correct trajectory IDs are assigned.

[0037] A multi-object pedestrian tracking method based on feature association and feature enhancement specifically includes the following steps:

[0038] Step 1: Obtain the original surveillance video containing the detection target, and perform frame slicing on the surveillance video to slice it into several frame image sequences;

[0039] Step 2: Use OpenCV to read the frame image of the original surveillance video. If it is the first read, read the first frame image. If the read is successful, execute Steps 3 - 7; when the read of the frame image fails, it means that every frame of the video has been read, and execute Step 8;

[0040] Step 3: Use OpenCV to scale the frame image obtained in Step 2 to the fixed size required by the backbone network of the neural network to obtain the scaled image L;

[0041] Step 4: Input the scaled image L in Step 3 into the DLA34 backbone network of the neural network to extract its feature map;

[0042] Step 5: Use the feature correlation module to perform self - attention and mutual - attention enhancement on the feature map extracted in Step 4 to obtain the detection - associated feature map and the ReID - associated feature map; The feature correlation module is as Figure 2 shown.

[0043] Step 5.1: Define the output feature map of the backbone network as F, where F ∈ R C*H / 4*W / 4 , R represents the set of real numbers, H is the height of the image L, W represents the width of the image L, and C is the number of channels of the feature map; Generate the strong feature map F1 by performing max - pooling operation on F, and generate the weak feature map F2 by performing average - pooling on F, where {F1, F2} ∈ R C*H ' *W ', H' = 10, W' = 6; where, H' and W' are the height and width of the feature maps F1 and F2; Perform convolution operations on the strong feature map F1 and the weak feature map F2 respectively to generate the strong convolutional feature map T1 and the weak convolutional feature map T2; Deform the strong convolutional feature map T1 and the weak convolutional feature map T2 to a fixed size to obtain the strong standard feature map M1 and the weak standard feature map M2, {M1, M2} ∈ R C*N ', where N' = H' * W'; Obtain the strong self - attention map SA1 by normalizing the cross - product of M1 with its transpose matrix, and obtain the weak self - attention map SA2 by normalizing the cross - product of M2 with its transpose matrix, as shown in the following formula (1):

[0044]

[0045] where, C represents the number of channels of the feature map, i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; k represents the label of the strong and weak feature maps for this method. When k = 1, represents the element value of the i - th row and j - th column of the strong self - attention map, represents the i - th element value of the strong standard feature map, represents the value of the j-th element of the strong standard feature map; when k = 2, represents the value of the element in the i-th row and j-th column of the weak self-attention map, represents the value of the i-th element of the weak standard feature map, represents the value of the j-th element of the weak standard feature map, exp is the exponential function with the natural constant e as the base, ∑ is the summation operation, and · represents the matrix multiplication operation;

[0046] Step 5.2: Perform a cross product on the transposed matrices of the strong standard feature map M1 and the weak standard feature map M2, and normalize the cross product result to obtain the strong mutual attention map SC1, as shown in the following formula (2):

[0047]

[0048] where C represents the number of channels of the feature map, and i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; represents the value of the i-th element of the strong standard feature map, represents the value of the j-th element of the weak standard feature map;

[0049] Step 5.3: Perform a cross product on the transposed matrices of the strong standard feature map M1 and the weak standard feature map M2, and transpose and normalize the cross product result to obtain the weak mutual attention map SC2, as shown in the following formula (3);

[0050]

[0051] where C represents the number of channels of the feature map, and i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; represents the value of the i-th element of the weak standard feature map, represents the value of the j-th element of the strong standard feature map;

[0052] Step 5.4: Add the strong self-attention map SA1 obtained in Step 5.1 and the strong mutual attention map SC1 in Step 5.2 to obtain the fused strong attention map SME1, and add SA2 and SC2 to obtain the weak attention map SME2; to prevent the feature map after attention enhancement from losing the original information, multiply, add, normalize, and transform the feature map F after deformation with the strong attention map SME1 to obtain the detection correlation feature map FM1, and multiply, add, normalize, and transform the deformed feature map F with the weak attention map SME2 to obtain the ReID correlation feature map FM2;

[0053] Step 6: Input the detection correlation feature map FM1 obtained in Step 5 into the Detection branch to obtain the center point coordinates of each detected object in image L and the width and height of the detection box; Image L is obtained by scaling in Step 3; Input the ReID correlation feature map FM2 obtained in Step 5 into the ReID branch with enhanced spatial attention and channel attention to obtain the feature vector corresponding to each detected object;

[0054] Step 6.1: Use the Detection branch to perform 3 different parallel convolution operations on the detection correlation feature map FM1 obtained in Step 5 to obtain the heat map, width and height, and center point offset of the detected object; The scale of the heat map is (W / 4, H / 4, 1), the scale of the width and height is (W / 4, H / 4, 4), and the scale of the center point offset is (W / 4, H / 4, 2), where W and H are the width and height of image L, and the size in this network is fixed at (1088, 608). The heat map represents the probability value of each pixel position in the feature map FM1 corresponding to whether it is the center point of the object. If it is closer to the object center, its value is closer to 1, and conversely, if it is farther from the object center, its value is closer to 0. The center point coordinates of the detected object in image L can be obtained from the heat map, and the center point coordinates are corrected according to the center point offset. The corrected result is used as the final center point coordinates, and the width and height of the prediction box of the detected object can be obtained by combining the width and height corresponding to this center point coordinate;

[0055] Step 6.2: Adopt the ReID branch with enhanced spatial attention and channel attention, where the ReID branch is as shown in Figure 3 First, use convolution and pooling operations on the ReID correlation feature map FM2 obtained in Step 5 to obtain the spatial hierarchical weight SW, and multiply the spatial hierarchical weight SW by the ReID correlation feature map FM2 to obtain the spatially attention-enhanced feature map SF; Use convolution and pooling operations on the spatially attention-enhanced feature map SF to obtain the channel hierarchical weight CW, and multiply the channel hierarchical weight CW by the above-obtained spatially attention-enhanced feature map SF to obtain the channel attention-enhanced feature map SCF, and SCF is the final feature vector corresponding to each detected object;

[0056] Step 7: Package the center point coordinates of the detected object, the width and height of the detection box, and the feature vector corresponding to each detected object obtained in Step 6 into object detection information, and use the DeepSort method to perform object tracking on all detected objects in the current image L to obtain the trajectory ID and coordinate position corresponding to each detected object in the current image L, that is, the tracking result; Then jump to Step 2 to obtain the next frame of the image. If the acquisition is successful, continue to execute Steps 3-7. If the acquisition fails, it means that all frames of the image have been acquired, and directly execute Step 8;

[0057] In this embodiment, this step specifically uses the DeepSort algorithm to match the detection target with the trajectory to obtain the tracking result. The main method is as follows:

[0058] In the initial state, that is, when there is no existing trajectory in the trajectory pool, the detection result is used as the initial trajectory, and the feature vector corresponding to the detection result is used as the feature vector of the corresponding trajectory. An ID is assigned to each trajectory, and the initialized trajectory is added to the trajectory pool. The trajectory is a description of the historical appearance position of the target, and the trajectory pool is a set of historical trajectories.

[0059] If there are trajectories in the trajectory pool, calculate the appearance distance between the trajectories in the trajectory pool and the current detection result to obtain a cost matrix. Calculate the mean and variance according to the historical appearance position of the trajectory, use the Kalman filter algorithm to predict the next appearance position of each trajectory, and update the cost matrix according to the deviation between the predicted position and the actual position of the detection result. According to the cost matrix, use the Hungarian matching algorithm to match the target and the trajectory pool.

[0060] For the trajectories and targets that are not successfully matched, perform secondary matching. Calculate the cost through the overlap degree IOU to obtain a cost matrix, and use the Hungarian matching algorithm to match the unmatched targets and trajectories.

[0061] For the detection results that are not successfully matched in step 2 and step 3, they are considered as newly emerged targets, and their trajectories are initialized and new IDs are assigned and added to the trajectory pool. The trajectories that are not successfully matched are identified as the lost state in the current frame. If the number of consecutive lost frames exceeds a certain threshold, which is fixed at 30 in this method, it is considered that the trajectory has left, and this trajectory is removed from the trajectory pool.

[0062] Step 8: Create a txt file and write the tracking results of all frame images obtained in step 7 into the txt file to obtain the multi-target tracking results of pedestrians.

Claims

1. A pedestrian multi-object tracking method based on feature association and feature enhancement, characterized in that It includes the following steps: Step 1: Obtain the original surveillance video containing the detection target, perform frame cutting on the surveillance video, and cut it into several frame image sequences; Step 2: Use OpenCV to read the frame images of the original surveillance video, read the first frame image. If the reading is successful, execute Steps 3-7; when the frame image reading fails, it means that each frame of the video has been read and tracked, and execute Step 8; Step 3: Use OpenCV to scale the frame image obtained in Step 2 to the fixed size required by the backbone network of the neural network to obtain the scaled image L; Step 4: Input the image L obtained in Step 3 into the DLA34 backbone network of the neural network to extract its feature map; Step 5: Adopt a feature association module to enhance the self-attention and mutual attention of the feature map extracted in Step 4 to obtain a detection association feature map and a ReID association feature map; Step 6: Input the detection association feature map FM1 obtained in Step 5 into the Detection branch to obtain the central point coordinates where each detection target appears in the image L and the width and height of the detection box; The image L is obtained by scaling in Step 3; Input the ReID association feature map FM2 obtained in Step 5 into the ReID branch with spatial attention and channel attention enhancement to obtain the feature vector corresponding to each detection target; Step 7: Package the central point coordinates of the detection target, the width and height of the detection box, and the feature vector corresponding to each detection target obtained in Step 6 into target detection information, use the DeepSort method to perform target tracking on all detection targets appearing in the current image L, and obtain the trajectory ID and coordinate position corresponding to each detection target in the current image L, that is, the tracking result; then jump to Step 2 to obtain the next frame image. After successful acquisition, execute Steps 3-7. If the acquisition fails, it means that all frame images have been acquired, and directly execute Step 8; Step 8: Create a txt file and write the tracking results of all frame images obtained in Step 7 into the txt file to obtain the multi-object tracking result of pedestrians.

2. The multi-object pedestrian tracking method based on feature association and feature enhancement according to claim 1, wherein Step 5 is specifically: Step 5.1: Define the output feature map of the backbone network as F, where F ∈ R C*H / 4*W / 4 , R represents the set of real numbers, H is the height of image L, W represents the width of image L, and C is the number of channels of the feature map; Generate a strong feature map F1 by performing max pooling on F, and generate a weak feature map F2 by performing average pooling on F, where {F1, F2} ∈ R C*H'*W' ; where, H' and W' are the height and width of feature maps F1 and F2 respectively; Perform convolution operations on the strong feature map F1 and the weak feature map F2 respectively to generate a strong convolutional feature map T1 and a weak convolutional feature map T2; Deform the strong convolutional feature map T1 and the weak convolutional feature map T2 to a fixed size to obtain a strong standard feature map M1 and a weak standard feature map M2, {M1, M2} ∈ R C*N' , where N' = H' * W'; Multiply M1 by the transpose matrix of itself and normalize to obtain a strong self-attention map SA1, multiply M2 by the transpose matrix of itself and normalize to obtain a weak self-attention map SA2, as shown in the following formula (1): Among them, C represents the number of channels of the feature map, and i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; k represents the label of the strong and weak feature maps. When k = 1, represents the element value at the i-th row and j-th column of the strong self-attention map, represents the i-th element value of the strong standard feature map, represents the j-th element value of the strong standard feature map; when k = 2, represents the element value at the i-th row and j-th column of the weak self-attention map, represents the i-th element value of the weak standard feature map, represents the j-th element value of the weak standard feature map, exp is the exponential function with the natural constant e as the base, ∑ is the accumulation operation, and · represents the matrix multiplication operation; Step 5.2: Perform a cross product on the transposed matrices of the strong standard feature map M1 and the weak standard feature map M2, and normalize the cross product result to obtain the strong mutual attention map SC1, as shown in the following formula (2): Among them, C represents the number of channels of the feature map, and i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; represents the value of the i-th element of the strong standard feature map, represents the value of the j-th element of the weak standard feature map; Step 5.3: Perform a cross product on the transposed matrices of the strong standard feature map M1 and the weak standard feature map M2, and after transposing and normalizing the cross product result, obtain the weak mutual attention map SC2, as shown in the following formula (3); Among them, C represents the number of channels of the feature map, and i, j ∈ {1, 2, 3,..., C} represent the element indices in the strong and weak standard feature maps; represents the value of the i-th element of the weak standard feature map, represents the value of the j-th element of the strong standard feature map; Step 5.4: Add the strong self-attention map SA1 obtained in Step 5.1 and the strong mutual attention map SC1 in Step 5.2 to obtain the fused strong attention map SME1, and add SA2 and SC2 to obtain the weak attention map SME2; to prevent the feature map from losing its original information after attention enhancement, multiply, add, normalize, and transform the feature map F and then multiply it with the strong attention map SME1 to obtain the detection association feature map FM1, and multiply, add, normalize, and transform the deformed feature map F with the weak attention map SME2 to obtain the ReID association feature map FM2.

3. A multi-object pedestrian tracking method based on feature association and feature enhancement according to claim 1, characterized in that Step 6 is specifically: Step 6.1: Use the Detection branch to perform three different parallel convolution operations on the detection correlation feature map FM1 obtained in Step 5 to obtain the heat map of the detection target, the width, height, and center point offset; according to the heat map, the center point coordinates where the detection target appears in the image L can be obtained. The center point coordinates are corrected according to the center point offset, and the corrected result is used as the final center point coordinates. Combining the width and height corresponding to this center point coordinates, the width and height of the prediction box of the detection target can be obtained; Step 6.2: Adopt the ReID branch with enhanced spatial attention and channel attention. First, use convolution and pooling operations on the ReID correlation feature map FM2 obtained in Step 5 to obtain the spatial hierarchical weight SW. Multiply the spatial hierarchical weight SW by the ReID correlation feature map FM2 to obtain the spatially attention-enhanced feature map SF; use convolution and pooling operations on the spatially attention-enhanced feature map SF to obtain the channel hierarchical weight CW. Multiply the channel hierarchical weight CW by the above-obtained spatially attention-enhanced feature map SF to obtain the channel attention-enhanced feature map SCF. SCF is the final feature vector corresponding to each detection target.

Citation Information

Patent Citations

  • ReID feature-based strong data association integrated real-time multi-target tracking method

    CN112487934A

  • Unmanned aerial vehicle video multi-target tracking method based on attention feature fusion

    CN113807187A