Pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction

By improving the YOLOv5s object detection algorithm and combining the ResNet-50 feature extraction network and CBAM attention mechanism, the problem of multi-object tracking of pedestrians in complex scenarios is solved, and accurate tracking and trajectory analysis of multiple pedestrians is achieved.

CN119991744AInactive Publication Date: 2025-05-13SHANDONG UNIV OF TECH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510078795.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively track multiple pedestrian targets in complex scenarios, resulting in increased difficulty in trajectory analysis and intelligent management of urban public security.

Method used

By improving the YOLOv5s object detection algorithm, deformable convolution and optimized loss functions are introduced, and combined with the ResNet-50 feature extraction network and CBAM attention mechanism, effective processing of pedestrian multi-objective tracking is achieved.

Benefits of technology

It is possible to accurately track multiple pedestrian targets in complex scenarios, improving the accuracy of trajectory analysis and the efficiency of urban public security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991744A_ABST
    Figure CN119991744A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, belongs to the field of machine vision, and improves a YOLOv5s target detection algorithm. An improved YOLOv5s-DCN is used as a detection algorithm, and a target detection frame, a target ID and target position information are output; predicting a target track according to the position information and the motion information of the previous frame of target; resNet-50 is used as a feature extraction network reference model, and improvement and training are carried out; performing cascade matching on a detection frame output by the detection algorithm and a prediction frame in a confirmation state in a previous frame by using the trained feature extraction model; performing IOU matching on the detection frame and the prediction frame which are not successfully matched with the prediction frame in an unconfirmed state in the previous frame; and processing the detection frame, the prediction frame and the prediction track after IOU matching. According to the pedestrian multi-target tracking method based on the improved YOLOv5 and feature extraction provided by the invention, a plurality of pedestrian targets in a complex scene can be effectively tracked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and in particular to a pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction. Background Art

[0002] Vision is the primary way of human perception. Visual information is the main way for people to understand the world and interact with the outside world. Research on the interpretation and processing of visual information is crucial. Computer vision originated in the 1950s. It aims to imitate the working method of the human visual system, so that computers can obtain visual information from the real world and analyze and process it.

[0003] With the rapid development of artificial intelligence in recent years, real-life production, industrial manufacturing, safety supervision and other real-life work tasks are gradually moving towards the intelligent era. Multi-target tracking is an important branch of computer vision. Its main function is to identify the position of the same target in continuous image frames, link the position of the target in different frames to form a trajectory, so that the computer can determine the starting point and end point of each target. Multi-target tracking is a multidisciplinary research field, which includes components such as target detection, target re-identification, data matching and motion prediction. Combined with control theory, sensor technology, image processing, pattern recognition and other multi-field technologies, it constitutes a comprehensive topic with great research significance and value. With the development of society, the population density in various scenes in the city is increasing. Accurately and effectively analyzing the trajectory and behavior of pedestrians in complex scenes is of great significance to realizing the intelligent management of urban public security and promoting the development of social informatization. Multi-target tracking is very useful for statistical analysis tasks in complex scenes, so it is widely used in pedestrian data analysis in various scenes. Summary of the invention

[0004] The purpose of the present invention is to provide a pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, which is used to track multiple pedestrian targets in multiple complex scenes.

[0005] To achieve the above object, the present invention provides a pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, comprising the following steps:

[0006] S1. Improve the YOLOv5s target detection algorithm by introducing deformable convolution and optimizing loss function to obtain YOLOv5s-DCN.

[0007] S2, perform target detection on the first frame, use the improved YOLOv5s-DCN as the detection algorithm, and output the target detection box, target ID and target location information;

[0008] S3, using the Kalman filter to predict the target trajectory based on the position information and motion information of the target in the previous frame, and outputting the confirmed trajectory and the uncertain trajectory;

[0009] S4. Use ResNet-50 as the feature extraction network benchmark model, introduce the CBAM attention mechanism to improve the feature extraction model, and use public datasets for training;

[0010] S5, performing target detection on the target in the second frame, and using the trained feature extraction model to perform cascade matching on the detection frame output by the detection algorithm and the prediction frame in the confirmed state in the previous frame;

[0011] S6, performing IOU matching between the detection frame and the prediction frame that have not been successfully matched after the cascade matching and the prediction frame that is in an unconfirmed state in the previous frame;

[0012] S7. Process the detection box, prediction box and prediction trajectory after IOU matching:

[0013] The successfully matched detection frame and prediction frame enter the trajectory update of the next frame; the unsuccessfully matched detection frame enters the trajectory update of the next frame as the new target in the current frame; the unsuccessfully matched predicted trajectory, if it is in an unconfirmed state, the trajectory is deleted; if it is in a confirmed state, the number of adaptations of the current trajectory is determined. If it exceeds the threshold, the trajectory is deleted. If it does not exceed the set threshold, the trajectory is updated for the next frame.

[0014] Preferably, the deformable convolution content in S1 is as follows:

[0015] The operation process of deformable convolution is expressed as:

[0016]

[0017] Where: P n is an integer, representing the offset of each point of the convolution output relative to each point on the receptive field; ω is the convolution kernel, R = {(-1,-1), (-1,0), ..., (0,1), (1,1)} represents the feature vector; p0 is the current pixel; ΔP n is the two-dimensional offset;

[0018] The third version of the deformable convolution DCNv3 using the multi-group mechanism is expressed as:

[0019]

[0020] Where: p0 is the current pixel; G represents the number of aggregation groups; m gk ∈R represents the modulation scalar of the kth sampling point in the gth group and is normalized along dimension K by the softmax function; x gis the input feature map of the slice; p k is the grid sampling position in the gth group; Δp gk For p k The offset at the position.

[0021] Preferably, the loss function in S1 is EIOU, and the calculation process of EIOU loss function is:

[0022] L EIoU =1-IoU+L dis +L asp (3)

[0023]

[0024] Where: L EIoU is the EIOU loss function; L dis is the distance loss; L asp is the edge length loss; b p and b gt Represent the center points of the predicted box and the real box respectively; ρ 2 (b p ,b gt ) is the Euclidean distance between the center points of the real box and the predicted box; c is the minimum diagonal length of the outer rectangle of the real box and the predicted box; w p 、w gt 、h p and h gt They are respectively represented as the width of the real box, the width of the predicted box, the height of the real box and the height of the predicted box; c w and c h The width and height of the minimum outer rectangle of the true box and the predicted box.

[0025] Preferably, the steps for improving the YOLOv5s target detection algorithm are as follows:

[0026] Data enhancement at the input end uses a mosaic approach to enhance the input data to improve the generalization and robustness of the model;

[0027] After data enhancement, the feature map is sliced ​​through the feature extraction backbone network Backbone, and the sliced ​​feature maps are stacked in the channel dimension. The CSP module solves the problem of large amount of calculation caused by gradient data. The upper part passes through the CBL module and a layer of standard convolution, and the lower part undergoes standard convolution. The output results of the two parts are then spliced ​​together. After that, the SPP module extracts features from images of different sizes and outputs feature maps of the same scale, which increases the generalization ability of the model.

[0028] Subtract the mean envelope from x(t) to get the intermediate signal;

[0029] The backbone network output is subjected to feature fusion through Neck;

[0030] The Neck output passes through the Head module to perform multi-scale target detection on the extracted feature maps.

[0031] Preferably, the CBAM attention mechanism steps in S4 are as follows:

[0032] The input feature map first passes through the channel attention module to obtain the feature map of the target in the channel domain. The result is multiplied by the input feature map and then input into the spatial attention module to obtain the feature map of the target in the spatial domain. The result is multiplied by the input feature map to obtain the output feature map of CBAM. The calculation process of CBAM feature map can be expressed as:

[0033]

[0034] Where: F is the input feature map; F' is the channel attention output feature map; F″ is the CBAM output feature map; W c and W s Represent channel weight and spatial weight respectively;

[0035] After global maximum pooling and global average pooling, each channel of the input feature map F is dimensionally upgraded and dimensionally reduced through a multi-layer perceptron sharing a fully connected layer to obtain a channel weight vector. The channel weight vectors of the two pooling methods are added together and activated using a sigmoid function to obtain a comprehensive channel weight coefficient W. c , the channel attention mechanism can be expressed as:

[0036]

[0037] Where σ is the Sigmoid activation function; W1 and W are the weight matrices in MLP respectively; AvgPool and MaxPool are the average pooling and maximum pooling operations respectively; and are the features after average pooling and maximum pooling respectively; W1∈R C ×C / r , W0∈R C / r×C , r is the dimension reduction factor;

[0038] The input feature map is subjected to global maximum pooling and global average pooling respectively, and then the processing results are stacked together and passed through a shared convolution layer. The convolution kernel size of this convolution layer is 7×7 and the number of channels is 2. After the convolution extracts the features, the sigmoid function is used for activation to obtain the comprehensive channel weight coefficient W s , spatial attention can be expressed as:

[0039]

[0040] Where F' is the spatial attention input feature; f 7×7 Indicates that a 7×7 convolution kernel is used for convolution; and They are the features after average pooling and maximum pooling respectively.

[0041] Therefore, the present invention adopts the above-mentioned pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, and can effectively track multiple pedestrian targets in complex scenes through improved YOLOv5 and feature extraction.

[0042] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of an implementation of an embodiment of the present invention;

[0044] Figure 2 This is a specific structural diagram of the YOLOv5s-DCN detection algorithm according to an embodiment of the present invention;

[0045] Figure 3 This is the feature extraction network structure diagram of fusion CBAM. DETAILED DESCRIPTION

[0046] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] See also Figure 1-Figure 2 , a pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, comprising the following steps:

[0048] S1. Improve the YOLOv5s target detection algorithm by introducing deformable convolution and optimizing the loss function.

[0049] The deformable convolution is as follows:

[0050] (1) The operation process of deformable convolution can be expressed as:

[0051]

[0052] Where: P n is an integer, representing the offset of each point of the convolution output relative to each point on the receptive field, ω is the convolution kernel, R = {(-1,-1), (-1,0), ..., (0,1), (1,1)} represents the feature vector; ΔPn is the two-dimensional offset;.

[0053] (2) DCNv3 is the third version of deformable convolution, which introduces a multi-group mechanism. The spatial aggregation process is grouped, and different groups have different offsets and adjustment scales during the sampling process, forming multiple spatial aggregation modes to obtain more features. Weights are shared between convolution neurons. The convolution weights are divided into two parts: depth weights and point weights. The depth part is adjusted by the initial position perception, and the point part is shared. The modulation scalars are normalized along the sampling points. The elements along the sampling points are normalized using softmax, and the sum of all modulation scalars is 1, making the model training process more stable.

[0054] In summary, DCNv3 can be expressed as:

[0055]

[0056] Where: p0 is the current pixel, G represents the number of aggregation groups, m gk ∈R represents the modulation scalar of the kth sampling point in the gth group and is normalized along dimension K by the softmax function. g is the input feature map of the slice. k is the grid sampling position in group g, Δp gk For p k The offset at the position.

[0057] The loss function is EIOU as the loss function, and the EIOU loss calculation process is:

[0058] L EIoU =1-IoU+L dis +L asp (3)

[0059]

[0060] Where: L EIoU is the EIOU loss function; L dis is the distance loss; L asp is the edge length loss; b p and b gt Represent the center points of the predicted box and the real box respectively; ρ 2 (b p ,b gt ) is the Euclidean distance between the center points of the real box and the predicted box; c is the minimum diagonal length of the outer rectangle of the real box and the predicted box; w p 、w gt 、h p and h gt They are respectively represented as the width of the real box, the width of the predicted box, the height of the real box and the height of the predicted box; cw and c h The width and height of the minimum outer rectangle of the true box and the predicted box.

[0061] The steps to improve the YOLOv5s target detection algorithm are as follows:

[0062] (1) Data enhancement at the input end: Mosaic method is used to enhance the input data to improve the generalization ability and robustness of the model.

[0063] (2) After data enhancement, the feature map is sliced ​​through the feature extraction backbone network Backbone, and the sliced ​​feature maps are stacked in the channel dimension. The CSP module solves the problem of large amount of calculation caused by gradient data. The upper part passes through the CBL module and a layer of standard convolution, and the lower part undergoes standard convolution. The output results of the two parts are then spliced ​​together. After that, the SPP module is used to extract features from images of different sizes and output feature maps of the same scale, which increases the generalization ability of the model.

[0064] (3) Subtract the mean envelope from x(t) to obtain the intermediate signal.

[0065] (4) The backbone network output is subjected to feature fusion through Neck.

[0066] (5) The Neck output passes through the Head module to perform multi-scale target detection on the extracted feature map.

[0067] S2. Perform target detection on the first frame, use the improved YOLOv5s-DCN as the detection algorithm, and output the target detection box, target ID and target location information.

[0068] S3. Use the Kalman filter to predict the target trajectory based on the position information and motion information of the target in the previous frame, and output the confirmed trajectory and the uncertain trajectory.

[0069] S4. Use ResNet-50 as the feature extraction network benchmark model, introduce the CBAM attention mechanism to improve the feature extraction model, and use public datasets for training.

[0070] like Figure 3 , the steps of CBAM attention mechanism are as follows:

[0071] (1) The input feature map first passes through the channel attention module to obtain the feature map of the target in the channel domain. The result is multiplied by the input feature map and then input into the spatial attention module to obtain the feature map of the target in the spatial domain. The result is multiplied by the input feature map to obtain the output feature map of CBAM. The CBAM feature map calculation process can be expressed as:

[0072]

[0073] Where: F is the input feature map; F' is the channel attention output feature map; F″ is the CBAM output feature map; W c and W s Represent channel weight and spatial weight respectively.

[0074] (2) After global maximum pooling and global average pooling, each channel of the input feature map F is dimensionally upgraded and dimensionally reduced through a multi-layer perceptron sharing a fully connected layer to obtain a channel weight vector. The channel weight vectors of the two pooling methods are added together and activated using a sigmoid function to obtain a comprehensive channel weight coefficient W. c , the channel attention mechanism can be expressed as:

[0075]

[0076] Where σ is the Sigmoid activation function; W1 and W are the weight matrices in MLP respectively; AvgPool and MaxPool are the average pooling and maximum pooling operations respectively; and are the features after average pooling and maximum pooling respectively; W1∈R C ×C / r , W0∈R C / r×C , r is the dimensionality reduction factor.

[0077] (3) Perform global maximum pooling and global average pooling operations on the input feature map, and then stack the processing results together and pass them through a shared convolution layer. The convolution kernel size of this convolution layer is 7×7 and the number of channels is 2. After the convolution extracts the features, the sigmoid function is used for activation to obtain the comprehensive channel weight coefficient W. s , spatial attention can be expressed as:

[0078]

[0079] Where F' is the spatial attention input feature; f 7×7 Indicates that a 7×7 convolution kernel is used for convolution; and They are the features after average pooling and maximum pooling respectively.

[0080] S5. Detect the target in the second frame, and use the trained feature extraction model to perform cascade matching on the detection box output by the detection algorithm and the prediction box in the confirmed state in the previous frame.

[0081] S6. Perform IOU matching on the detection boxes and prediction boxes that have not been successfully matched after the cascade matching and the prediction boxes that are in an unconfirmed state in the previous frame.

[0082] S7. Process the detection box, prediction box and prediction trajectory after IOU matching:

[0083] The successfully matched detection frame and prediction frame enter the trajectory update of the next frame; the unsuccessfully matched detection frame enters the trajectory update of the next frame as the new target in the current frame; the unsuccessfully matched predicted trajectory, if it is in an unconfirmed state, the trajectory is deleted; if it is in a confirmed state, the number of adaptations of the current trajectory is determined. If it exceeds the threshold, the trajectory is deleted. If it does not exceed the set threshold, the trajectory is updated for the next frame.

[0084] Therefore, the present invention adopts the above-mentioned pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, and can effectively track multiple pedestrian targets in complex scenes through improved YOLOv5 and feature extraction.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction, characterized in that: The following steps are involved: S1. Improve the YOLOv5s target detection algorithm by introducing deformable convolution and optimizing loss function to obtain YOLOv5s-DCN. S2, perform target detection on the first frame, use the improved YOLOv5s-DCN as the detection algorithm, and output the target detection box, target ID and target location information; S3, using the Kalman filter to predict the target trajectory according to the position information and motion information of the target in the previous frame, and outputting the confirmed trajectory and the uncertain trajectory; S4. Use ResNet-50 as the feature extraction network benchmark model, introduce the CBAM attention mechanism to improve the feature extraction model, and use public datasets for training; S5, performing target detection on the target in the second frame, and using the trained feature extraction model to perform cascade matching on the detection frame output by the detection algorithm and the prediction frame in the confirmed state in the previous frame; S6, performing IOU matching between the detection frame and the prediction frame that have not been successfully matched after the cascade matching and the prediction frame that is in an unconfirmed state in the previous frame; S7. Process the detection box, prediction box and prediction trajectory after IOU matching: The successfully matched detection boxes and prediction boxes enter the trajectory update of the next frame; the unsuccessfully matched detection boxes enter the trajectory update of the next frame as the new targets in the current frame; For the predicted trajectory that is not successfully matched, if it is in the unconfirmed state, the trajectory is deleted; if it is in the confirmed state, the number of adaptations of the current trajectory is determined. If it exceeds the threshold, the trajectory is deleted. If it does not exceed the set threshold, the trajectory of the next frame is updated.

2. A pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction according to claim 1, characterized in that: The content of deformable convolution in S1 is as follows: The operation process of deformable convolution is expressed as: Where: P n is an integer, representing the offset of each point of the convolution output relative to each point on the receptive field; ω is the convolution kernel, R = {(-1,-1), (-1,0), ..., (0,1), (1,1)} represents the feature vector; p0 is the current pixel; ΔP n is the two-dimensional offset; The third version of the deformable convolution DCNv3 using the introduction of a multi-group mechanism is expressed as: Where: p0 is the current pixel; G represents the number of aggregation groups; m gk ∈R represents the modulation scalar of the kth sampling point in the gth group and is normalized along dimension K by the softmax function; x g is the input feature map of the slice; p k is the grid sampling position in the gth group; Δp gk For p k The offset at the position.

3. A pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction according to claim 2, characterized in that: The loss function in S1 is EIOU, and the calculation process of EIOU loss function is: L EIoU =1-IoU+L dis +L asp (3) Where: L EIoU is the EIOU loss function; L dis is the distance loss; L asp is the edge length loss; b p and b gt Represent the center points of the predicted box and the real box respectively; ρ 2 (b p ,b gt ) is the Euclidean distance between the center points of the real box and the predicted box; c is the minimum diagonal length of the outer rectangle of the real box and the predicted box; w p 、w gt 、h p and h gt They are respectively represented as the width of the real box, the width of the predicted box, the height of the real box and the height of the predicted box; c w and c h The width and height of the minimum outer rectangle of the true box and the predicted box.

4. A pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction according to claim 3, characterized in that: The steps to improve the YOLOv5s target detection algorithm are as follows: Data enhancement at the input end uses mosaic method to enhance the input data; After data enhancement, the feature map is sliced ​​through the feature extraction backbone network Backbone, and the sliced ​​feature maps are stacked in the channel dimension. The upper part passes through the CSP module, and the CBL module and a layer of standard convolution are performed on the lower part. The output results of the upper and lower parts are then spliced ​​together. The SPP module extracts features from images of different sizes and outputs feature maps of the same scale. Subtract the mean envelope from x(t) to get the intermediate signal; The backbone network output is subjected to feature fusion through Neck; The Neck output passes through the Head module to perform multi-scale target detection on the extracted feature maps.

5. A pedestrian multi-target tracking method based on improved YOLOv5 and feature extraction according to claim 4, characterized in that: The steps of the CBAM attention mechanism in S4 are as follows: The input feature map first passes through the channel attention module to obtain the feature map of the target in the channel domain. The result is multiplied by the input feature map and then input into the spatial attention module to obtain the feature map of the target in the spatial domain. The result is multiplied by the input feature map to obtain the output feature map of CBAM. The calculation process of CBAM feature map is expressed as: Where: F is the input feature map; F' is the channel attention output result feature map; F″ is the feature map of CBAM output result; W c and W s Represent channel weight and spatial weight respectively; After global maximum pooling and global average pooling, each channel of the input feature map F is dimensionally upgraded and dimensionally reduced through a multi-layer perceptron sharing a fully connected layer to obtain a channel weight vector. The channel weight vectors of the two pooling methods are added together and activated using a sigmoid function to obtain a comprehensive channel weight coefficient W. c , the channel attention mechanism is expressed as: Where σ is the Sigmoid activation function; W1 and W are the weight matrices in MLP respectively; AvgPool and MaxPool are the average pooling and maximum pooling operations respectively; and They are the features after average pooling and maximum pooling respectively; W1∈R C×C / r , W0∈R C / r×C , r is the dimension reduction factor; The input feature map is subjected to global maximum pooling and global average pooling respectively, and then the processing results are stacked and passed through a shared convolution layer with a convolution kernel size of 7×7 and a channel number of 2. After the convolution extracts the features, the sigmoid function is used for activation to obtain the comprehensive channel weight coefficient W. s , spatial attention is expressed as: Where F' is the spatial attention input feature; f 7×7 Indicates that convolution is performed using a 7×7 convolution kernel; and They are the features after average pooling and maximum pooling respectively.

Citation Information

Patent Citations

  • Residual network target tracking method and system based on attention mechanism

    CN116645399A

  • Target detection method based on dense scene

    CN116935435A

  • Pedestrian multi-target tracking method combined with instance segmentation

    CN116993775A

  • Pedestrian multi-target tracking method based on deep learning

    CN117237411A

  • Video multi-target tracking method based on domain adaptive feature fusion

    CN117541625A