Multi-target tracking method and system based on random channel adaptive attention mechanism
By introducing a random channel adaptive attention mechanism in multi-objective tracking technology, explicitly weighting the importance of target features, the challenges of occlusion and real-time tracking are solved, and fast and accurate multi-objective tracking is achieved.
Patent Information
- Application Number
- CN202210501399.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing multi-objective tracking technology is prone to large position displacement due to obstacles during target movement, resulting in the inability to continuously and stably track after reproduction, and it is difficult to achieve fast and accurate target detection and tracking in real-time application scenarios.
A multi-objective tracking method based on the random channel adaptive attention mechanism is adopted. Through four steps: random, extrusion, excitation and weighting, the importance of target features is explicitly displayed and weighted, the significance of the target features is enhanced, and the distinctive feature extraction capability of the model is used, and a lightweight deep aggregation neural network is used for feature extraction and tracking.
The anti-occlusion and long-term tracking capabilities of multi-objective tracking are improved, and the fast and accurate tracking of multi-objective videos is achieved, achieving a good balance of speed and accuracy.
Smart Images

Figure CN114842388B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, in particular to the field of video multi-target tracking, and specifically relates to a multi-target tracking method and system based on a random channel adaptive attention mechanism. Background Art
[0002] With the application of deep learning technology in tasks related to visual multi-target tracking and its great success, the current deep learning-based visual multi-target tracking technology framework is mostly a two-stage model, that is, the DBT model that realizes multi-target tracking based on data association based on target detection results. In addition, from the perspective of the structure of deep neural networks, sub-modules in DBT such as feature extraction can be integrated into the target detection network, and joint detection and tracking based on the fusion of sub-modules in DBT, that is, the JDT model, is realized. Using a deep network framework to realize visual multi-target tracking is a new trend in the past two years.
[0003] The general process of DBT is mainly composed of three main modules: target detection, feature extraction, and data association. The JDT algorithm framework that has emerged in the past two years has reduced the complexity of the DBT framework while also improving the accuracy of multi-target tracking. Its strategy is to integrate some functional modules to a certain extent based on DBT, reduce the algorithm complexity brought by staged processing, and increase the coupling between functional modules.
[0004] At present, DBT algorithms based on deep learning have made huge gains in target tracking by embedding deep learning modules in traditional sub-modules, such as target detection, feature extraction, and data association, but this has also brought about a loss in algorithm speed; deep learning neural networks are the key to the development of JDT algorithms. By integrating multi-module joint learning, online end-to-end multi-target tracking algorithms are realized. Among them, detector performance plays a key role in the tracking performance of JDT algorithms. At present, multi-target tracking is generally developing from DBT algorithms to JDT, achieving a balance between algorithm accuracy and speed in stages. However, current multi-target tracking still has problems such as frequent occlusion of multiple targets and the trade-off between accuracy and speed. Specifically, it can be divided into the following points:
[0005] 1) During the target movement, it is often partially or completely blocked by obstacles. The position of pedestrians before and after the blockage is greatly displaced, which makes it impossible to continue to track stably after reappearance;
[0006] 2) When the tracking algorithm is deployed to actual application scenarios, the real-time requirements of the algorithm need to be considered. Implementing a real-time target detection and tracking algorithm is also a major challenge. Summary of the invention
[0007] The purpose of the present invention is to provide a multi-target tracking method and system based on a random channel adaptive attention mechanism, and to establish an intelligent model that can quickly and accurately realize multi-target tracking in video surveillance, so as to solve the problem that the target is often partially or completely blocked by obstacles during movement, and there is a large displacement of the pedestrian position before and after the blockage, resulting in the inability to continue to track stably after reproduction.
[0008] The technical solution adopted by the present invention is:
[0009] A multi-target tracking method based on random channel adaptive attention mechanism mainly includes four steps: randomization, squeezing, excitation and weighting, where:
[0010] Random refers to the random activation of channel attention. A gating mechanism is added to the features of each target during transmission and aggregation to achieve the random activation of the spatial adaptive attention module. By artificially setting the random ratio, the feature channels of multiple targets are partially enhanced.
[0011] Squeezing refers to the global average pooling operation. After the adaptive attention module is started, the feature maps of multiple targets transmitted are subjected to global average pooling, thereby completing the compression of the feature maps. In this way, the feature maps of each target have a global receptive field.
[0012] Excitation refers to using the fully connected layer to restore the features of each target to its original dimension, and then calculate the weight of each channel;
[0013] Weighting means to weight the importance of different channels of the feature maps of each original target by using the calculated weights to strengthen the role of significant features in the image;
[0014] After weighting, the model outputs a more accurate dense heat map of the center points of each target, then predicts the offset of the center points of each target between adjacent video frames, and finally uses a greedy association algorithm to achieve multi-target tracking.
[0015] As mentioned above, a multi-target tracking method based on a random channel adaptive attention mechanism is proposed, in which: the random module strengthens the learning ability of the model by setting random numbers, so that important feature information receives more attention; on this basis, squeezing, excitation and weighting make the feature importance explicitly displayed on the weight, thereby strengthening the model's ability to extract significant features. Furthermore, a lightweight deep aggregation neural network model is used for feature extraction, and the result heat map of the previous frame, current frame and previous frame of the video sequence is used as the input of the neural network, and the tracking mark and trajectory of the target in the video sequence is used as the output of the network. The weight obtained above is used as the significant feature of the target for identification and tracking.
[0016] The present invention also provides a multi-target tracking system based on a random channel adaptive attention mechanism, which comprises:
[0017] The random module is used to randomly start the spatial adaptive attention module when the features of each target are transmitted and aggregated through the gating mechanism, and to partially enhance the feature channels of multiple targets by setting random ratios;
[0018] The squeezing module is used to perform global average pooling on the feature maps of multiple targets passed in after the adaptive attention module is started, thereby completing the compression processing of the feature maps;
[0019] The excitation module is used to restore the features of each target to its original dimension using the fully connected layer, and then calculate the weights of each channel;
[0020] The weighting module is used to weight the importance of different channels of the feature map of each target by using the calculated weights to strengthen the role of significant features in the image;
[0021] The tracking module is used to output the center point of each target after weighting, predict the offset of the center point of each target between adjacent video frames, and realize multi-target tracking using a greedy association algorithm.
[0022] The beneficial effects of the present invention are:
[0023] (1) The present invention has strong anti-occlusion and long-term tracking capabilities, and can well solve the problem of multi-target tracking tasks.
[0024] (2) The present invention realizes fast and accurate tracking of multiple targets in video, achieving a good balance between speed and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flow chart of the multi-target tracking method based on random channel adaptive attention mechanism;
[0026] Figure 2 This is a framework diagram of a center-based tracker;
[0027] Figure 3 This is a diagram of the multi-target tracking network architecture based on random channel adaptive attention;
[0028] Figure 4 This is the module structure diagram of the random channel adaptive attention mechanism;
[0029] Figure 5 Implementation example graph for multiple object tracking. DETAILED DESCRIPTION
[0030] The following is an introduction to a multi-target tracking method based on a random channel adaptive attention mechanism provided by the present invention in conjunction with the accompanying drawings and embodiments:
[0031] like Figure 1 As shown in the figure, a multi-target tracking method based on random channel adaptive attention mechanism mainly includes four modules: random, squeeze, excitation and weighted. Among them, random is to add a gating mechanism during feature transmission and feature aggregation to realize the random start of the spatial adaptive attention module, and the random ratio needs to be set manually; squeeze is to perform global average pooling on the feature map transmitted after the module is started; excitation is to restore the feature to the original dimension using the fully connected layer, and then calculate the weight of each channel; weighted is to weight the importance of different channels of the original feature map by the calculated weight to obtain the salient features of the image.
[0032] The multi-target tracking network based on random channel adaptive attention involved in the present invention is based on a center point-based tracker, which belongs to a detection and tracking joint learning framework. The main idea is that each object is marked with a point at the center of its bounding box, namely the center point, and then the multi-target tracking is completed by identifying and associating the center point in the adjacent video frames. Among them, the identification and positioning of the target center point uses a center point-based detector. First, the detector inputs two adjacent video frames and a "heat map" used to represent the center point position of the target in the previous video frame. Then, the detector outputs the center point position of the target in the current video frame to complete the identification and positioning of each target. Then the center point-based tracker enables the detector to learn the offset of the center point of each target in the two adjacent video frames, and the offset is from the center point position of the target in the current frame to the center point position of the target in the previous frame. Finally, the trajectory association of each target can be completed by using a simple greedy algorithm based on the predicted center point offset. The center point-based tracker only associates objects in adjacent frames without reinitializing lost long-range tracking to increase speed and accuracy.
[0033] like Figure 2 As shown in the figure, the implementation principle of the center point tracker is: input the previous and next frame images, as well as the heat map rendered by the target center point tracked in the previous frame, and pass through the multi-target tracking network (i.e. Figure 2The output is the detection heat map of the current frame, the size feature map, and the offset feature map of this frame relative to the previous frame. Finally, the data can be associated through the nearest distance matching to obtain the target identification number. The target attributes obtained by the multi-target tracking network are b = (p, s, w, id), which are the center point position, width and height, confidence and identification number of the target. Among them, the detection heat map contains the feature information of the center point position of each target, the size feature map contains the feature information of the length and width of each target bounding box, and the offset feature map contains the feature information of the change of the center point position of each target between adjacent video frames.
[0034] The center-point-based tracker predicts the two-dimensional offset of the target in this frame relative to the previous frame, which is equivalent to the result of motion modeling in the tracking algorithm. It calculates the distance error between the target position of the previous frame after the offset is transformed to this frame and the target position detected in the current frame. The target data can be associated using the greedy method of minimum distance to obtain the target identification number, thereby completing the multi-target tracking task.
[0035] Figure 3 This is the overall framework of the multi-target tracking network based on random channel adaptive attention designed by the present invention. At the current time t, the network inputs the image of the current frame Previous frame image And the tracking target of the previous frame, which is represented by the result heat map of the previous frame in …represents the position and identification number of each target in the previous frame. Here, a represents the time range hyperparameter for forward retrieval, which can be adjusted manually. It means that by learning a certain time range in the forward vicinity, the anti-occlusion ability of the model can be improved, thereby strengthening the long-term tracking ability of the model. After input, after simple preprocessing, it is fused into a feature and passed to the random channel adaptive attention deep aggregation convolutional neural network. According to the output of the network, the center position, scale and offset of the target in the current frame are predicted, thereby completing the target tracking. Among them, deep aggregation refers to the features obtained by downsampling the image at different multiples, and then restoring it to the specified resolution through upsampling, and finally fusing the features at different resolutions, so as to obtain more semantic and dimensional information, improve the accuracy of feature extraction and the stability of feature fusion.
[0036] The present invention adds a random channel adaptive attention module on the basis of the existing multi-target tracking network, which is mainly to randomly let the model pay attention to the relationship between channels, hoping that the model can automatically learn the importance of different random channel features and calibrate their importance. This attention mechanism allows the model to pay more attention to channel features with large amounts of information and suppress those unimportant channel features. The module can be divided into a random gating unit and an attention module. Among them, the attention module includes three submodules: Squeeze, Excitation and Scale.
[0037] The specific implementation of the random channel adaptive attention module is as follows Figure 4 As shown. First, for a width of w, a length of H, and a number of channels of C 1 The original input feature x performs a convolution operation to obtain the feature map u that needs to be recalibrated. At this time, the number of channels of the feature map u becomes C 2 Then initialize a random identifier of the input feature map. The random identifier takes a value of 0 to 1, indicating that only a part of the channel features of the video frame are used to complete the calculation. Different values need to be set for different tasks. At the same time, the feature map u is input into the random gating unit, and the random identifier is used to indicate whether the attention module is started or not. Here, the idea of the "gate" mechanism is borrowed, and the random identifier is used to implement threshold filtering to determine whether the current input needs to be weighted by the attention mechanism. In actual application scenarios, different random identifiers should be set according to the differences in the data sets, so as to adjust the feature map into u. 1 and u 2 Two parts, feature u 2 The number of channels becomes This step is a "random" operation.
[0038] Feature map u 1 No operation is required, just skip it. Feature map u 2 The feature is compressed by the global average pooling operation, and the length is This step is a "squeeze" operation, which outputs a set of local descriptors that can express the entire image information:
[0039]
[0040] Where: H, W represent the height and width of the feature map, u c represents the cth channel of the feature map, u c (i,j) represents the pixel in the i-th row and j-th column in the c-th channel, z c is the output of the extrusion operation, F sq Represents the squeeze operation function.
[0041] Then the global features obtained by the squeezing operation are first passed through a fully connected layer to transform the dimension from Reduced to Use the ReLU function for activation, and then go through a fully connected layer to increase the dimension to the original dimension The Sigmoid activation function is used to obtain the weight coefficient of each channel (channel refers to a single-layer feature map) in the video frame feature, that is, the "excitation" operation. Its calculation formula is shown in formula (2):
[0042] s c =F ex (z c ,W)=σ(g(z c ,W))=σ(W 2 δ(W 1 z c )) (2)
[0043] Where: W 1 and W 2 is a fully connected operation, z c is the output of the squeeze operation, δ is the ReLU activation function, σ is the Sigmoid activation function, s c is the output of the excitation operation, F ex represents the excitation operation function, and W represents the overall fully connected operation.
[0044] Finally, the weight coefficient obtained by the excitation operation is multiplied by the feature map u to recalibrate the feature importance and update the feature map, which is the "weighted" operation. The calculation formula is shown in formula (3):
[0045]
[0046] Where: s c is the weight of the cth channel of the feature map, is the output of the weighted operation, F scale Represents the weighted operation function.
[0047] The random channel adaptive attention module of the present invention ensures that the feature layer α (α∈(0,1)) in the model is calibrated for importance through this method through the control of random identifiers and gate mechanisms, thereby enabling the model to have more powerful learning capabilities and extract more representative features.
[0048] The original multi-target tracking method uses the IOU (Intersection over Union) loss to evaluate the overlap between the predicted frame and the true frame. IOU has certain defects because it does not consider the situation that the predicted frame and the true frame do not intersect at all. Based on this, Zhaohui Zheng et al. proposed two evaluation criteria, DIOU (distance intersection over union) and CIOU (complete intersection over union). Among them, complete intersection over union (CIOU) is currently the best evaluation criterion and is widely used. Its formula is:
[0049]
[0050] Among them, α is a positive hyperparameter used to weigh the influence weight between aspect ratio and IOU. α is calculated as follows:
[0051]
[0052] Among them, v represents the consistency of the aspect ratio of the target bounding box, which is calculated as follows:
[0053]
[0054] Among them, ρ represents the Euclidean distance calculation function, b represents the predicted position of the target center point, and b gt Indicates the true position of the target center point, w gt represents the true width of the target bounding box, h gt represents the true height of the target bounding box, w represents the predicted width of the target bounding box, and h represents the predicted height of the target bounding box.
[0055] In the implementation of the present invention, the network input scale is set to 3×1088×608, the network output feature map scale is 272×152, and the Batch Size is set to 8. The learning rate is initialized to 0.0001 and reduced to 10% at the 25th epoch of training. The training cycle is set to 40 epochs, and the model optimizer used is Adam. In order to shorten the total time of network training, Backbone in the present invention uses a pre-trained model trained based on the COCO dataset.
[0056] Implementation examples such as Figure 5 ,After inputting the original video sequence, the bounding box of the target is drawn in the video frame and numbered as the unique ,identification of the target.
[0057] In other embodiments of the present invention, in addition to using CIOU, other loss functions such as GIOU may also be used.
[0058] Based on the same inventive concept, another embodiment of the present invention provides a multi-target tracking system based on a random channel adaptive attention mechanism implemented by the method of the present invention, which includes:
[0059] The random module is used to randomly start the spatial adaptive attention module when the features of each target are transmitted and aggregated through the gating mechanism, and to partially enhance the feature channels of multiple targets by setting random ratios;
[0060] The squeezing module is used to perform global average pooling on the feature maps of multiple targets passed in after the adaptive attention module is started, thereby completing the compression processing of the feature maps;
[0061] The excitation module is used to restore the features of each target to its original dimension using the fully connected layer, and then calculate the weights of each channel;
[0062] The weighting module is used to weight the importance of different channels of the feature map of each target by using the calculated weights to strengthen the role of significant features in the image;
[0063] The tracking module is used to output the center point of each target after weighting, predict the offset of the center point of each target between adjacent video frames, and realize multi-target tracking using a greedy association algorithm.
[0064] Based on the same inventive concept, another embodiment of the present invention provides an electronic device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.
[0065] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, CD), which stores a computer program. When the computer program is executed by a computer, it implements the various steps of the method of the present invention.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-target tracking method based on random channel adaptive attention mechanism, It is characterized in that The following steps are involved: Perform random operations, that is, when the features of each target are transmitted and aggregated, the spatial adaptive attention module is randomly activated through the gating mechanism, and the partial enhancement of the feature channels of multiple targets is achieved by setting random ratios; Perform a squeezing operation, that is, after the adaptive attention module is started, perform a global average pooling process on the feature maps of multiple targets passed in, thereby completing the compression process of the feature maps; Perform the excitation operation, that is, use the fully connected layer to restore the features of each target to the original dimension, and then calculate the weight of each channel; Perform weighted operation, that is, weight the importance of different channels of the feature map of each target by using the calculated weights to strengthen the role of significant features in the image; After weighting, the center point of each target is output, the offset of the center point of each target between adjacent video frames is predicted, and the greedy association algorithm is used to achieve multi-target tracking; The random operation includes: first, for a random number with a width of w, a length of H, and a number of channels of C 1 The original input feature x performs a convolution operation to obtain the feature map u that needs to be recalibrated. At this time, the number of channels of the feature map u becomes C 2 ; Then initialize a random identifier of the input feature map. The random identifier takes a value from 0 to 1, indicating that only part of the channel features of the video frame are used to complete the calculation. Different values need to be set for different tasks. At the same time, the feature map u is input into the random gating unit, and the random identifier is used to indicate whether the spatial adaptive attention module is started or not.
2. The method according to claim 1, It is characterized in that The random operation strengthens the learning ability of the model by setting random numbers, so that important feature information receives more attention; on this basis, the squeezing operation, the excitation operation and the weighted operation make the feature importance explicitly displayed in the weight, thereby strengthening the model's significant feature extraction ability.
3. The method according to claim 1, It is characterized in that The extrusion operation is achieved by the following formula: Among them, H and W represent the height and width of the feature map, and u c represents the cth channel of the feature map, u c (i,j) represents the pixel in the i-th row and j-th column in the c-th channel, z c is the output of the extrusion operation, F sq Represents the squeeze operation function.
4. The method according to claim 1, It is characterized in that The excitation operation includes: passing the global features obtained by the squeezing operation through a fully connected layer to reduce the dimension, activating it with a ReLU function, then passing it through a fully connected layer to increase the dimension to the original dimension, and using a Sigmoid activation function to obtain the weight coefficients of each channel in the video frame feature.
5. The method according to claim 4, It is characterized in that The calculation formula of the excitation operation is: s c =F ex (With c ,W)=σ(g(z c ,W))=σ(W 2 δ(W 1 With c )) Among them, W 1 and W 2 is a fully connected operation, z c is the output of the squeeze operation, δ is the ReLU activation function, σ is the Sigmoid activation function, s c is the output of the excitation operation, F ex represents the excitation operation function, and W represents the overall fully connected operation.
6. The method according to claim 1, It is characterized in that The calculation formula of the weighted operation is: Among them, s c is the weight of the cth channel of the feature map, is the output of the weighted operation, F scale Represents the weighted operation function.
7. A multi-target tracking system based on a random channel adaptive attention mechanism using the method of any one of claims 1 to 6, It is characterized in that include: The random module is used to randomly start the spatial adaptive attention module when the features of each target are transmitted and aggregated through the gating mechanism, and to partially enhance the feature channels of multiple targets by setting random ratios; The squeezing module is used to perform global average pooling on the feature maps of multiple targets passed in after the adaptive attention module is started, thereby completing the compression processing of the feature maps; The excitation module is used to restore the features of each target to the original dimension using the fully connected layer, and then calculate the weight of each channel; The weighting module is used to weight the importance of different channels of the feature map of each target by using the calculated weights to strengthen the role of significant features in the image; The tracking module is used to output the center point of each target after weighting, predict the offset of the center point of each target between adjacent video frames, and realize multi-target tracking using a greedy association algorithm.
8. An electronic device, It is characterized in that The invention comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Unsupervised video target segmentation method and system based on long-term and short-term time sequence characteristics
CN114973071A
Target tracking method based on gated attention mechanism and space-time memory network
CN119131085A