Improved multi-target vehicle tracking method and system based on LCBlock and CBAM attention mechanism
By using LC_Block and CBAM attention mechanisms to improve the feature extraction network in multi-target vehicle tracking, the problem of reduced detection accuracy in vehicle tracking is solved, and a more efficient and lightweight vehicle tracking method is realized, which improves detection and tracking speed and cost-effectiveness.
Patent Information
- Application Number
- CN202510280453.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
AI Technical Summary
The existing CPU-based lightweight detection model has reduced detection accuracy in multi-target vehicle tracking, making it difficult to achieve lightweight model on the basis of ensuring detection accuracy.
The multi-objective vehicle tracking method is improved by using LC_Block and CBAM attention mechanisms, and the multi-objective vehicle tracking method is improved by improving feature extraction backbone network, introducing a CBAM channel space hybrid attention mechanism, and integrating the residual structure with the CBAM attention mechanism to improve detection accuracy and efficiency.
It significantly improves detection and tracking speed, reduces model complexity and resource consumption, improves the overall cost-effectiveness of the algorithm, and provides a more efficient and lightweight solution for real-time object detection and tracking tasks.
Smart Images

Figure CN120163847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a multi-object vehicle tracking method and system improved based on the LC_Block and CBAM attention mechanisms. Background Technique
[0002] In the current development stage of deep learning, detection-based tracking technology plays a crucial role in the field of multi-object tracking. By integrating object detection and trajectory matching, this technology has the ability to automatically identify new objects and remove disappearing objects, thus adapting to modern complex multi-object tracking scenarios. However, in practical applications, the tracking effect is often directly affected by the performance of the detection algorithm.
[0003] Currently, lightweight detection models based on CPU have become a research hotspot for solving this problem. Although some existing methods such as MobileNet, ShuffleNet, and GhostNet perform well in image classification tasks, their direct application in object detection often leads to a decrease in detection accuracy. Therefore, designing a method to achieve model lightweight while ensuring detection accuracy has become a major challenge in current research. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-object vehicle tracking method and system improved based on the LC_Block and CBAM attention mechanisms to solve the problems proposed in the above background technique.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A multi-object vehicle tracking method improved based on the LC_Block and CBAM attention mechanisms, the method comprising the following steps:
[0006] a) Obtain a dataset of vehicle tracking images;
[0007] b) Improve the feature extraction backbone network, specifically adopt the MKLDNN acceleration strategy, and combine the lightweight convolution operations in MobileNetV3, as well as depthwise separable convolution, Ghost Module, and Squeeze-and-Excitation module to reduce the computational amount and improve the model performance;
[0008] c) Introduce the CBAM channel-spatial hybrid attention mechanism, which includes a channel attention module CAM and a spatial attention module SAM. By strengthening the attention in the channel and spatial dimensions of the input feature map, the position information relationship is optimized, thereby improving the detection accuracy;
[0009] d) Integrate the residual structure with the CBAM attention mechanism. Reduce the number of channels through 1×1 convolution in the lower branch of the residual structure, and perform feature extraction on the other branch. After extracting features through multiple convolutional layers, superimpose the outputs of the two branches in channels, and introduce the CBAM attention mechanism after the superposition to obtain the relationship between the position information of the vehicle target and the channels in a more efficient way, improving the tracking accuracy.
[0010] Preferably, the processing process of the channel attention module CAM includes:
[0011] Perform global max pooling and global average pooling on the input feature map respectively to obtain two one-dimensional vectors;
[0012] Send the two vectors into the multi-layer perceptron neural network MLP respectively;
[0013] Merge the outputs of the MLP through element-wise summation to obtain the final channel attention map.
[0014] Preferably, the processing process of the spatial attention module SAM includes:
[0015] Multiply the channel attention map output by the CAM module with the input initial feature element-wise to obtain the input initial feature map of the SAM;
[0016] Perform global average pooling and global max pooling on the input initial feature map;
[0017] Perform a channel-based concatenation operation on the two obtained feature maps;
[0018] After a 7×7 convolution operation, generate spatial attention features through Sigmoid activation;
[0019] Multiply the feature with the input feature to obtain the finally generated feature.
[0020] Preferably, the fusion method of the residual structure and the CBAM attention mechanism is specifically:
[0021] Introduce the CBAM attention mechanism after the concat operation of the residual structure, and improve the tracking accuracy on the CPU without adding high-cost equipment.
[0022] Preferably, the method is applicable to the vehicle multi-object tracking scenario with high demand for computing resources. By improving the LC_Block and the CBAM attention mechanism, the accuracy and efficiency of vehicle tracking are effectively improved.
[0023] The multi-object vehicle tracking system improved based on the LC_Block and the CBAM attention mechanism is applied to the multi-object vehicle tracking method improved based on the LC_Block and the CBAM attention mechanism. The system includes:
[0024] A data collection module that obtains a dataset of vehicle tracking images;
[0025] A network improvement module that improves the feature extraction backbone network. Specifically, it adopts the MKLDNN acceleration strategy and combines lightweight convolution operations in MobileNetV3, as well as depthwise separable convolution, Ghost Module, and Squeeze-and-Excitation module to reduce the computational amount and improve the model performance;
[0026] A position relationship optimization module that introduces the CBAM channel-spatial hybrid attention mechanism. This mechanism includes a channel attention module CAM and a spatial attention module SAM. By strengthening the attention in the channel and spatial dimensions of the input feature map, it optimizes the position information relationship to improve the detection accuracy;
[0027] A channel stacking module that fuses the residual structure with the CBAM attention mechanism. In the lower branch of the residual structure, the number of channels is reduced through 1×1 convolution, and the other branch performs feature extraction. After extracting features through multiple convolutional layers, the outputs of the two branches are channel-stacked, and the CBAM attention mechanism is introduced after the stacking to obtain the relationship between the position information and channels of the vehicle target in a more efficient way and improve the tracking accuracy.
[0028] Preferably, the processing process of the channel attention module CAM in the position relationship optimization module includes:
[0029] Performing global max pooling and global average pooling on the input feature map respectively to obtain two one-dimensional vectors;
[0030] Feeding the two vectors into a multi-layer perceptron neural network MLP respectively;
[0031] Merging the outputs of the MLP through element-wise summation to obtain the final channel attention map.
[0032] Preferably, the processing process of the spatial attention module SAM in the position relationship optimization module includes:
[0033] Element-wise multiplying the channel attention map output by the CAM module with the input initial feature to obtain the input initial feature map of the SAM;
[0034] Performing global average pooling and global max pooling on the input initial feature map;
[0035] Performing a channel-based concatenation operation on the two obtained feature maps;
[0036] After a 7×7 convolution operation, generating spatial attention features through Sigmoid activation;
[0037] Multiply this feature with the input feature to obtain the finally generated feature.
[0038] Preferably, the fusion method of the residual structure in the channel stacking module and the CBAM attention mechanism is specifically as follows:
[0039] Introduce the CBAM attention mechanism after the concat operation of the residual structure, and improve the tracking accuracy on the CPU without adding high-cost devices.
[0040] Preferably, the system is applicable to the vehicle multi-object tracking scenario with high demand for computing resources, and effectively improves the accuracy and efficiency of vehicle tracking by using the improvements of the LC_Block and the CBAM attention mechanism.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] The multi-object vehicle tracking method and system based on the improvement of the LC_Block and the CBAM attention mechanism proposed by the present invention, by precisely selecting and integrating the CBAM and the LCNet module, not only achieves significant improvements in multiple key performance indicators, especially in terms of improving the detection and tracking speed and reducing the model complexity. These improvements not only enhance the adaptability and robustness of the model in practical applications, but also greatly reduce the resource consumption and improve the overall cost performance of the algorithm, providing a more efficient and lightweight solution for real-time object detection and tracking tasks. It should be noted that the number of parameters of the algorithm is reduced by 78.3MB, which means that the network model is smaller and the cost performance is improved.
[0043] By introducing the CBAM residual module, the present invention significantly enhances the performance of the model in detection and tracking tasks. Specifically, the channel attention mechanism of the CBAM residual module effectively focuses on the key channel information in the image features, improves the discriminability of the features, and thus greatly improves the recall rate (R) and the multi-object tracking accuracy (MOTA). The spatial attention mechanism captures the spatial position information in the image, further refines the target localization, and significantly enhances the detection and tracking accuracy of the model.
[0044] In view of the problem that the detection rate, tracking rate and model volume have not been significantly improved after the introduction of the CBAM module, the present invention creatively integrates the LCNet module. By adopting the H-Swish activation function and optimizing the position of the SE module to the end of the network, the LCNet module effectively improves the operation efficiency of the network, and the detection rate and tracking rate are significantly improved by 12.5% and 17% respectively. In addition, by adding a 1×1 convolution with 1280 dimensions after the global average pooling (GAP) layer of the network, not only the fitting ability of the model is enhanced, but also the network structure is greatly simplified. The number of model parameters decreases by 79.61MB, only 10% of the original model, significantly reducing the model volume and improving the practicability and deployment flexibility of the model. Description of the Drawings
[0045] Figure 1 Schematic diagram of the CBAM model structure introduced by the present invention;
[0046] Figure 2 Schematic diagram of the CBAM residual structure after improvement of the present invention. Detailed Implementation Manner
[0047] In order to clearly and completely describe the objectives, technical solutions of the present invention and make the advantages more clear, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0048] Embodiment 1, please refer to Figures 1 to 2 , the present invention provides a technical solution: a multi-object vehicle tracking method improved based on the LC_Block and CBAM attention mechanisms. The method mainly includes two parts: a structure reparameterization backbone network and a spatial attention mechanism to improve detection accuracy. The main steps are as follows:
[0049] Obtain the dataset UA-DETRAC of vehicle tracking images.
[0050] Improve the feature extraction backbone network and utilize the MKLDNN (Intel hardware-optimized deep learning inference library) acceleration strategy. First, lightweight convolutional operations in MobileNetV3 are adopted, such as Depthwise Separable Convolutions, Ghost Module, and Squeeze-and-Excitation (SE) module, to reduce the computational amount and improve the model performance. The specific measures are as follows: the model structure of the network is divided into DW / h-swish module, PW / h-swish module, GAP module, and FC module. The combination of Depthwise Convolution and Pointwise Convolution with the h-swish activation function is used to enhance the non-linear characteristics while reducing the computational cost, thereby making the network easier to converge. The global average pooling layer is used to reduce the dimension of the feature map by taking the average of all values in the feature map. It is used to reduce the model parameters and improve the robustness of the model to spatial translation. The Fully Connected Layers are used to accelerate convergence and alleviate the vanishing gradient problem. The combination with the h-sigmoid activation function is used to generate probability outputs for final regression and classification while reducing the computational cost.
[0051] Introduce the CBAM (Convolutional Block Attention Module) channel-spatial hybrid attention mechanism to optimize the structure of the position information relationship, such as Figure 1。This module contains two sub-modules, the Channel Attention Module (CAM) and the Spatial Attention Module (SAM), and can enhance attention in both the spatial and channel dimensions. The specific process of CAM is as follows: First, for the input feature map F (H×W×C, where H, W, and C represent the height, width, and channels of the initial feature map respectively), it is passed through global max pooling and global average pooling to obtain two one-dimensional vectors of H×W×1. Then, these two one-dimensional vectors are respectively fed into a multi-layer perceptron neural network (MLP). The channel attention map calculated by the MLP is merged through element-wise summation to obtain the final channel attention map. SAM pays more attention to the spatial position information of image features, which also makes up for the deficiency of CAM. In spatial attention, the channel attention map finally obtained by CAM is multiplied element-wise with the input initial feature to obtain the input initial feature map of SAM. The input feature map is subjected to global average pooling and global max pooling based on channels to obtain two feature maps. Then, these two feature maps are concatenated based on channels, and then passed through a 7×7 convolution operation to reduce the dimension to 1 channel, that is, H×W×1. After passing through the Sigmoid activation, the spatial attention feature is generated. Finally, this feature is multiplied with the input feature of this module to obtain the finally generated feature. Step four is to fuse the residual structure with the CBAM attention mechanism. The lower branch of the residual structure reduces the number of channels through a 1×1 convolution; the other branch containing the residual structure mainly performs feature extraction. First, it undergoes a 1×1 convolution operation for dimension reduction, and then passes through a 1×1 convolution and a 3×3 convolution to extract features. Finally, the extracted feature information undergoes a 1×1 convolution operation for dimension increase and is channel-added to the left branch for output. We introduce the convolutional attention mechanism after its concat to obtain the relationship between the position information of the vehicle target and the channels in a more efficient way. Without introducing high-cost equipment, the tracking accuracy can be improved on the CPU. The improved CBAM residual structure diagram is as shown in Figure 2 。
[0052] The overall network structure table of the algorithm is as follows:
[0053]
[0054]
[0055] Example 2, based on Example 1, a multi-target vehicle tracking system improved based on the LC_Block and CBAM attention mechanisms is proposed, which is applied to the multi-target vehicle tracking method improved based on the LC_Block and CBAM attention mechanisms. The system includes:
[0056] Data collection module, which obtains a dataset of vehicle tracking images;
[0057] Network improvement module, which improves the feature extraction backbone network. Specifically, it adopts the MKLDNN acceleration strategy and combines lightweight convolutional operations in MobileNetV3, as well as depthwise separable convolution, Ghost Module, and Squeeze-and-Excitation module to reduce the computational amount and improve the model performance;
[0058] Position relationship optimization module, which introduces the CBAM channel-spatial hybrid attention mechanism. This mechanism includes a channel attention module CAM and a spatial attention module SAM. By strengthening the attention in the channel and spatial dimensions of the input feature map, it optimizes the position information relationship, thereby improving the detection accuracy;
[0059] Channel stacking module, which fuses the residual structure with the CBAM attention mechanism. In the lower branch of the residual structure, the number of channels is reduced through 1×1 convolution, and the other branch performs feature extraction. After extracting features through multiple convolutional layers, the outputs of the two branches are channel-stacked, and the CBAM attention mechanism is introduced after the stacking to obtain the relationship between the position information and channels of the vehicle target in a more efficient way, improving the tracking accuracy.
[0060] The processing process of the channel attention module CAM in the position relationship optimization module includes: performing global max pooling and global average pooling on the input feature map respectively to obtain two one-dimensional vectors; sending the two vectors into a multi-layer perceptron neural network MLP respectively; merging the outputs of the MLP through element-wise summation to obtain the final channel attention map. Preferably, the processing process of the spatial attention module SAM in the position relationship optimization module includes: multiplying the channel attention map output by the CAM module element-wise with the input initial feature to obtain the input initial feature map of the SAM; performing global average pooling and global max pooling on the input initial feature map; performing a channel-based concatenation operation on the two obtained feature maps; after a 7×7 convolution operation, generating a spatial attention feature through Sigmoid activation; multiplying this feature with the input feature to obtain the finally generated feature.
[0061] The fusion method of the residual structure and the CBAM attention mechanism in the channel stacking module is specifically: introducing the CBAM attention mechanism after the concat operation of the residual structure, and improving the tracking accuracy on the CPU without adding high-cost devices.
[0062] The system is applicable to vehicle multi-object tracking scenarios with high requirements for computing resources. By using the improvements of the LC_Block and CBAM attention mechanisms, it effectively improves the accuracy and efficiency of vehicle tracking.
[0063] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An improved multi-target vehicle tracking method based on LC_Block and CBAM attention mechanism, characterized by: The method comprises the following steps: a) Obtain a dataset of vehicle tracking images; b) Improve the feature extraction backbone network by using the MKLDNN acceleration strategy and combining the lightweight convolution operation in MobileNetV3, as well as the depthwise separable convolution, Ghost Module, and Squeeze-and-Excitation modules to reduce the amount of computation and improve model performance; c) Introducing the CBAM channel-space hybrid attention mechanism, which includes the channel attention module CAM and the spatial attention module SAM. By strengthening the attention of the channel and spatial dimensions of the input feature map, the position information relationship is optimized, thereby improving the detection accuracy; d) The residual structure is integrated with the CBAM attention mechanism. The number of channels is reduced by 1×1 convolution in the lower branch of the residual structure, and the other branch performs feature extraction. After extracting features through multiple convolutional layers, the outputs of the two branches are superimposed on the channels, and the CBAM attention mechanism is introduced after superposition to obtain the relationship between the position information of the vehicle target and the channel in a more efficient way, thereby improving the tracking accuracy.
2. The multi-target vehicle tracking method based on the improved LC_Block and CBAM attention mechanism according to claim 1 is characterized in that: The processing of the channel attention module CAM includes: Perform global maximum pooling and global average pooling on the input feature map to obtain two one-dimensional vectors; Send the two vectors into the multi-layer perceptron neural network MLP respectively; The outputs of the MLP are combined by element-wise summation to obtain the final channel attention map.
3. The multi-target vehicle tracking method based on the improved LC_Block and CBAM attention mechanism according to claim 2 is characterized in that: The processing of the spatial attention module SAM includes: Multiply the channel attention map output by the CAM module by the input initial feature element by element to obtain the input initial feature map of SAM; Perform global average pooling and global maximum pooling on the input initial feature map; Perform channel-based concatenation on the two obtained feature maps; After a 7×7 convolution operation, the spatial attention feature is generated through Sigmoid activation; Multiply this feature with the input feature to get the final generated feature.
4. The multi-target vehicle tracking method based on the improved LC_Block and CBAM attention mechanism according to claim 1, characterized in that: The specific fusion method of the residual structure and the CBAM attention mechanism is as follows: The CBAM attention mechanism is introduced after the concat operation of the residual structure to improve the tracking accuracy on the CPU without adding high-cost equipment.
5. The multi-target vehicle tracking method based on the improved LC_Block and CBAM attention mechanism according to claim 1, characterized in that: The method is suitable for vehicle multi-target tracking scenarios with high computing resource requirements. It uses the improvements of LC_Block and CBAM attention mechanisms to effectively improve the accuracy and efficiency of vehicle tracking.
6. The improved multi-target vehicle tracking system based on LC_Block and CBAM attention mechanism is applied to the improved multi-target vehicle tracking method based on LC_Block and CBAM attention mechanism as described in any one of claims 1 to 5, characterized in that: The system comprises: A data collection module, which obtains a dataset of vehicle tracking images; The network improvement module improves the feature extraction backbone network by adopting the MKLDNN acceleration strategy and combining the lightweight convolution operation in MobileNetV3, as well as the depthwise separable convolution, Ghost Module, and Squeeze-and-Excitation modules to reduce the amount of computation and improve model performance. The position relationship optimization module introduces the CBAM channel-space hybrid attention mechanism, which includes the channel attention module CAM and the spatial attention module SAM. It optimizes the position information relationship by strengthening the attention of the channel and spatial dimensions of the input feature map, thereby improving the detection accuracy; The channel superposition module integrates the residual structure with the CBAM attention mechanism. The lower branch of the residual structure reduces the number of channels through 1×1 convolution, and the other branch performs feature extraction. After extracting features through multiple convolutional layers, the outputs of the two branches are superimposed on the channels, and the CBAM attention mechanism is introduced after superposition. This can obtain the relationship between the position information of the vehicle target and the channel in a more efficient way, thereby improving the tracking accuracy.
7. The improved multi-target vehicle tracking system based on LC_Block and CBAM attention mechanism according to claim 6, characterized in that: The processing process of the channel attention module CAM in the position relationship optimization module includes: Perform global maximum pooling and global average pooling on the input feature map to obtain two one-dimensional vectors; Send the two vectors into the multi-layer perceptron neural network MLP respectively; The outputs of the MLP are combined by element-wise summation to obtain the final channel attention map.
8. The multi-target vehicle tracking system based on LC_Block and CBAM attention mechanism improvement according to claim 6, characterized in that: The processing process of the spatial attention module SAM in the position relationship optimization module includes: Multiply the channel attention map output by the CAM module by the input initial feature element by element to obtain the input initial feature map of SAM; Perform global average pooling and global maximum pooling on the input initial feature map; Perform channel-based concatenation on the two obtained feature maps; After a 7×7 convolution operation, the spatial attention feature is generated through Sigmoid activation; Multiply this feature with the input feature to get the final generated feature.
9. The multi-target vehicle tracking system based on LC_Block and CBAM attention mechanism improvement according to claim 6, characterized in that: The fusion method of the residual structure in the channel superposition module and the CBAM attention mechanism is as follows: The CBAM attention mechanism is introduced after the concat operation of the residual structure to improve the tracking accuracy on the CPU without adding high-cost equipment.
10. The multi-target vehicle tracking system based on LC_Block and CBAM attention mechanism improvement according to claim 6, characterized in that: The system is suitable for vehicle multi-target tracking scenarios with high computing resource requirements. It uses improvements to the LC_Block and CBAM attention mechanisms to effectively improve the accuracy and efficiency of vehicle tracking.