A Traffic Sign Detection Method Based on Transformer and Cross-Dimensional Attention
By introducing the Transformer module and the Cross-dimensional Attention Module (ECSI) into the traffic sign detection model, the problem of insufficient network receptive field and small target characteristics being flooded with background information is solved, and high-precision traffic sign detection is achieved.
Patent Information
- Application Number
- CN202211583330.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-09
AI Technical Summary
The existing traffic sign detection model cannot effectively learn the long-range dependence between features in the initial stage, and the features of small targets are easily overwhelmed by complex background information during the feature fusion stage, resulting in poor detection results.
The Transformer module is used to expand the early receptive field of the network, combine the cross-dimensional attention module (ECSI) to enhance the communication and connection between features, and realize multi-scale object detection through the LG-YOLOv5 feature fusion network and detection head design.
The accuracy and recall rate of traffic sign detection are improved, and high-precision detection is achieved with small parameters, which is better than existing algorithms, especially the mAP on the TT100K dataset reaches 90.5%.
Smart Images

Figure CN115830575B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and relates to the improvement of target detection models, image target detection and simulation implementation. Background Art
[0002] As one of the basic tasks of computer vision, the main task of target detection is to locate the targets of interest in the input image and determine the category to which each target belongs. It has been applied in various scenarios, such as traffic sign detection and autonomous driving. In recent years, with the continuous development of deep learning, traffic sign detection research has received extensive attention. Since the images collected by vehicles are extremely vulnerable to factors such as light and weather during the image acquisition process, resulting in unclear or distorted images; in addition, traffic signs often account for less than 1% of the entire image, making the detection and recognition of traffic signs more challenging than ordinary target detection tasks.
[0003] Currently, there is relatively little research on algorithms specifically for traffic sign detection. Directly using general target detection methods is likely to cause missed detection of small targets, and the effect is not good. Therefore, the present invention designs two modules to enhance traffic sign detection. The specific method is as follows: to solve the problem that the shallow layer of the network lacks sufficient effective receptive fields, a Transformer module is designed to strengthen the expression of shape features and capture the dependencies between pixels at a long distance, enabling the network to obtain a global receptive field at the initial stage; to solve the problem that while enhancing one-dimensional features, the attention mechanism will lose a large amount of information in other dimensions, a cross-dimensional attention module is designed to improve the cross-dimensional interaction ability of the model in channels and space, strengthen the expression of small target features and suppress redundant background features. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] Aiming at the deficiencies of the prior art, the present invention provides a traffic sign detection method based on Transformer and cross-dimensional attention. It solves the problems that the current traffic sign detection model cannot learn rich long-range dependencies between features at the initial stage and the effective features of small targets are easily submerged in complex background information during the feature fusion stage.
[0006] (2) Technical Solutions
[0007] To achieve the above objectives, the present invention proposes a traffic sign detection method based on Transformer and cross-dimensional attention. This network can model rich local and global features to enhance the communication and connection between information and generate more discriminative features. First, in order to enable the network to have a larger receptive field in the initial stage, a Transformer module is embedded in the backbone network to capture the dependencies between distant pixels. Second, in order to reduce the information loss caused during the use of the attention mechanism, the present invention proposes a cross-dimensional attention module (Enchancing Channel and Spaital Interaction, ECSI), which is composed of a low-order global attention module (Low-order Global Attention, LGA) and a high-order global attention module (High-order Global Attention, HGA). LGA aims to learn the importance of different features under the local receptive field, while HGA aims to learn the importance of different features under the global receptive field. Finally, four feature maps of different scales are obtained, and object detection is performed on these four feature maps of different scales respectively to obtain the final detection result.
[0008] A traffic sign detection method based on Transformer and cross-dimensional attention according to the present invention includes the following steps:
[0009] S1. First, a Transformer module is added to the backbone network, aiming to expand the effective receptive field in the initial stage of the network to learn the semantic relationships between different features. To reduce the computational complexity, the present application only embeds the Transformer module in the sixth and ninth layers of the backbone network, and then four feature maps of different levels are obtained;
[0010] S2. Inject the four feature maps of different levels obtained in S1, namely the output features of the 2nd, 4th, 6th, and 9th layers, into the feature fusion network. In the feature fusion network, LG-YOLOv5 uses the CARAFE upsampling operator to reduce information loss, and adds the cross-dimensional attention ECSI module to the top-down and bottom-up information propagation paths to selectively inject semantic and detailed information to avoid the influence of interfering features. Then, a Transformer module is added at the position of the output detection head to learn the context relationships between different scale features, and finally four feature maps P2, P3, P4, and P5 are obtained for detecting the target position.
[0011] S3. Then, the four feature maps of different resolutions obtained in S2 are respectively input into a convolutional layer to predict targets of different scales. P2, P3, P4, and P5 respectively represent the prediction results on the feature maps downsampled by 4, 8, 16, and 32 times.
[0012] (III) Advantageous Effects
[0013] The present invention provides a traffic sign detection method based on Transformer and cross-dimensional attention. It has the following advantageous effects:
[0014] By using the Transformer-based module, the present invention solves the problems of small receptive fields in the initial stage of the network and insufficient learning of context relationships. The cross-dimensional attention module solves the problem that small targets have few effective features and are easily diluted, highlighting the effective features.
[0015] The model proposed by the present invention has good performance. Experimental results on the TT100K dataset show that compared with existing advanced algorithms, the algorithm of the present invention achieves the best trade-off between the number of parameters and accuracy. With only 26M parameters, the mAP reaches 90.5%. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is the overall framework structure diagram of the present invention;
[0017] Figure 2 is the structure diagram of the Transformer module constructed by the present invention;
[0018] Figure 3 is the structure diagram of the cross-dimensional attention module constructed by the present invention;
[0019] Figure 4 is the visualization result diagram of the heat map of the Transformer module of the present invention;
[0020] Figure 5 is the visualization result diagram of the attention of the cross-dimensional attention module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The technical method in the present invention will be clearly and completely described below with reference to the accompanying drawings. A traffic sign detection method based on Transformer and cross-dimensional attention, the specific implementation steps are as follows:
[0022] (S1): Add a detection head and a feature fusion method.
[0023] To help locate small targets, the present invention adds a new branch specifically for detecting small targets, and adds the ECSI module to the top-down and bottom-up feature fusion paths to selectively inject semantic and detailed information to avoid the influence of interfering features. The low-level features and high-level features in the present invention respectively select the 2nd, 4th, 6th, and 9th layer features in the LG-YOLOv5 backbone network to fuse multi-scale features.
[0024] (S2): Design the Transformer module.
[0025] In order to capture the dependencies between long-distance features while aggregating local features, so as to learn the semantic relationships between different targets and improve the accuracy of object detection, this application proposes the LGIM module. As Figure 2 (a) shows, the LGIM module consists of two branches. The left branch is used to aggregate local features, and the right branch is used to capture the dependencies between long-distance features. The left branch consists of a 3×3 convolution, BN, and the SiLu activation function. After passing through a 1×1 convolution, a BN layer, and the SiLu activation function, the right branch inputs the Mixing Local Global Transformer (MLGT) to learn the inductive bias and the dependencies between long-distance pixels. The feature maps obtained from the input features through the two branches are first concatenated in the channel direction, and then passed through a 1×1 convolution operation with BN and the SiLu activation function to enable full interaction and fusion of global and local information.
[0026] Since the self-attention module in the Transformer only models the spatial relationships and lacks the connection between channels. Therefore, this application proposes the channel-enhanced self-attention module CE-(S)WSA to learn the channel-level dynamic attention weights and strengthen the channel modeling ability in the self-attention mechanism. As Figure 2 (c) shows, CE-(S)WSA mainly consists of two branches. One branch generates a similarity matrix spatially, and the other branch generates the channel-level dynamic weights. Among them, the calculation process of the similarity matrix is similar to that of the Swin Transformer. The generation process of the channel-level dynamic attention weights is as follows: The input feature x first passes through a 3×3 depth convolution to model the cross-window feature relationships; then the global average pooling is used to generate the dynamic weights for each channel; then two 1×1 convolutions are used to fully learn the dependencies between different channels; finally, the channel-level dynamic weight vector is obtained through the Sigmoid activation function. The combination of the channel-level dynamic attention weights and the Transformer is completed using matrix multiplication operations. The calculation process of CE-(S)WSA is as shown in formulas (1-5) below:
[0027] X = GAP(DConv 3×3 (x)) (1)
[0028] Y = GeLu(BN(Conv(X))) (2)
[0029] Z = Sigmoid(BN(Conv(Y))) (3)
[0030]
[0031]
[0032] Among them, DConv 3×3 refers to a 3×3 depth convolution, GAP represents global average pooling, Conv represents a 1×1 convolution, Z represents the generated channel attention weight, and V represents the feature vector after embedding the channel weight. refers to the Reshape operation, and RPC is relative position encoding. Compared with the way of self-attention external embedded convolution, CE-(S)WSA not only alleviates the problem of insufficient information interaction between windows in the window-based self-attention mechanism, but also endows the self-attention module with the ability of channel modeling, so as to learn richer global semantic information.
[0033] (S3): Design a cross-dimensional attention module.
[0034] To strengthen the interaction of features at different levels in the channel and spatial dimensions, this application proposes the ECSI module. As Figure 3 (c) shows, the ECSI module consists of a Low-order Global Attention (LGA) structure and a High-order Global Attention (HGA) structure. The LGA aims to learn the importance of different features under the local receptive field, while the HGA aims to learn the importance of different features under the global receptive field.
[0035] As Figure 3 (c) shows, the construction process of the LGA is as follows. The input feature x first passes through a 3×3 depth convolution to aggregate local spatial information, and then two fully connected layers are used to achieve cross-dimensional interaction of local features in the spatial and channel dimensions. Among them, the first MLP compresses the number of feature channels from C to C / 4 for global feature encoding; the second fully connected layer restores the number of channels to C. This way of first reducing the dimension and then increasing the dimension not only improves the fitting ability of the network but also greatly reduces the number of parameters brought by the fully connected layer. Subsequently, the weight size of different features under the local receptive field is obtained through the Sigmoid activation function. Finally, the low-order cross-dimensional global attention weight is multiplied element-wise with the input feature to obtain the enhanced feature map x'.
[0036] As Figure 3As shown in (c), the output feature x' of the LGA serves as the input feature of the HGA. To enhance the interaction of global information while reducing the number of parameters, the HGA mainly uses grouped convolution and channel recombination modules. To obtain high-order spatial dependence relationships, the feature x' is first evenly divided by channels and input into two branches. One of the branches uses a 7×7 depth convolution to aggregate spatial information under a larger receptive field, and then element-wise multiplied with the feature of the other branch. After obtaining a feature map containing first-order spatial dependence relationships, it is element-wise multiplied with it again to obtain a feature map containing second-order spatial dependence relationships, and then the number of channels is adjusted through a 1×1 convolution. To obtain the importance of features at different scales under the global receptive field, this application sends the feature containing high-order spatial dependence relationships into two consecutive 5×5 grouped convolutions to learn global context information, and then obtains the importance of features at different scales of the global receptive field through the Sigmoid activation function. Since grouped convolution hinders the ability to model channels within different groups, this application uses a channel recombination operation to enhance the channel information interaction between groups. Channel recombination realizes the information interaction between channels in a parameter-free manner by rearranging the channel order. Finally, the global cross-dimensional attention map obtained after passing through the channel recombination module is element-wise multiplied with the original input feature x' to obtain a feature containing second-order spatial dependence relationships and cross-dimensional interaction information at different levels. Compared with the SE and CBAM attention modules in Figure 3 (a) and Fig. (b), the ECSI not only retains important semantic and location information, but also aggregates important features at different levels to model global context relationships.
[0037] The effects of the present invention will be described in detail below in combination with experimental data and visualization heatmaps.
[0038] Table 1 compares the detection speed and accuracy of the method proposed in the present invention with other methods on the TT100K dataset. It can be seen from the experimental results in Table 1 that compared with classical methods such as Faster-RCNN, SSD, and YOLOv3, the accuracy of the method in this application has increased by 27%-40%. Compared with the baseline network YOLOv5m, the mAP of the method in this application has increased by 8.5 while only adding 4.8M parameters. Compared with the latest detection algorithms DE-DETR and YOLOX, the accuracy of the method in this application has increased by 26% and 12% respectively, and the convergence speed has increased by 2 times and 5 times respectively. Compared with the CAB algorithm, the method in this application greatly surpasses the traditional CNN model in terms of accuracy. Compared with the TSR-SA detection algorithm, the method in this application does not use an additional dataset and achieves higher accuracy with fewer parameters. Compared with the latest MDCOD algorithm, the detection speed of the method in this application is ten times faster than it, and the accuracy only drops by 2% while reducing 24% of the parameters. In summary, the method in this application achieves the best trade-off between the number of parameters and accuracy compared with various advanced detection algorithms. Under the condition that the number of parameters is only 26M, the mAP reaches 90.5%.
[0039] Table 1 Comparison with Advanced Methods on the TT100K Dataset
[0040]
[0041] Figure 4 Illustrates the effectiveness of the LGIM module in the present invention. Comparing Figure 4 the middle column and the right column, it can be seen that after passing through the LGIM module, the network can effectively focus on all traffic signs of different scales and shapes in the figure. The network with the LGIM module added can focus on the complete area of each traffic sign, rather than just focusing on the local area of the target, and eliminates the high response values of other background features outside the target area. Thus, it can be seen that the LGIM module proposed in this application enables the network to have a larger effective receptive field, can adaptively learn the dependence relationship between different scale targets, and therefore greatly improves the recall rate and accuracy of the network.
[0042] Figure 5 Illustrates the effectiveness of the ECSI module in the present invention. The visualization of the feature map corresponding to Baseline is as Figure 5 shown in the middle column. The network not only has a high activation value for the target area, but also has a certain degree of activation in the background area within the red frame. The feature distribution is relatively messy and prone to false detection. The features improved by the ECSI module are as Figure 5As shown in the right column, the model only has a high response to the target area, and the interference features in the background area are significantly suppressed. It can be seen that the ECSI module proposed in this application enables the network to have the ability to cross-dimensionally interact on different hierarchical features, can efficiently learn the importance of features at different scales, so that the network can accurately focus on the effective area features and suppress the interference features, effectively solving the problem that small target information in traffic sign detection is easily submerged by complex features.
[0043] The present invention proposes a local and global information interaction module and an enhanced channel and spatial interaction module. Among them, the local and global information interaction module fully interacts the local features and the global features, increases the network receptive field, and improves the global semantic expression ability of the shallow features. The enhanced channel and spatial interaction module enhances the information interaction between the channel and spatial dimensions at different levels, effectively alleviating the problem that the effective features of small targets are interfered and diluted by background information. A large number of experiments on TT100K show that compared with the current advanced traffic sign detection algorithms, the method of this application achieves the best trade-off between parameters and accuracy.
[0044] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A traffic sign detection method based on Transformer and cross-dimensional attention, characterized in that: It includes the following steps: S1. First, add the LGIM module based on Transformer to its backbone network and Neck network of YoLov5, and add the cross-dimensional attention ECSI module to the Neck network of YoLov5 to form the LG-YOLOv5 model. Embed the LGIM module based on Transformer only in the sixth and ninth layers of the backbone network of YoLov5, and then obtain four feature maps at different levels; The LGIM module consists of two branches. The left branch is used to aggregate local features, and the right branch is used to capture the dependencies between long-range features. The left branch consists of a 3×3 convolution, BN, and SiLu activation function. After passing through a 1×1 convolution, BN layer, and SiLu activation function, the right branch inputs a local-global hybrid Transformer to learn the inductive bias and the dependencies between long-range pixels. The feature maps obtained from the two branches of the input features are first concatenated in the channel direction, and then passed through a 1×1 convolution operation with BN and SiLu activation functions to fully interact and fuse the global and local information; S2. Inject the four feature maps at different levels obtained in S1, that is, the output features of the 2nd, 4th, 6th, and 9th layers, into the feature fusion network. In the feature fusion network, LG-YOLOv5 uses the CARAFE upsampling operator to reduce information loss, and adds the cross-dimensional attention ECSI module to the top-down and bottom-up information propagation paths to selectively inject semantic and detailed information to avoid the influence of interfering features; Then, add the LGIM module at the position of the output detection head to learn the context relationship between features at different scales, and finally obtain four feature maps P2, P3, P4, and P5 for detecting the target position; The cross-dimensional attention ECSI module consists of a low-order global attention LGA structure and a high-order global attention HGA structure; S3. Then input the four feature maps with different resolutions obtained in S2 into a convolutional layer to predict targets at different scales. P2, P3, P4, and P5 respectively represent the prediction results on the feature maps downsampled by 4, 8, 16, and 32 times.
2. The traffic sign detection method based on Transformer and cross-dimensional attention according to claim 1, characterized in that: Channel attention is embedded in the Transformer structure, which has the ability to model spatial and channel relationships simultaneously.
3. A traffic sign detection method based on Transformer and cross-dimensional attention according to claim 1, characterized in that: The cross-dimensional attention fuses spatial and channel information and retains the important information of each dimension.