Traffic Sign Detection Method and System Based on Class Pooling Connection Attention Fusion
By introducing pool-like connected attention fusion technology in traffic sign detection, using the scalable channel attention module and the multi-head self-attention module, the detection problems of the existing technology under light and weather changes are solved, and more efficient and accurate traffic sign detection is achieved.
Patent Information
- Application Number
- CN202311828056.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-12-28
AI Technical Summary
The existing traffic sign detection methods are difficult to accurately detect under factors such as light changes, bad weather and motion blur, and traditional methods rely on manual feature extraction and are inefficient.
A traffic sign detection method based on pooled connected attention fusion is proposed. By inserting the scalable channel attention module into the Transformer, combining the multi-head self-attention module and patch merging operations, a hierarchical feature extraction network is constructed to realize multi-scale feature extraction and fusion.
Improve the accuracy and efficiency of traffic sign detection, especially under different light and weather conditions, the detection accuracy is improved by about 7%, while reducing computing complexity and memory consumption.
Smart Images

Figure CN117541903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of traffic sign detection and deep learning, and particularly to a traffic sign detection method and system based on class pooling connection attention fusion. Background Art
[0002] Traffic sign detection, as one of the key technologies for autonomous driving and high-definition map environment perception, is of great significance for providing road information judgment and real-time safety warning for vehicles. Due to different road conditions and natural environments, the results of traffic sign detection are still restricted by many factors such as light changes, bad weather, and motion blur, which greatly increases the difficulty of this task.
[0003] Most traditional traffic sign detection methods rely on manually extracting features from color information and geometric shapes. However, problems such as the proportion change and occlusion of the traffic sign area during the transmission of the sensor in motion have hindered the practical application of these methods.
[0004] In order to balance accuracy and efficiency, most advanced object detection algorithms have started to use deep convolutional neural networks instead of manual feature extraction. Although the classic two-stage detection model has high detection accuracy, its complex structure leads to low detection efficiency. Compared with the two-stage model, the single-stage model has a relatively simple structure, so its detection efficiency is higher, but its detection accuracy is not satisfactory.
[0005] Recently, new models based on Transformer have shown that the end-to-end standard transformer has achieved encouraging results in various computer vision tasks and has quickly become the backbone model. However, Transformer regards the image as a sequence and lacks the ability to obtain channel dimension information when modeling visual features in local windows and scale transformation. As the network depth deepens, the information between each channel is gradually lost, and the global information is also lost. Therefore, Transformer cannot be directly applied to complex traffic sign detection tasks. Summary of the Invention
[0006] In view of the above situation, the main purpose of the present invention is to propose a traffic sign detection method and system based on class pooling connection attention fusion to solve the above technical problems.
[0007] The present invention proposes a traffic sign detection method based on class pooling connection attention fusion, and the method includes the following steps:
[0008] Step 1: Insert the scalable channel attention module into the Transformer based on the multi-head self-attention module and serialize it with the multi-head self-attention module to obtain a dual attention module. Package the PM module, the PC module, and at least one pair of dual attention modules in a serial manner into a class pooling connection attention module, and use several class pooling connection attention modules to construct a hierarchical feature extraction network;
[0009] Step 2: Preprocess the traffic image to obtain a two-dimensional image, encode the two-dimensional image into an initial one-dimensional sequence, and input it into the hierarchical feature extraction network;
[0010] Step 3: Use the scalable channel attention module to extract features from the input features to obtain a channel attention feature map, use the multi-head self-attention module to extract features from the channel attention feature map to obtain a dual attention fusion feature map, and pass the dual attention fusion feature map through a multi-layer perceptron to increase non-linear expression to obtain the final one-dimensional sequence;
[0011] Step 4: Reshape the final one-dimensional sequence into a two-dimensional image to obtain the two-dimensional feature map of the first stage;
[0012] Perform dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fuse the dimensionality reduction results to obtain the one-dimensional feature output of the first stage;
[0013] Step 5: Use the one-dimensional feature output of the first stage as the input feature of the next stage, and repeat Steps 3 to 4 iteratively several times to obtain two-dimensional feature maps output at different levels through different levels of the feature extraction network, and obtain several two-dimensional feature maps of different sizes;
[0014] Step 6: Fuse the features of several two-dimensional feature maps of different sizes and send them to the classification and regression stage for prediction to achieve traffic sign detection.
[0015] The present invention also proposes a traffic sign detection system based on class pooling connection attention fusion, wherein the system applies a traffic sign detection method based on class pooling connection attention fusion as described above, and the system includes:
[0016] A network construction module for:
[0017] Insert the scalable channel attention module into the Transformer based on the multi-head self-attention module and serialize it with the multi-head self-attention module to obtain a dual attention module. Package the PM module, the PC module, and at least one pair of dual attention modules in a serial manner into a class pooling connection attention module, and use several class pooling connection attention modules to construct a hierarchical feature extraction network;
[0018] A data preprocessing module for:
[0019] Preprocess the traffic image to obtain a two-dimensional image, encode the two-dimensional image into an initial one-dimensional sequence, and input it into a hierarchical feature extraction network;
[0020] A feature extraction module for:
[0021] Use a scalable channel attention module to extract features from the input features to obtain a channel attention feature map, use a multi-head self-attention module to extract features from the channel attention feature map to obtain a feature map with dual attention fusion, and pass the feature map with dual attention fusion through a multi-layer perceptron to increase non-linear expression to obtain the final one-dimensional sequence;
[0022] Reshape the final one-dimensional sequence into a two-dimensional image to obtain a two-dimensional feature map in the first stage, perform dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fuse the dimensionality reduction results to obtain a one-dimensional feature output in the first stage. Use the one-dimensional feature output in the first stage as the input feature for the next stage, and obtain two-dimensional feature maps output at different levels through different levels of the feature extraction network to obtain several two-dimensional feature maps of different sizes;
[0023] A feature fusion module for:
[0024] Fuse the feature maps of several two-dimensional feature maps of different sizes and send them to the classification and regression stage for prediction to achieve traffic sign detection.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] 1. By introducing scalable channel attention into Transformer, the scalable channel attention can perform feature scaling and reweighting on each channel, selectively enhance the features containing useful information, and suppress unimportant features, so as to solve the problem that Transformer lacks the ability to obtain channel dimension information, and at the same time strengthen the attention of Transformer to the target area and reduce the attention to background information to achieve traffic sign detection.
[0027] 2. By establishing a multi-scale feature extraction backbone model with dual attention fusion, the highest mAP is achieved under different meta-architectures after training for a small number of epochs. Among them, the mAP 50 accuracy of the CascadeRCNN as the meta-architecture reaches 84.0%, which is about 7% higher than that of the baseline model, indicating the effectiveness of the proposed multi-scale feature extraction backbone model with dual attention fusion.
[0028] 3. The present invention uses a PC module to replace the original convolution operation. The PC module and the PM module perform a dimensionality reduction operation through patch merging, and then fuse the dimensionality reduction results. The patch merging operation is similar to pooling. Using the PC module to replace the original convolution operation can make the image feature structure consistent, prevent network degradation while enhancing feature fusion, reduce the computational complexity and memory consumption, and avoid the problem of fusing different structural features.
[0029] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of a traffic sign detection method based on pooling-like connection attention fusion proposed by the present invention;
[0031] Figure 2 It is a model framework diagram of a traffic sign detection method based on pooling-like connection attention fusion of the present invention;
[0032] Figure 3 It is a comparison schematic diagram of the double attention module model of the present invention and the original Transformer structure;
[0033] Figure 4 It is a structural diagram of a traffic sign detection system based on pooling-like connection attention fusion of the present invention;
[0034] Figure 5 It is a schematic diagram of three major categories of traffic signs;
[0035] Figure 6 It is a schematic diagram of traffic image enhancement;
[0036] Figure 7 It is a schematic diagram of the PR curve of ResNet-101 under the Faster RCNN meta-architecture and the present invention as the backbone;
[0037] Figure 8 It is a schematic diagram of the PR curve of ResNet-101 under the RetinaNet meta-architecture and the present invention as the backbone;
[0038] Figure 9 It is a schematic diagram of the PR curve of ResNet-101 under the Cascade RCNN meta-architecture and the present invention as the backbone;
[0039] Figure 10 It is an ablation comparison schematic diagram under the RetinaNet meta-architecture;
[0040] Figure 11It is a schematic diagram of ablation comparison under the meta-architecture of Faster RCNN;
[0041] Figure 12 It is a schematic diagram of ablation comparison under the meta-architecture of Cascade RCNN;
[0042] Figure 13 It is a comparison of feature heat maps under three meta-architectures.
[0043] Among them, PR is the English abbreviation for recall rate and precision. Specific implementation manner
[0044] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0045] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited by this.
[0046] Please refer to Figures 1 to 3 , the embodiment of the present invention provides a traffic sign detection method based on class pooling connection attention fusion. The method includes the following steps:
[0047] Step 1: Insert a scalable channel attention module (SCAB) into a Transformer based on a multi-head self-attention module and serialize it with the multi-head self-attention module to obtain a dual attention module (DAB). Package the PM module, PC module, and at least one pair of dual attention modules in a serial manner into a class pooling connection attention module (PAB), and use several class pooling connection attention modules to construct a hierarchical feature extraction network (AFPC-T);
[0048] In the above solution, the multi-head self-attention module includes a window multi-head self-attention module and a sliding window multi-head self-attention module. The multi-head self-attention modules in each pair of dual attention modules respectively adopt the window multi-head self-attention module and the sliding window multi-head self-attention module.
[0049] The process of feature extraction by each pair of dual attention modules has the following relational formula:
[0050] ;
[0051] Among them, Represents the window multi-head self-attention module, Represents the sliding window multi-head self-attention module, Represents the scalable channel attention module, Represents the feature output after the fusion of the window multi-head self-attention module and the scalable channel attention module in the current dual attention module, Represents the feature output after the fusion of the sliding window multi-head self-attention module and the scalable channel attention module in the next dual attention module, Represents the multi-layer perceptron, Represents the output feature of the MLP module in the current dual attention module, Represents the output feature of the MLP module in the next dual attention module, Represents the normalization layer function, Represents the input feature of the dual attention module.
[0052] As Figure 3 shown, (a) is the Transformer structure, and (b) is the Transformer structure after adding the scalable channel attention module. Since the dual attention module needs to be arranged in pairs, two consecutive dual attention module structures are shown in this part.
[0053] Step 2: Preprocess the traffic image to obtain a two-dimensional image, encode the two-dimensional image into an initial one-dimensional sequence, and input it into the hierarchical feature extraction network;
[0054] Given a two-dimensional image of size divide it into one-dimensional sequences of size each, where represents the feature height, represents the feature width;
[0055] Perform a linear projection on the one-dimensional sequence to obtain an initial one-dimensional sequence of size where represents the feature dimension.
[0056] Step 3: Use the scalable channel attention module to extract features from the input feature to obtain a channel attention feature map, use the multi-head self-attention module to extract features from the channel attention feature map to obtain a dual attention fusion feature map, and pass the dual attention fusion feature map through a multi-layer perceptron to increase the non-linear expression to obtain the final one-dimensional sequence;
[0057] In the above solution, the method of using the scalable channel attention module to extract features from the input feature to obtain a channel attention feature map specifically includes the following steps:
[0058] Normalize the initial one-dimensional sequence first and then send it into the channel attention module. In the channel attention module, the relationship between channels is modeled to obtain the channel attention weights. The process of modeling the relationship between channels in the channel attention module has the following relational expressions:
[0059] ;
[0060] Among them, represents the Sigmoid activation function, represents the channel attention weight, represents the feature tensor, represents the balance factor that balances the influence of channel attention. The value of the balance factor is set to 0.1, represents a one-dimensional convolution. The convolution kernel size of the one-dimensional convolution is 3, represents the global average pooling operation. The process of performing the global average pooling operation on the feature tensor has the following relational expressions:
[0061] ;
[0062] Among them, represents a certain point on the feature tensor;
[0063] According to the feature channel weights, the feature maps under each channel of the feature tensor are rescaled to obtain the final output, that is, the channel attention feature map. The calculation process of the channel attention feature map has the following relational expressions:
[0064] ;
[0065] Among them, represents the result of multiplying the channel attention weight and the feature map channel by channel, represents the output of channel attention, that is, the feature map with channel attention, and , represents the dimension of feature map, represents the dimension of channel attention weight, and , represents the operation of multiplying the channel attention weight and the feature map channel by channel.
[0066] In the above scheme, the method of using the multi-head self-attention module to extract features from the channel attention feature map to obtain the double-attention fusion feature map specifically includes the following steps:
[0067] After the channel attention feature map is encoded by the sequence encoder, it becomes a one-dimensional sequence feature. Then, after relative position encoding, it is sent to the normalization layer for normalization. After that, a one-dimensional sequence feature with channel attention is obtained through the scalable channel attention module. The one-dimensional sequence feature is divided into local windows of size of , where represents the side length of the local window;
[0068] Using the local window features for linear mapping to obtain queries, keys, and values, the process of using the local window features for linear mapping has the following relationship:
[0069] ;
[0070] where represents a certain local window feature, and , , , represent the query matrix, key matrix, and value matrix respectively, represent the query vector, key vector, and value vector respectively, represents the sequence mapping dimension, represents the number of heads of self-attention, The relationship of is
[0071] Performing multi-head self-attention calculation on the queries, keys, and values to model the local window feature dependence relationship using the multi-head self-attention mechanism MHSA, obtaining the relative dependence relationship of the local space, that is, the feature map of double attention fusion. The process of performing multi-head self-attention calculation on the queries, keys, and values has the following relationship:
[0072] ;
[0073] where represents the transpose operation, represents the relative position encoding.
[0074] In the above solution, the process of increasing the non-linear expression of the feature map of double attention fusion through the multi-layer perceptron has the following relationship:
[0075] ;
[0076] where and represent two different learned linear transformations respectively, , , represents the dimension of the linear transformation, represents the activation function.
[0077] Step 4: Perform dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fuse the dimensionality reduction results to obtain the one-dimensional feature output of the first stage;
[0078] The specific steps of the above solution are as follows:
[0079] There is the following relational expression in the process of performing dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fusing the dimensionality reduction results to obtain the one-dimensional feature output of the first stage:
[0080] ;
[0081] Among them, represents the one-dimensional feature output of the th class pool connection attention module, represents the feature extraction operation of the dual attention module, represents the operation for performing patch merging on the final one-dimensional sequence, represents the operation for performing patch merging on the initial one-dimensional sequence.
[0082] In the PM module or the PC module, the one-dimensional sequence will be reshaped into a two-dimensional feature, then the width and height will be halved, and the dimension will become twice the original. After that, the two-dimensional feature is converted back into a one-dimensional sequence.
[0083] The present invention uses the PC module to replace the original convolution operation. The dimensionality reduction operation is realized by the PC module and the PM module in the way of patch merging, and then the dimensionality reduction results are fused. The operation of patch merging is similar to pooling. Using the PC module to replace the original convolution operation can make the image feature structure consistent, prevent network degradation while enhancing feature fusion, reduce the computational complexity and memory consumption, and avoid the problem of fusing different structural features at the same time.
[0084] Step 5: Use the one-dimensional feature output of the first stage as the input feature of the next stage, and repeat Steps 3 to 4 iteratively several times to obtain the two-dimensional feature maps output by different levels through different levels of the feature extraction network, and obtain several two-dimensional feature maps of different sizes;
[0085] As Figure 2 shown, the class pool connection attention module has four layers. The size of the feature map (F1) generated in the first stage is , the size of the feature map (F2) generated in the second stage is , the size of the feature map (F3) generated in the third stage is , and the size of the feature map (F4) generated in the fourth stage is .
[0086] Step 6: Fuse the feature maps of several two-dimensional features of different sizes and send them to the classification and regression stages for prediction to achieve traffic sign detection.
[0087] Please refer to Figure 4 , the embodiment of the present invention also provides a traffic sign detection system based on class pooling connection attention fusion. Among them, the system applies a traffic sign detection method based on class pooling connection attention fusion as described above. The system includes:
[0088] Network construction module, used for:
[0089] Insert the scalable channel attention module into the Transformer based on the multi-head self-attention module and connect it in series with the multi-head self-attention module to obtain a dual attention module. Package the PM module, PC module, and at least one pair of dual attention modules in a serial manner as a class pooling connection attention module, and use several class pooling connection attention modules to construct a hierarchical feature extraction network;
[0090] Data preprocessing module, used for:
[0091] Preprocess the traffic image to obtain a two-dimensional image, encode the two-dimensional image into an initial one-dimensional sequence, and input it into the hierarchical feature extraction network;
[0092] Feature extraction module, used for:
[0093] Use the scalable channel attention module to extract features from the input features to obtain a channel attention feature map, use the multi-head self-attention module to extract features from the channel attention feature map to obtain a dual attention fusion feature map, and increase the non-linear expression of the dual attention fusion feature map through a multi-layer perceptron to obtain the final one-dimensional sequence;
[0094] Reshape the final one-dimensional sequence into a two-dimensional image to obtain the two-dimensional feature map of the first stage. Perform dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fuse the dimensionality reduction results to obtain the one-dimensional feature output of the first stage. Use the one-dimensional feature output of the first stage as the input feature of the next stage, and obtain the two-dimensional feature maps output at different levels through different levels of the feature extraction network to obtain several two-dimensional feature maps of different sizes;
[0095] Feature fusion module, used for:
[0096] Fuse the feature maps of several two-dimensional features of different sizes and send them to the classification and regression stages for prediction to achieve traffic sign detection.
[0097] This embodiment uses Figure 5 to identify the object recognized by the present invention, and uses Figure 6To identify the feature enhancements made, the steps include S201 to S202.
[0098] Please refer to Figure 5 shown, (a) Prohibitory sign; (b) Mandatory sign; (c) Warning sign. The number of photos exceeds 100,000, and the resolution is 2048×2048 pixels.
[0099] To improve the detection effect, this paper deleted unlabeled and duplicate traffic sign images from the dataset and selected 42 traffic sign categories, with more than 100 images in each category. Among them, there are 6105 training images and 3071 test images.
[0100] Please refer to Figure 6 , in addition, to improve the prediction performance of the model, data augmentation techniques were also used to expand the dataset. As Figure 6 shown, in the figure, (a) is the original image, (b) is the image with adjusted brightness, (c) is the image with added noise, (d) is the flipped image. Therefore, through at least one or more effects such as brightness change, adding noise, and flipping, each category has more than 500 instances. After data augmentation, the final training dataset contains 17704 images. Table 1 shows the final number of training and test images.
[0101] Table 1 Final number of images
[0102]
[0103] S201. Design a method for enhanced images.
[0104] Specific enhancement methods include brightness change, adding noise, and flipping.
[0105] S2011. Select an enhancement method.
[0106] S2011a. Data augmentation is generally divided into geometric augmentation and color augmentation. There are also many types of color augmentation, and the most common one is brightness transformation. Usually, the brightness range covered by the collected data may not be sufficient, and the robustness to brightness has almost become a basic requirement for deep learning. Using brightness enhancement is becoming more and more important
[0107] S2011b. The basic idea of the noise data augmentation principle is to add a certain amount of noise to the original data, thereby generating new data samples, which can effectively increase the size of the dataset, thereby improving the generalization ability and robustness of the model.
[0108] S2011c. Image flipping is a way of data augmentation by rotating and flipping the image to achieve image amplification.
[0109] S202. Specific enhancement method selection;
[0110] S202a. Brightness adjustment is enhanced through the following formula:
[0111] ;
[0112] Among them, represents the pixel value before adjustment, represents the pixel value after adjustment. is used to adjust the brightness, takes values in [-1, 1], is used to adjust the contrast, takes values in [1, 89]. When = 0, only the contrast is adjusted; when c = 0, only the brightness is adjusted.
[0113] S202b. Adding noise is enhanced through the following formula:
[0114] Gaussian noise: The noise is distributed on each pixel point, and the amplitude value is random, and the distribution approximately conforms to the normal distribution.
[0115] ;
[0116] Among them, represents the Gaussian white noise value at time t, represents the amplitude of the noise, represents the time offset of the noise, represents the standard deviation of the noise.
[0117] S202c. Image flipping is enhanced through the following formula:
[0118] ;
[0119] Among them, represents the pixel point coordinates in the original image, represents the corresponding pixel point coordinates in the image after rotation transformation.
[0120] Please refer to Figure 7 , Figure 8 and Figure 9 , using the above dataset as the input image, respectively using ResNet-50, ResNet-101, PVT-T, and Swin-T to replace the hierarchical feature extraction network of the present invention for comparative experiments to verify the performance of the present invention, and the results are shown in Table 2.
[0121] It should be noted that in order to fully verify the performance of the model, three representative meta-architectures and ResNet-101 were used as the baseline model to evaluate the performance of the present invention. The meta-architectures mainly include two two-stage models, FasterRCNN and Cascade RCNN, and a single-stage model, RetinaNet. Specifically, the backbone of these frameworks was constructed using the present invention. Comparative experiments were conducted with ResNet-50, ResNet-101, PVT-T, and Swin-T.
[0122] Please refer to Figure 7 , Figure 8 and Figure 9 , in order to study the relationship between the precision and recall of the hierarchical feature extraction network of the present invention and that with ResNet-101 as the backbone under the same classifier, the horizontal axis in the figure is the recall and the vertical axis is the precision. The higher the precision, the more it indicates that the positive samples predicted by the model are true positive samples; the higher the recall, the more it indicates that the model can correctly predict the true samples. The PR curves under different IoU thresholds (i.e., from 0.5 to 0.85) are shown, and the PR curves with IoU thresholds of 0.9 and 0.95 are excluded. The PR curves with Faster RCNN, RetinaNet, and Cascade RCNN as the meta-architectures are respectively as Figure 7 , Figure 8 and Figure 9 shown.
[0123] It can be seen from the results that the meta-architecture with the hierarchical feature extraction network of the present invention as the backbone is superior to that with ResNet-101 as the backbone in terms of both recall and precision. In addition, the area of the PR curve with the hierarchical feature extraction network of the present invention as the backbone is also significantly larger than that with ResNet-101 as the backbone, which also reflects that the average precision with the hierarchical feature extraction network of the present invention as the backbone is superior to that with ResNet-101 as the backbone.
[0124] As can be seen from Table 2 below, among all the meta-architecture methods, the model with the hierarchical feature extraction network of the present invention as the backbone outperforms the baseline model, and the FPS does not decrease significantly. Compared with the baseline model, the maximum increase in mAP50 of the model based on RetinaNet is about 7%. In addition, its AP small increases by about 3%, AP medium increases by about 6%, and AP large increases by about 8%. Although there is a significant improvement in RetinaNet, Cascade RCNN achieves the best result, with its mAP50 reaching 84.0% and mAP75 reaching 78.7%. The experimental results show that the mAP at different object sizes is greatly improved with a slight decrease in FPS, which reflects the balance between detection accuracy and inference speed to a certain extent.
[0125] Table 2 Performance comparison. Among them, mAP50 and mAP75 represent IoU of 0.5 and 0.75 respectively. Small, medium, and large represent mAP corresponding to small, medium, and large object groups respectively.
[0126]
[0127] Please refer to Figures 10 to 12 , and an ablation experiment was conducted on the model constructed by the present invention and Swin-T to verify the effectiveness of the present invention.
[0128] As shown in the following table, ablation was carried out to further verify the effectiveness of the present invention. By adding the scalable channel attention module (SCAB) and PC module to the baseline model one by one, their effects were demonstrated. By adding the scalable channel attention module to activate more important dimensions, the performance of Faster RCNN, Cascade RCNN, and RetinaNet has been significantly improved, especially in the large, medium, and small ranges. Cascade RCNN+SCAB increases its mAP50, mAP75, small, and medium from 83.8%, 78.5%, 45.1%, and 74.3% to 85.0%, 79.8%, 47.2%, and 75.1% respectively. The effectiveness of SCAB was demonstrated with a slight decrease in FPS.
[0129] In order to explore the role of each module, the effect of PC was also evaluated. PC improves the performance of the detector to a certain extent. After adopting Faster RCNN+SCAB+PC, its mAP 50 、mAP 75Small, medium, and large increased from 78.6%, 73.4%, 37.2%, 71.9%, and 74.9% to 80.4%, 75.3%, 40.7%, 73.2%, and 75.6% respectively. The experimental results show that both SCAB and PC improve the performance of the present invention, and their combination achieves the best performance.
[0130] Table 3 Ablation comparison. +SCAB means adding a scalable channel attention module to the hierarchical feature extraction network constructed by Swin-T. +SCAB, +PC means adding a PC module while adding a scalable channel attention module to the hierarchical feature extraction network constructed by Swin-T.
[0131]
[0132] Please refer to Figure 10 , RetinaNet, Figure 11 , Faster RCNN, Figure 12 , Cascade RCNN as shown, under different meta-architectures, adding different networks and their effects on loss and number of training times. It can be found that the training time to reach the same loss is significantly reduced after adding SCAB and PC, indicating that all networks are effective.
[0133] Please refer to Figure 13 as shown, for the comparison of heatmaps generated by the present invention and Swin-T as the backbone model under three meta-architectures of (a) FasterRCNN, (b) RetinaNet, and (c) Cascade RCNN. Among each left and right pair, the left group is Swin-T and the right group is the present invention. It can be seen that the model based on the present invention pays less attention to background information, indicating better effects under attention fusion.
[0134] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0135] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A traffic sign detection method based on class pooling connection attention fusion, characterized in that, the method comprises the following steps: Step 1: Insert a scalable channel attention module into the Transformer based on the multi-head self-attention module and serialize it with the multi-head self-attention module to obtain a dual attention module. Package the PM module, the PC module, and at least a pair of dual attention modules in a serial manner as a class pooling connection attention module, and use a number of class pooling connection attention modules to construct a hierarchical feature extraction network; Step 2: Preprocess the traffic image to obtain a two-dimensional image, encode the two-dimensional image into an initial one-dimensional sequence, and input it into the hierarchical feature extraction network; Step 3: Use the scalable channel attention module to extract features from the input features to obtain a channel attention feature map, use the multi-head self-attention module to extract features from the channel attention feature map to obtain a dual attention fusion feature map, and pass the dual attention fusion feature map through a multi-layer perceptron to increase non-linear expression to obtain the final one-dimensional sequence; Step 4: Reshape the final one-dimensional sequence into a two-dimensional image to obtain the two-dimensional feature map of the first stage; Perform dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fuse the dimensionality reduction results to obtain the one-dimensional feature output of the first stage; Step 5: Use the one-dimensional feature output of the first stage as the input feature of the next stage, and repeat Steps 3 to 4 iteratively for a number of times to obtain two-dimensional feature maps output at different levels through different levels of the feature extraction network, and obtain a number of two-dimensional feature maps of different sizes; Step 6: Fuse the two-dimensional feature maps of different sizes and send them to the classification and regression stage for prediction to achieve traffic sign detection; The PC module and the PM module adopt the method of patch merging; The following relationship exists in the process of performing dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively and fusing the dimensionality reduction results to obtain the one-dimensional feature output of the first stage: ; Among them, represents the one-dimensional feature output of the th class pool connection attention module, represents the feature extraction operation of the dual attention module, represents the operation for patch merging on the final one-dimensional sequence, represents the operation for patch merging on the initial one-dimensional sequence.
2. A traffic sign detection method based on class pooling connection attention fusion according to claim 1, characterized in that, in Step 1, the multi-head self-attention module includes a window multi-head self-attention module and a sliding window multi-head self-attention module, and the multi-head self-attention modules in each pair of dual attention modules adopt the window multi-head self-attention module and the sliding window multi-head self-attention module respectively.
3. A traffic sign detection method based on class pooling connection attention fusion according to claim 2, characterized in that, in Step 1, the following relationship exists in the process of feature extraction of each pair of dual attention modules: ; Among them, represents the window multi-head self-attention module, represents the sliding window multi-head self-attention module, represents the scalable channel attention module, represents the feature output after the fusion of the window multi-head self-attention module and the scalable channel attention module in the current dual attention module, represents the feature output after the fusion of the sliding window multi-head self-attention module and the scalable channel attention module in the next dual attention module, represents the multi-layer perceptron, represents the output feature of the MLP module in the current dual attention module, represents the output feature of the MLP module in the next dual attention module, represents the normalization layer function, represents the input feature of the dual attention module.
4. A traffic sign detection method based on class pooling connection attention fusion according to claim 3, characterized in that, in Step 2, the method of encoding the two-dimensional image into an initial one-dimensional sequence specifically includes the following steps: Given a two-dimensional image of size , divide it into one-dimensional sequences of size , where represents the feature height and represents the feature width. Perform a linear projection on the one-dimensional sequence to obtain an initial one-dimensional sequence of size , where represents the feature dimension.
5. A traffic sign detection method based on class pooling connection attention fusion according to claim 4, characterized in that, In step 3, the method for extracting features from the input features using the scalable channel attention module to obtain the channel attention feature map specifically includes the following steps: Normalize the initial one-dimensional sequence first and then send it into the channel attention module, that is, the scalable channel attention module. Use the channel attention module to model the relationship between channels to obtain the channel attention weights. The process of using the channel attention module to model the relationship between channels has the following relational expression: ; Among them, represents the Sigmoid activation function, represents the channel attention weight, represents the feature tensor, represents the balance factor for balancing the influence of channel attention, represents one-dimensional convolution, represents the global average pooling operation. There is the following relational expression in the process of performing the global average pooling operation on the feature tensor: ; Among them, represents a certain point on the feature tensor; According to the feature channel weights, rescale the feature maps under each channel of the feature tensor to obtain the final output, that is, the channel attention feature map. The calculation process of the channel attention feature map has the following relational expression: ; Among them, represents the result of multiplying the channel attention weights by the feature map channel by channel, represents the output of channel attention, that is, the feature map with channel attention, and , represents a feature map with a dimension of , represents the channel attention weights with a dimension of , and , represents the operation of multiplying the channel attention weights by the feature map channel by channel.
6. A traffic sign detection method based on class pooling connection attention fusion according to claim 5, characterized in that In step 3, the method for extracting features from the channel attention feature map using the multi-head self-attention module to obtain the double attention fusion feature map specifically includes the following steps: The channel attention feature map becomes a one-dimensional sequence feature after sequence encoding, and then is fed into a normalization layer after relative position encoding for normalization. After that, a one-dimensional sequence feature with channel attention is obtained through a scalable channel attention module. The one-dimensional sequence feature is divided into local windows of size where represents the side length of the local window; Perform a linear mapping using the local window features to obtain the query, key, and value. The process of performing a linear mapping using the local window features has the following relational expression: ; Among them, represents a certain local window feature, and , , , represent the query matrix, the key matrix, and the value matrix respectively, represent the query vector, the key vector, and the value vector respectively, represents the sequence mapping dimension, represents the number of heads of self-attention, has the relationship of ; Perform multi-head self-attention calculations on the query, key, and value to model the local window feature dependence relationship using the multi-head self-attention mechanism MHSA to obtain the relative dependence relationship in the local space, that is, the double attention fusion feature map. The process of performing multi-head self-attention calculations on the query, key, and value has the following relational expression: ; Among them, represents the transpose operation, represents the relative position encoding.
7. A traffic sign detection method based on class pooling connection attention fusion according to claim 6, characterized in that In step 3, the process of increasing the non-linear expression of the double attention fusion feature map through a multi-layer perceptron has the following relational expression: ; Among them, and respectively represent two different learned linear transformations, , , represent the dimensions of the linear transformation, represents the activation function.
8. A traffic sign detection method based on class pooling connection attention fusion according to claim 7, characterized in that The hierarchical feature extraction network has four layers. The size of the feature map generated in the first stage is , the size of the feature map generated in the second stage is , the size of the feature map generated in the third stage is , and the size of the feature map generated in the fourth stage is .
9. A traffic sign detection system based on class pooling connection attention fusion, characterized in that The system applies a traffic sign detection method based on class pooling connection attention fusion as described in any one of claims 1 to 8. The system includes: A network construction module for: Insert the scalable channel attention module into the Transformer based on the multi-head self-attention module and serialize it with the multi-head self-attention module to obtain a double attention module. Package the PM module, PC module, and at least one pair of double attention modules in a serial manner as a class pooling connection attention module, and use several class pooling connection attention modules to construct a hierarchical feature extraction network; A data preprocessing module for: Preprocess the traffic image to obtain a two-dimensional image, encode the two-dimensional image into an initial one-dimensional sequence, and input it into the hierarchical feature extraction network; A feature extraction module for: Extract features from the input features using the scalable channel attention module to obtain the channel attention feature map, extract features from the channel attention feature map using the multi-head self-attention module to obtain the double attention fusion feature map, and increase the non-linear expression of the double attention fusion feature map through a multi-layer perceptron to obtain the final one-dimensional sequence; Reshape the final one-dimensional sequence into a two-dimensional image to obtain the two-dimensional feature map of the first stage. Perform dimensionality reduction operations on the final one-dimensional sequence and the initial one-dimensional sequence respectively, and fuse the dimensionality reduction results to obtain the one-dimensional feature output of the first stage. Use the one-dimensional feature output of the first stage as the input feature of the next stage, and obtain the two-dimensional feature maps output by different levels through different levels of the feature extraction network to obtain several two-dimensional feature maps of different sizes; A feature fusion module, which is used for: Fuse the two-dimensional feature maps of several different sizes and send them to the classification and regression stage for prediction to achieve traffic sign detection.