A Yolov5 Object Detection Method Based on Cross-Stage Routing Attention Module and Residual Information Fusion Module
By improving the cross-stage routing attention module and residual information fusion module of the Yolov5 network, the small object detection capability is enhanced, and the accuracy and speed problems of traffic sign detection in harsh environments are solved, and high-precision and robust traffic sign detection are achieved.
Patent Information
- Application Number
- CN202310865846.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-07-14
AI Technical Summary
When facing factors such as bad weather, light changes, and occlusion, the existing traffic sign detection algorithm has problems such as low detection accuracy and slow speed, especially the detection capabilities of small targets, and the traditional methods are poorly robust and cannot effectively meet the real-time needs of autonomous driving.
The improved Yolov5 object detection method based on the cross-stage routing attention module and residual information fusion module is adopted. By optimizing the backbone, neck and decoupling network of the Yolov5 network, the detection ability of small targets is enhanced, multi-scale feature information is fused, and the two-branch prediction decoupling head and Focal EIOU loss function are used to improve detection accuracy and robustness.
It has achieved high-precision and high-rootability detection of small targets such as traffic signs, which can effectively deal with light changes and the impact of disturbances, improve detection speed and accuracy, and is suitable for the real-time needs of autonomous driving.
Smart Images

Figure CN116721398B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and more specifically relates to a Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module. Background Art
[0002] With the rapid rise of intelligent transportation, vehicle autonomous driving technology has gradually become a research hotspot. Research institutions such as ADT that study autonomous driving technology combine multiple technologies such as computer technology, artificial intelligence, and sensors. By perceiving information about itself and the surrounding environment, relying on the perceived information, making decision judgments, and implementing commands to control the vehicle, so as to achieve the purpose of autonomous driving. Traffic sign detection and recognition is an important task for realizing autonomous driving. However, affected by factors such as bad weather, light changes, damaged shapes and faded colors of traffic signs, and occlusion, at the same time, traffic signs are small object detections, occupying limited pixels in the image and carrying limited information, and the model can only capture very little appearance information, which makes traffic sign detection a challenging task in the field of computer vision. In recent years, the problem of traffic sign detection has received extensive attention, and a large number of scholars have conducted in-depth research on traffic sign detection.
[0003] In the current development technologies of deep learning such as CNN convolutional neural networks, object detection models based on deep learning are generally divided into two-stage and single-stage. Two-stage detectors divide the detection problem into two stages. First, candidate regions are generated by the method of region proposal for regions of interest to determine the background and foreground, and then the proposed candidate regions are accurately classified and precisely located. Usually, the accuracy is high, but the detection speed is slow and the process is relatively cumbersome. Representative algorithms include: Faster-RCNN, Cascade-RCNN, VFNET, CenterNet, etc. Single-stage detectors do not need to generate candidate regions and directly generate classification results and position coordinate information through objects. Therefore, their detection speed is fast and the method is simple, but the detection accuracy is somewhat lost. Representative algorithms include: YOLO, FCOS, DETR, EffiencentDet, etc. The detection speed of single-stage detectors is superior to that of two-stage detectors on the premise of little loss of accuracy, and can meet the real-time performance of traffic sign detection. Therefore, the detection method studied in the present invention is mainly based on the single-stage object YOLO (You Only Look Once).
[0004] In deep learning, the Attention mechanism is a method similar to the human visual neural network. Humans can effectively find significant regions in complex scenes by focusing attention on key points. However, Hard Attention selects information based on maximum sampling or random sampling, and cannot be trained using the backpropagation algorithm. Therefore, Soft Attention is widely used in computer vision. The SE Attention converts channels into vectors through global average pooling, and then the feature map is weighted according to the output of the side network. The CBAM Attention mechanism uses spatial feature relationships to supplement the attention channel relationships. The spatial attention part is calculated through average pooling and maximum pooling of channels. Convolutional attention has low computational overhead, but cannot capture position information and cannot learn the sequential relationships in the sequence. The self-attention mechanism provides position information for the image and calculates the similarity between points in the feature map, reducing the dependence on external information. It can not only learn the sequential order information of the image but also focus more on important regions in the size context, effectively capturing the correlation between data and features.
[0005] Currently, traditional traffic sign detection algorithms, such as those based on geometric information like color and shape, and histograms of oriented gradient, have defects such as high overhead costs, poor robustness, and lack of real-time performance. These methods cannot effectively meet the requirements of autonomous driving technology for accuracy and detection speed. For traffic sign detection algorithms based on convolutional neural networks, as the number of network layers increases, the receptive field of the extracted feature map becomes larger, and the deep semantic information of the image becomes stronger. However, the texture and spatial position information contained in the shallow feature map are lost, resulting in the loss of information about small objects in the feature map. In traditional single-stage object detectors, usually only the last layer of the feature map is used for regression and localization. This leads to less effective information about small targets on the last feature map, reducing the detection ability for small targets such as traffic signs. There is a semantic gap between different layers, and due to the limitation of unidirectional information flow transmission, high-level semantic information and low-level spatial information are not fully utilized.
[0006] In view of the above problems existing in the existing traffic sign small target detection algorithms, it is urgent to design a new improved Yolov5 object detection method. Summary of the Invention
[0007] (1) Technical Problem
[0008] Based on the problems existing in the existing traffic sign object detection algorithms, the present invention proposes a Yolov5 object detection method based on a cross-stage routing module and a residual information fusion module. By optimizing and improving the structures of the backbone, neck, and decoupled network of the original Yolov5, the detector can fully fuse multi-scale feature information and enhance the detection of small objects such as traffic signs. It pays attention to the use of shallow feature maps, which is more conducive to the detection of small objects. At the same time, only two detection decoupled heads are used to achieve higher accuracy. This method can effectively cope with various challenges such as illumination changes, deformations, scale changes, and the influence of interference objects, providing high-precision and high-robustness object detection, which helps the model to make accurate inferences.
[0009] (2) Technical solutions
[0010] The present invention provides a Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module. This method improves the original Yolov5 network, specifically including:
[0011] (1) Cross-stage routing attention module
[0012] In the Yolov5 network, the C3 modules of the 6th and 8th layers in the backbone network are replaced with cross-stage routing attention modules, and the feature map information of the 2nd, 4th, 6th, and 9th layers is sequentially used as input signals P2, P3, P4, and P5 and input into the neck network; the composition method of the cross-stage routing attention module is as follows: first, the number of channels of the feature map is divided into two parts. The first part enhances the feature information through an attention mechanism, and the other part is output-merged with the enhanced feature through cross-stage, and finally, a residual structure is used for local semantic enhancement;
[0013] (2) Multi-scale feature fusion method
[0014] In the neck network, through the backbone network, i represents the i-th layer feature map extracted from the backbone network, i ∈ {0, 1, 3, 4}, and f0 to f3 respectively correspond to the input signals P5 to P2, where C i ∈ {1024, 512, 256, 128}; S i represents the output result after passing through the multi-level information fusion module, and its mathematical formula through the multi-scale feature fusion network is expressed as:
[0015] S0 = f0
[0016] S i = MRI(f i , S i-1 ) i = 1, 4
[0017] Si = MRI(f i , S i-1 , f i+1 ) for i = 2, 3
[0018] S i = MRI(S i-3 , S i-4 ) for i = 5, 7
[0019] S i = MRI(S i-3 , S i-4 , S i-1 ) for i = 6
[0020] Among them, the meaning of the MRI function is a multi-scale feature fusion module function, which can fuse the multi-scale features of each parameter based on splicing and upsampling;
[0021] (3) Dual-branch prediction decoupling head
[0022] In the decoupling head, the two-layer outputs of the feature maps corresponding to the relatively shallower P2 and P3 layers in the neck network are used to perform dual-branch prediction decoupling output respectively as the final prediction result.
[0023] Preferably, the detection target of the Yolov5 object detection method based on the cross-stage routing attention module and the residual information fusion module is a traffic sign.
[0024] On the other hand, the present invention also discloses a Yolov5 object detection system based on a cross-stage routing attention module and a residual information fusion module, characterized in that it includes:
[0025] At least one processor; and
[0026] At least one memory communicatively connected to the processor, wherein:
[0027] The memory stores program instructions executable by the processor, and the processor can execute the Yolov5 object detection method based on the cross-stage routing attention module and the residual information fusion module as described in any one of the above.
[0028] On the other hand, the present invention also discloses a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the Yolov5 object detection method based on the cross-stage routing attention module and the residual information fusion module as described in any one of the above.
[0029] (III) Beneficial effects
[0030] Compared with the prior art, the Yolov5 object detection method based on the cross-stage routing module and the residual information fusion module has the following advantages:
[0031] (1) The technical solution of the present invention is an improvement based on the Yolov5 network model. The present invention proposes a Yolov5 object detection method based on the cross-stage routing attention module (CSB module) and multi-scale residual information fusion based on the MRI module. By strengthening the feature extraction of the feature map, it can effectively capture the dependency relationships in the data, establish the interaction between global information and local information, and quickly focus on the most relevant regions; the residual information fusion module can enhance the semantic information of the fused multi-scale feature map. By using the method of fusing deeper receptive fields with shallow receptive fields, it can effectively improve the accuracy of traffic sign object detection, and make full use of the different feature information between multi-scale feature maps. The three cooperate with each other to enable the detector of the present invention to obtain better small object detection effects.
[0032] (2) In addition, during the object detection process, the backbone network is used to extract features from the input image. By gradually reducing the size of the feature map, it enhances the semantic information of the features, captures information such as the structure, texture, and edges in the image, and then the extracted features are sent into the neck feature fusion network. The feature fusion network is used to further process the features extracted by the backbone network and perform cross-scale feature fusion operations on the feature maps from different backbone network levels. The present invention fuses feature maps of different scales through methods such as upsampling and splicing, and then inputs the obtained feature map into the residual information fusion module proposed by the present invention to enhance the semantic information of the fused multi-scale feature map, and fuses the features with rich high-resolution semantic information and accurate low-resolution spatial information to obtain comprehensive features. Finally, the decoupled dual-branch prediction head performs classification and bounding box regression on the basis of enhanced features, and finally can quickly generate more accurate detection results for small objects such as traffic signs. And through experimental comparison, it can be seen that the present invention has obtained the best performance in various indicators, which proves the superior performance of the method proposed by the present invention for traffic sign object detection. Description of the Drawings
[0033] Figure 1 is the network structure schematic diagram of YOLOV5 in the prior art;
[0034] Figure 2 is the overall structure diagram of the Yolov5 object detector based on the cross-stage routing attention module and the residual information fusion module in the present invention;
[0035] Figure 3 is the structure diagram of the BRA attention mechanism in the cross-stage routing attention module CSB in the present invention;
[0036] Figure 4 It is the structural diagram of the cross-stage routing module (CSB module) in the present invention;
[0037] Figure 5 It is the structural diagram of the MRI module in the present invention;
[0038] Figure 6 It is the simplified structural diagram of the neck network based on multi-scale bidirectional feature fusion in the present invention;
[0039] Figure 7 It is the schematic structural diagram of the Couple Head and Decouple Head in the present invention. Specific Embodiments
[0040] The following combines the accompanying drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0041] In order to improve the detection accuracy and speed of traffic signs to obtain better detection effects, the network structure of the present invention is improved on the basis of the YOLOV5 network in the prior art. Compared with Figure 1 the original YOLOV5 network in the prior art, the main solutions adopted in the present invention are mainly divided into three parts of improvement content. In the last two layers of the backbone network, a CSB structure based on small targets (i.e., cross-stage routing attention module) is selected to replace the original C3 module, and a BRA attention mechanism is introduced in the CSB structure to strengthen the features of small object targets in the deep features of the backbone network; secondly, for traffic signs, the present invention improves the original PANET in the neck, designs a new fusion network to replace the neck network in Yolov5, and uses the MRI module to fuse multi-scale features. Finally, in the last prediction part of the network, only two shallow decoupled detection heads are used to replace the original decoupled head in YOLO, and at the same time, the Focal EIOU loss function is introduced to replace the original CIOU loss function in YOLOV5.
[0042] See Figure 2 As shown, the new Yolov5 object detection method based on the cross-stage routing module and residual information fusion module of the present invention specifically includes the following improvements:
[0043] I. Cross-stage routing attention module (abbreviation: CSB module)
[0044] The present invention designs a cross-stage routing attention module, as Figure 2As shown in the figure, in the original Yolov5 network, the C3 modules in the 6th and 8th layers of the backbone network are replaced with CSB modules, and the information of the 2nd, 4th, 6th, and 9th (SPPF layer) layers is input into the neck network as input signals P2, P3, P4, and P5 in sequence. The CSB module is used to enhance the feature extraction of small targets of traffic signs. This method can effectively capture the dependence relationship in the data. The routing attention method used in the present invention can establish the interaction between global information and local information, quickly focus on the most relevant areas, and has high parallelism, which is beneficial to the subsequent training and inference of the model.
[0045] Specifically, as Figure 4 shown, the composition method of the CSB module in the present invention is as follows: First, the number of channels of the feature map is divided into two parts. The first part enhances the feature information through the attention mechanism, and the other part is output-merged with the enhanced feature through cross-stage, and finally the residual structure is used for local enhancement of semantics.
[0046] Specifically, the attention mechanism in the other part of the channels of the CSB module of the present invention can be either the bi-level routing attention module in the prior art or preferably the BRA attention mechanism specially improved and designed in the present invention as Figure 3 shown, so that after cooperating with the cross-stage structure of the cross-stage routing attention module CSB, it has the functions of quickly focusing on the most relevant areas and having high parallelism.
[0047] As Figure 3 shown, the composition method of the BRA attention mechanism improved and designed by the present invention for the bi-level routing attention module is as follows:
[0048] In the BRA attention mechanism, any unit X ∈ R C×H×W is taken as the input, and Y ∈ r C×H×W is taken as the output. First, the present invention divides the feature map into non-overlapping regions of P×P blocks, and then flattens these regions in the spatial dimension to obtain feature vectors, and then the obtained feature vectors are input and deduced through linear mapping to obtain where C is the number of channels of the feature map, H and W are the width and height of the feature map respectively, and P is the number of blocks divided by the feature map;
[0049] Q = X r W Q ,K = X r W K ,V = X r W V
[0050] Here, W Q , KQ , V Q ∈R C×C are all parameter matrices, representing the linear mapping weight matrices of Query, Key, and Value in the mapped image. Then, the average value of the regions divided by the eigenvectors of Q and K is calculated to obtain Q r ,
[0051] A r = Mean(Q) × Mean(K)
[0052] Then Q r and K r are multiplied to calculate the most relevant similarity to obtain S. The top K operator is used on the obtained similarity matrix to retain the indices of the top K regions with the closest relationships, resulting in the region routing index I k ; Finally, the adjacent miniatures are connected to the original K and V;
[0053] K g = gather(I K , K), V g = gather(I K , V)
[0054] where
[0055] Then the obtained K g and V g are used for self-attention calculation with the original Q. At the same time, the obtained V r is used for residual calculation with the result of self-attention through depthwise separable convolution;
[0056]
[0057] BRA(X) = Attention(Q, K g , V g ) + DwConv(V g )
[0058] The final result O ∈ R C×H×W is obtained, where Mean is the mean function, softmax is the normalization exponential function, gather is the dimension concatenation function, and DwConv is the depthwise separable convolution function.
[0059] It can be seen from this that the improved attention module BRA of the present invention has made certain modifications to the attention structure of the original double-layer routing attention module, and deleted the depthwise separable convolution and the residual of self-attention in the last stage. Such a method filters out irrelevant regions at the coarse region level, retains the routing of some relevant regions through the adjacency matrix, finds two semantically related regions through routing, continuously queries relevant regions through this method and combines vectors, without distracting the attention of other irrelevant tokens. Therefore, it has good performance and high computational efficiency. And traffic signs are small target objects. In the feature map of the deep network in the backbone network, the number of pixels of small target objects it contains is small. The ultimate purpose of designing the cross-stage routing attention module CSB of the present invention is to enable the model to pay attention to the small target adjacent objects in the deep feature map, so that it can quickly pay attention to more effective regions.
[0060] II. Multi-scale Feature Fusion Method
[0061] As Figure 2 shown, for the four-level input signals P2 to P5 of the backbone network, the neck network performs multi-scale feature fusion. Using multi-level feature maps of different sizes can increase the semantic information and location information of features. Especially when detecting small target objects, the improvement effect is significant.
[0062] In this regard, the present invention designs a novel multi-scale bidirectional feature fusion network to achieve feature map fusion. Different scales of feature maps are used as inputs. Shallow features focus on details and location information, which is helpful for positioning. Deep features contain rich semantic information and are more helpful for classification. The present invention changes feature maps of different scales into the same size through convolution operations, so that feature maps of different scales and different numbers of channels can also share the same channel dimension, which is more conducive to multi-scale feature fusion. However, it also brings a small amount of additional computational overhead. Nevertheless, the present invention still ensures the real-time performance of traffic sign targets, and the detection accuracy based on the improved network structure of lightweight YOLOV5 has been greatly improved, proving the effectiveness of the method of the present invention.
[0063] See Figure 6 It can be known that through the backbone network, i represents the i-th layer feature map extracted from the backbone network, i ∈ {0, 1, 3, 4}, and f0 to f3 respectively correspond to the input signals P5 to P2, where C i ∈ {1024, 512, 256, 128}; S i represents the output result after passing through the multi-level information fusion module, and its mathematical formula after passing through the multi-scale feature fusion network is expressed as:
[0064] S0 = f0
[0065] S i = MRI(f i ,S i-1 ), i = 1, 4
[0066] S i = MRI(f i ,S i-1 ,f i+1 ), i = 2, 3
[0067] S i = MRI(S i-3 ,S i-4 ), i = 5, 7
[0068] S i = MRI(S i-3 ,S i-4 ,S i-1 ), i = 6
[0069] Among them, the meaning of the MRI function is a multi-scale feature fusion module function, which can fuse multi-scale features of each parameter based on splicing and upsampling.
[0070] Since feature maps of different scales have different spatial information and semantic information, problems such as information mismatch and loss, context information mismatch, information overlap and loss often occur during scale fusion, making it difficult for the model to understand the semantic information in the image and resulting in a decline in model performance. The original C3 module in YOLOV5 improves the receptive field and depth of the network, but it cannot effectively combine the fusion of feature maps of different scales, and it has no clear mechanism to promote the reuse of cross-layer features. After the fusion features are realized through the splicing operation for feature splicing, there may be insufficient fusion of information in the shallow features in the deep features, resulting in the loss of small target information. In view of the deficiencies of the C3 module, the present invention designs a multi-level information fusion module of P2 to P5, and replaces the original network fusion module with a multi-level information fusion module after the splicing operation. The purpose is to enhance the semantic information between feature maps of different scales, and at the same time further increase the receptive field of the feature map, so that the feature mapping contains more semantic information, reduce the loss of information of small targets such as traffic signs in the deep features, enable the network to learn the feature information of small targets such as traffic signs, and improve the accuracy of model detection. The multi-level information fusion module has a deeper network, and its method can better fit and learn features, and is more helpful for gradient transmission.
[0071] See Figure 5 It can be seen that the definition of the MRI function for multi-scale fusion in the present invention is as follows:
[0072] MRI(X) = Concat(Conv(X), ResBlock(X)) + Conv(X)
[0073] ResBlock(X) = SiLU(Conv(X) + Conv(X)) + X
[0074] Among them, ResBlock is the residual block function, Concat is the connection block function, Conv is the basic convolution block function, SiLU is the silu activation function, and X is the feature map information.
[0075] First of all, the present invention changes the number of channels of the fused feature map through two 1x1 convolutions, making it half of the original channels, and then strengthens the semantic information of the fused feature map by stacking multiple residual fusion feature modules. At the same time, it increases its receptive field to take into account both global information and local information, which is more conducive to the acquisition of context information and improves the information loss in multi-scale information fusion. At the same time, the present invention stacks the residuals of the double-branch spliced image and the input original feature map to strengthen the semantic information of feature maps of different scales, alleviates the information loss that is prone to occur in multi-scale fusion, and effectively improves the robustness of the model.
[0076] III. Double-branch prediction decoupling head
[0077] The present invention studies the GTSDB and CCTSDB datasets and finds that there are a large number of small and medium-sized target traffic sign instances in these traffic sign datasets, and traffic sign object detection has strong requirements for the real-time performance of the task. Therefore, referring to Figure 1 and Figure 7 it can be seen that the present invention directly removes the detection layer for detecting large objects, and at the same time uses the outputs of the feature maps corresponding to the relatively shallower P2 and P3 layers for double-branch prediction decoupling; its Decouple Head decoupling head only uses the first two shallow detection heads output by MRI in the neck network as the final prediction result, and it can be a conventional decoupling head. The detection head used in the present invention is generated from low-level high-resolution feature maps, which combines the semantic information of deep feature maps and is more sensitive to small objects such as traffic signs. It can not only improve the detection accuracy of small objects, but also reduce data redundancy, realize the lightweight of the detection model, and ensure the real-time performance of the detection task.
[0078] Subsequent experimental data and research can show that classification pays more attention to the texture content of the target, localization pays more attention to the edge information of the target, and at the same time, considering the balance between the representation ability of relevant operators and the computational overhead on the hardware, the present invention adopts a decoupled detection head structure to replace the coupled head in Yolov5, thereby accelerating the network convergence speed and improving the accuracy.
[0079] IV. Focal EIOU loss function
[0080] The prediction head in the dual-branch prediction decoupling receives the feature vector and predicts the class, confidence, and bounding box, which correspond to the three parts of the loss in Yolov5, namely the classification loss, the bounding box regression loss, and the object confidence loss. The loss function of Yolov5 can be expressed by the following formula
[0081]
[0082] where λ, μ, and φ are weight parameters, N represents the number of detection heads, h represents the number of objects assigned to the prior boxes by the labels, S represents the number of grids into which the image is segmented, L Box represents the bounding box loss, L obj represents the confidence loss, and L Cls represents the classification loss. Among them, both the classification loss and the object confidence loss are calculated by the binary cross-entropy loss, and their formulas are expressed as follows
[0083]
[0084] y is the predicted value of the model, is the true value of the label, and the bounding box regression loss is calculated by CIOU. Traffic sign detection is a challenging task, and there are some difficult samples such as bad weather and traffic signs being occluded and overlapped in the traffic sign dataset. However, the balance problem of easy and difficult samples is not considered in the CIOU loss calculation, and it is easy for the model to predict more and easier classes during the training process, thus affecting the final detection result. To solve the above problems, the present invention uses Focal EIOU instead of CIOU as the bounding box regression loss calculation. EIOU consists of three parts: the overlapping part, the center distance loss, and the width and height loss. EIOU is shown by the following formula
[0085]
[0086] where ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted box and the true box, h and w represent the width and height of the predicted box, h gt , w gt represent the width and height of the true box, h c , w C are the width and height of the smallest enclosing box covering the two boxes. The calculation formula of Focal EIOU is shown by the following formula
[0087] L Focal-EIO = IOU r L EIOU
[0088] Among them, IOU is the value of the intersection over union of the bounding regression box, and γ is a parameter that controls the degree of outlier suppression.
[0089] It can be seen from this that the present invention designs a cross-stage routing attention module CSB for enhancing feature extraction of small targets of traffic signs. This method can effectively capture the dependencies in the data. The routing attention method used in the present invention can establish the interaction between global information and local information, quickly focus on the most relevant regions, and has high parallelism, which is beneficial to the training and inference of the model. In addition, for the small target of traffic signs, the present invention replaces the network in the original neck structure, designs a novel bidirectional feature fusion network method, and designs an MRI module for enhancing the semantic information of the fused multi-scale feature maps. On the premise of ensuring real-time performance, the detection accuracy is improved and the additional computational overhead is reduced. Finally, for the small target of traffic signs, in order to make full use of the shallow features, compared with the original Yolov5 model, the present invention uses a shallower feature map as the input of the multi-scale feature map, and only uses two shallow detection heads as the final prediction result. In addition, the Focal EIOU loss function can be used instead of CIOU in the double-branch prediction decoupling to calculate the bounding box regression loss to improve the accuracy of the detection result.
[0090] In another embodiment, in order to verify the performance of the Yolov5 object detection method based on the cross-stage routing module and the residual information fusion module proposed above in the present invention, the improved method proposed in the present invention is verified on three datasets: TT-100K, CCTSDB, and GTSDB.
[0091] Table 1 Detailed data of comparison with other trackers on the TT-100K dataset
[0092] Method Input size Backbone AP50 mAP M2Det 800×800 ResNet50 65.6 29.4 Faster R-CNN+FPN 1024×1024 ResNet50 93.1 59.2 Cascade R-CNN 1024×1024 ResNet50 94.4 61.3 RetinaNet 1024×1024 ResNet50 91.3 65.3 EfficientDet-d4 1024×1024 EfficientDet-B4 79.9 61.3 Libra R-CNN 1024×1024 ResNet50 92.4 67.3 ATSS 1024×1024 ResNet50 91.8 66.7 YOLOv3 640×640 Darknet-53 94.0 66.9 YOLOv5 640×640 CSP V5 93.2 66.4 The method of the present invention 640×640 - 96.7 73.0
[0093] Table 2 Detailed data of comparison with other trackers on the GTSDB dataset
[0094]
[0095]
[0096] Table 3 Detailed data of comparison with other trackers on the CCTSDB dataset
[0097] Method Input size Parms Backbone mAP Faster R-CNN 1333×800 25.7M ResNet50 56.58 SSD 1333×800 18.7M ResNet50 49.20 RetinaNet 1333×800 28.3M ResNet50 57.78 Libra R-CNN 1333×800 32.4M ResNet50 61.35 Dynamic R-CNN 1333×800 33.2M ResNet50 60.01 Sparse-R-CNN 1333×800 124M ResNet50 59.65 YOLOv3 640×640 61.5M Darknet-53 82.40 YOLOv4 640×640 45.0M CSPDarknet-53 83.20 YOLOv5 640×640 46.5M Modified CSP V5 84.60 The method of the present invention 640×640 51.2M - 86.1
[0098] Among them, Table 1 shows the detailed data of the comparison between the method of the present invention and other trackers on the TT-100K dataset, Table 2 shows the detailed data of the comparison between the method of the present invention and other trackers on the GTSDB dataset, and Table 3 shows the detailed data of the comparison with other trackers on the CCTSDB dataset. It can be seen from the above Tables 1-3 that the method of the present invention has good performance scores in many evaluation indexes of TT-100K, CCTSDB and GTSDB, and its performance is significantly better than algorithms such as YOLOV5. Therefore, it is particularly suitable for the rapid and high-precision detection of many small targets such as traffic signs.
[0099] In addition, the above Yolov5 object detection method of the present invention based on the cross-stage routing module and the residual information fusion module can be converted into software program instructions, which can be implemented by using a software analysis system including a processor and a memory, or can also be implemented by computer instructions stored in a non-transitory computer-readable storage medium.
[0100] Finally, the method of the present invention is only a preferred implementation, and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module, characterized in that This method improves the original Yolov5 network, specifically including: (1) Cross-stage routing attention module In the Yolov5 network, the C3 modules of the 6th and 8th layers in the backbone network are replaced with cross-stage routing attention modules, and the feature map information of the 2nd, 4th, 6th, and 9th layers is used as input signals P2, P3, P4, and P5 to the neck network in sequence; the composition method of the cross-stage routing attention module is as follows: first, divide the number of channels of the feature map into two parts, the first part enhances the feature information through the attention mechanism, and the other part performs output merging with the enhanced feature through cross-stage, and finally, use the residual structure to perform local enhancement of semantics; (2) Multi-scale feature fusion method In the neck network, it is obtained through the backbone network i represents the i-th feature map extracted from the backbone network, i ∈ {0, 1, 3, 4}, and f0 to f3 respectively correspond to the input signals P5 to P2, where C i ∈ {1024, 512, 256, 128}; S i represents the result of the output after passing through the multi-level information fusion module, and its mathematical formula passing through the multi-scale feature fusion network is expressed as: S0 = f0 S i = MRI(f i , S i-1 ) i = 1, 4 S i = MRI(f i , S i-1 , f i+1 ) for i = 2, 3 S i = MRI(S i-3 , S i-4 ) i = 5, 7 S i = MRI(S i-3 , S i-4 , S i-1 ) i = 6 Among them, the meaning of the MRI function is the multi-scale feature fusion module function, which can fuse multi-scale features of each parameter based on splicing and upsampling; (3) Dual-branch prediction decoupling head In the decoupling head, the two-layer outputs of the feature maps corresponding to the relatively shallower P2 and P3 layers in the neck network are used to perform dual-branch prediction decoupling output as the final prediction result.
2. The Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module according to claim 1, wherein The attention mechanism in the cross-stage routing attention module is the improved BRA attention mechanism, and the composition method of the BRA attention mechanism is as follows: In the BRA attention mechanism, arbitrarily take a unit \(X\in\mathbb{R}\) C×H×W as the input, and \(Y\in\mathbb{R}\) C×H×W as the output. First, divide the feature map into non-overlapping regions of \(P\times P\) blocks, and then flatten these regions along the spatial dimension to obtain feature vectors. Then, input the obtained feature vectors into a linear mapping to derive where \(C\) is the number of channels of the feature map, \(H\) and \(W\) are the width and height of the feature map respectively, and \(P\) is the number of blocks into which the feature map is divided; Q = X r W Q , K = X r W K , V = X r W V The W here Q , W K , W V ∈ R C×C are all parameter matrices, representing the linear mapping weight matrices of Query, Key, and Value in the mapped image. Then, the average value of the regions divided by the eigenvectors of Q and K is obtained A r = Mean(Q) × Mean(K) Then multiply Q r by K r to calculate the most relevant similarity, obtaining S. Use the topK operator on the obtained similarity matrix to retain the indices of the top K regions with the closest relationships, obtaining the region routing index I k ; Finally, connect the adjacent miniatures with the original K and V; K g = gather(I K , K), V g = gather(I K , V) Among them Then K g and V g are used to perform self-attention calculation with the original Q, and at the same time, the obtained V r is used to perform residual calculation with the result of self-attention through depthwise separable convolution; BRA(X) = Attention(Q, K g , V g ) + DwConv(V g ) Get the final result \(O\in\mathbb{R}\) C×H×W , where Mean is the mean function, softmax is the normalization exponential function, gather is the dimension concatenation function, and DwConv is the depthwise separable convolution function.
3. The Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module according to claim 1, wherein The attention mechanism in the cross-stage routing attention module is a double-layer routing attention module.
4. The Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module according to claim 1, characterized in that, The definition of the MRI function is specifically as follows: MRI(X) = Concat(Conv(X), ResBlock(X)) + Conv(X) ResBlock(X) = SiLU(Conv(X) + Conv(X)) + X Among them, ResBlock is the residual block function, Concat is the connection block function, Conv is the basic convolution block function, SiLU is the silu activation function SigmoidLinear Unit, and X is the feature map information.
5. The Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module according to claim 1, characterized in that It also includes: (4) FocalEIOU loss function The prediction head in the dual-branch prediction decoupling receives the feature vector and predicts the category, confidence, and bounding box, which correspond to the three parts of the loss in Yolov5, namely the classification loss, the bounding box regression loss, and the object confidence loss. The loss function of Yolov5 can be expressed by the following formula Among them, λ, μ, and φ are weight parameters, N represents the number of detection heads, h represents the number of targets assigned to prior boxes by labels, S represents the number of grids into which the image is segmented, and L Box represents the bounding box loss, and L Obj represents the confidence loss, and L Cls represents the classification loss; Among them, both the classification loss and the object confidence loss are calculated by the binary cross-entropy loss, and its formula is expressed by the following formula The predicted value of the y model, is the true value of the label, and the bounding box regression loss is calculated by CIOU; Focal EIOU is used instead of CIOU to calculate the bounding box regression loss. EIOU consists of three parts: the overlapping part, the center distance loss, and the width and height loss. EIOU is shown by the following formula: where ρ 2 (b,b gt ) represents the Euclidean distance between the center points of the predicted box and the ground truth box, h and w represent the width and height of the predicted box, h gt ,w gt represent the width and height of the ground truth box, h c ,w C are the width and height of the smallest enclosing box covering the two boxes, and the calculation formula of Focal EIOU is shown as follows L Focal-EIOU = IOU r L EIOU Among them, IOU is the value of the intersection-over-union of the bounding regression box, and γ is the parameter that controls the degree of outlier suppression.
6. The Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module according to claim 1, characterized in that, The detection target of the Yolov5 object detection method based on the cross-stage routing attention module and the residual information fusion module is traffic signs.
7. A Yolov5 object detection system based on a cross-stage routing attention module and a residual information fusion module, characterized in that, It includes: At least one processor; And At least one memory communicatively connected to the processor, where: The memory stores program instructions executable by the processor, and the processor can execute the Yolov5 object detection method based on the cross-stage routing attention module and the residual information fusion module according to any one of claims 1 to 6 by invoking the program instructions.
8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the Yolov5 object detection method based on a cross-stage routing attention module and a residual information fusion module as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Small target detection method based on attention mechanism
CN114202672A
Multilevel semantic fusion cloud and cloud shadow detection method and device, and storage medium
CN114943876A