A Traffic Sign Detection Method Based on Multi-Scale Feature Fusion

By adopting a multi-scale feature fusion method based on YOLOv8 network in traffic sign detection, combining color difference characteristics and multi-view collaborative fusion characteristics, the problem of insufficient speed and accuracy of traffic sign detection in the prior art is solved, and more efficient and accurate traffic sign recognition is achieved.

CN119904842BActive Publication Date: 2025-06-10MO NI XUEDI (JIANGXI) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398954.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-10
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing traffic sign detection and identification algorithms have shortcomings in detection speed and accuracy, especially in complex traffic road environments, which are difficult to effectively detect and identify traffic signs.

Method used

The traffic sign detection method based on YOLOv8 network is adopted, and the color difference feature extraction module, backbone feature extraction module, multi-view collaborative fusion module and detection head are used to detect the traffic signs in combination with the color difference feature and the multi-view collaborative fusion feature to achieve efficient identification of traffic signs.

Benefits of technology

It improves the accuracy of the recognition of traffic signs, enhances the model's ability to express the color characteristics of traffic signs, reduces the impact of light on the recognition results, simplifies the model structure, reduces the amount of model parameters, and improves the efficiency of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904842B_ABST
    Figure CN119904842B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic sign detection method based on multi-scale feature fusion, including S1: constructing a data set, which includes a number of traffic sign pictures and corresponding labels; S2: constructing a detection model, importing the pictures in the data set in S1 into a color difference feature extraction module to obtain color difference features; S3: importing the color difference features in S2 into backbone feature extraction to obtain extraction features of different scales; S4: importing the extraction features in S3 into a multi-view collaborative fusion module to obtain multi-view assisted fusion features; S5: performing residual connection on the multi-view collaborative fusion features and the color difference features and inputting them into a detection head to obtain a prediction output; S6: optimizing model parameters. The present invention performs residual connection on the color difference features and the multi-view collaborative fusion features, thereby performing image completion in the recognition process by combining the color information and visual information of traffic sign pictures, so as to obtain more accurate detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and specifically relates to a traffic sign detection method based on multi-scale feature fusion. Background Art

[0002] The development of autonomous driving is of great significance for reducing traffic accidents, and traffic sign detection and recognition is an important part of autonomous driving. The significance of traffic sign detection and recognition is mainly reflected in the following aspects: improving traffic safety, the intelligence of traffic management, and the development of automatic timing and counting. However, how to efficiently, accurately, and automatically detect and recognize numerous traffic signs from complex traffic roads is a difficult problem at present.

[0003] Traffic sign detection and recognition is an important application scenario of object detection and recognition. In the development process of traffic sign detection and recognition technology, deep learning-based traffic sign detection and recognition algorithms have become the main research direction. While the detection and recognition accuracy is continuously improved, the number of parameters of the network model also increases accordingly, which in turn leads to a decrease in the detection speed, exacerbates the difficulty of model deployment, and there are also problems of insufficient detection and recognition accuracy caused by the small size characteristics of traffic signs and the complex traffic road environment background. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention proposes a traffic sign detection method based on multi-scale feature fusion, including the following steps:

[0005] Step S1: Construct a data set, which includes several traffic sign pictures and corresponding labels;

[0006] Step S2: Construct a detection model based on the YOLOv8 network. The model includes a color difference feature extraction module, a backbone feature extraction module, a multi-view collaborative fusion module, and a detection head. Import the pictures in the data set in Step S1 into the color difference feature extraction module to obtain color difference features;

[0007] Step S3: Import the color difference features in Step S2 into the backbone feature extraction module to obtain extraction features of different scales;

[0008] Step S4: Import the extraction features of different scales in Step S3 into the multi-view collaborative fusion module to obtain multi-view assisted fusion features; in this process, the multi-view assistance module performs data augmentation on the extraction features to obtain multi-view features from different perspectives and multi-distance features at different distances, and uses the multi-view features as supplementary semantics to supplement the multi-distance features to obtain multi-view collaborative features of different scales, and performs channel fusion on the multi-view collaborative features of different scales to obtain multi-view collaborative fusion features;

[0009] Step S5: Perform dimensionality conversion on the multi-view collaborative fusion features in Step S4 and the color difference features in Step S2, adjust the multi-view collaborative fusion features and the color difference features to a unified dimension, perform residual connection on the adjusted multi-view collaborative fusion features and the color difference features, and then input them into the detection head to obtain the prediction output;

[0010] Step S6: Construct a loss function and minimize the loss function to optimize the model parameters.

[0011] Further, Step S2 is specifically as follows:

[0012] Step S21: The color difference feature extraction module processes the red color difference map and the blue color difference map of the input image by assigning different weight coefficients based on the three-component color difference method, and calculates the red, blue, and yellow color difference maps of the image, expressed as:

[0013] ;

[0014] ;

[0015] ;

[0016] where, represents the intensity value of the red component of the input image at position , represents the intensity value of the blue component of the input image at position , represents the intensity value of the green component of the input image at position , represents the red color difference map, represents the blue color difference map, represents the yellow color difference map;

[0017] Step S22: Adaptively adjust the weights of the red color difference map, the blue color difference map, and the yellow color difference map according to the image features of the input image, and calculate the weights based on the image features ; the image features include the illumination, contrast, and color distribution of the input image;

[0018] Step S23: Use the adaptive weights to perform weighted stitching on the red color difference map, the blue color difference map, and the yellow color difference map to obtain the color difference features.

[0019] Further, Step S4 is specifically as follows:

[0020] Step S41: The multi-view assistance module performs data augmentation on the extracted features to respectively obtain perspective supplementary features from different perspectives and view supplementary features at different distances; wherein, the data augmentation includes operations of scaling, enlarging, diagonal flipping, and vertical flipping.

[0021] Step S42: Downsample the perspective supplementary features from different perspectives, align them with the original feature map, i.e., the corresponding extracted features, then perform convolution on the perspective supplementary features and the extracted features to compress the channel dimension, obtain perspective compressed features, then input the perspective compressed features into a multi-modal multiplication unit of three parallel tensors to calculate perspective attention weights, and based on the attention weights, splice the perspective supplementary features and the extracted features to obtain multi-perspective features.

[0022] Step S42: Downsample the distance supplementary features at different distances, align them with the original feature map, i.e., the corresponding extracted features, then perform convolution on the distance supplementary features and the extracted features to compress the channel dimension, obtain distance compressed features, then input the distance compressed features into a multi-modal multiplication unit of three parallel tensors different from those in Step S41 to calculate distance attention weights, and based on the attention weights, splice the distance supplementary features and the extracted features to obtain multi-distance features.

[0023] Step S43: Use the multi-perspective features as supplementary semantics to supplement the multi-view features to obtain multi-view collaborative features at different scales; specifically:

[0024] Successively process the multi-perspective features through channel attention and spatial attention to obtain supplementary features, splice the supplementary features and the multi-distance features in the channel dimension, and then further fuse them through convolution to obtain multi-view collaborative features at different scales.

[0025] Step S44: Perform channel fusion on the multi-view collaborative features at different scales to obtain multi-view collaborative fusion features.

[0026] Further, the process of obtaining multi-perspective features in Step S42 is expressed as:

[0027] ;

[0028] ;

[0029] ;

[0030] ;

[0031] ;

[0032] Wherein, represents the perspective supplementary features, represents the linear activation function, Indicates a splicing operation, Indicates a convolution operation, Indicates feature extraction, Indicates the perspective supplementary feature under the first perspective, Indicates the perspective supplementary feature under the second perspective, Indicates the perspective compression feature, Indicates the Sigmoid activation function, 、 and Indicates the weight matrix of the first multi - modal multiplication unit, Indicates the perspective attention weight of the first multi - modal multiplication unit, 、 and Indicates the weight matrix of the second multi - modal multiplication unit, Indicates the perspective attention weight of the second multi - modal multiplication unit, 、 and Indicates the weight matrix of the third multi - modal multiplication unit, Indicates the perspective attention weight of the third multi - modal multiplication unit, Indicates the multi - perspective feature, Indicates the element - wise product.

[0033] Furthermore, step S44 is specifically as follows:

[0034] Divide the multi - view collaborative features of different scales into n blocks along the channel dimension, and the number of channels in each block is the same; divide each block of multi - view collaborative features into two parts along the channel dimension again, one part is used for channel - based interaction with the multi - view collaborative features of adjacent blocks, and the other part is used for output; recombine and splice the outputs of all blocks to obtain the multi - view collaborative fusion feature.

[0035] Furthermore, the backbone feature extraction module in step S2 includes a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a third C2f module, a fifth convolutional layer, a fourth C2f module, and an SPPF module connected in series in sequence; among them, the second C2f module, the third C2f module, and the SPPF module respectively output extraction features of different scales.

[0036] Furthermore, the second C2f module, the third C2f module, and the fourth C2f module adopt deformable convolutions to perform dynamic sampling on the input.

[0037] 1) By extracting the color difference features of traffic sign images and using the color difference features as residual inputs for detection, the present invention can more comprehensively retain color information, enhance the model's ability to express the color features of traffic signs, and thus improve the recognition accuracy. Further, under different lighting conditions, the colors of traffic sign images may change. By using the color difference features as residual inputs, the model can better adapt to lighting changes, reduce the impact of lighting on the recognition results, help the model better focus on the color features of the traffic signs themselves, thereby enhancing the anti-interference ability of the model and also simplifying the model structure to a certain extent and reducing the number of model parameters. This makes the model easier to converge during the training process, with a shorter training time, and thus improves the efficiency of model training.

[0038] 2) The multi-view collaborative fusion module of the present invention performs implicit interaction and semantic association mining on features at different perspectives and distances through two-stage attention, well compensating for the differences between them, thereby making the distinction between traffic sign images and the background environment clearer.

[0039] 3) By performing residual connection on the color difference features and the multi-view collaborative fusion features, the present invention combines the color information and visual information of traffic sign images to perform image completion during the recognition process, thereby obtaining more accurate detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flowchart of the steps of a traffic sign detection method based on multi-scale feature fusion. DETAILED DESCRIPTION OF THE INVENTION

[0041] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, in the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "No. 1", "No. 2", "No. 3" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. The present invention will be further described below in conjunction with the specific embodiments.

[0042] Refer to Figure 1 , a traffic sign detection method based on multi-scale feature fusion, comprising the following steps:

[0043] Step S1: Construct a data set, which includes a number of traffic sign images and their corresponding labels;

[0044] Step S2: Build a detection model based on the YOLOv8 network. The model includes a color difference feature extraction module, a backbone feature extraction module, a multi-view collaborative fusion module, and a detection head. Import the pictures in the dataset in Step S1 into the color difference feature extraction module to obtain color difference features;

[0045] Step S3: Import the color difference features in Step S2 into the backbone feature extraction module to obtain extraction features of different scales;

[0046] Step S4: Import the extraction features of different scales in Step S3 into the multi-view collaborative fusion module to obtain multi-view assisted fusion features; in this process, the multi-view assistance module performs data augmentation on the extraction features to obtain multi-view features from different perspectives and multi-distance features at different distances, and uses the multi-view features as supplementary semantics to supplement the multi-distance features to obtain multi-view collaborative features of different scales. Channel fusion is performed on the multi-view collaborative features of different scales to obtain multi-view collaborative fusion features;

[0047] Step S5: Perform dimensional conversion on the multi-view collaborative fusion features in Step S4 and the color difference features in Step S2, adjust the multi-view collaborative fusion features and the color difference features to a unified dimension, perform residual connection on the adjusted multi-view collaborative fusion features and the color difference features, and then input them into the detection head to obtain the predicted output;

[0048] Step S6: Build a loss function and minimize the loss function to optimize the model parameters.

[0049] Further, Step S2 is specifically as follows:

[0050] Step S21: The color difference feature extraction module processes the red color difference map and the blue color difference map of the input picture based on the three-component color difference method by assigning different weight coefficients, and calculates the red, blue, and yellow color difference maps of the picture, expressed as:

[0051] ;

[0052] ;

[0053] ;

[0054] Among them, represents the intensity value of the red component of the input picture at position , represents the intensity value of the blue component of the input picture at position , represents the intensity value of the green component of the input picture at position , represents the red color difference map, represents the blue color difference map, Indicates the yellow color difference map;

[0055] Step S22: According to the image features of the input image Adaptive adjustment of the weights of the red color difference map, blue color difference map, and yellow color difference map, and calculation of the weights based on the image features The image features Include the illumination, contrast, and color distribution of the input image;

[0056] Step S23: Use the adaptive weights to perform weighted stitching on the red color difference map, blue color difference map, and yellow color difference map to obtain the color difference features.

[0057] Further, step S4 is specifically as follows:

[0058] Step S41: The multi-view assistance module performs data augmentation on the extracted features to respectively obtain the view supplementary features from different perspectives and the view supplementary features at different distances; among them, the data augmentation includes operations such as scaling, enlargement, diagonal flipping, and vertical flipping;

[0059] Step S42: Downsample the view supplementary features from different perspectives, align them with the original feature map, that is, the corresponding extracted features, then perform convolution on the view supplementary features and the extracted features to compress the channel dimension to obtain the view compressed features, then input the view compressed features into the multi-modal multiplication unit of three parallel tensors to calculate the view attention weights, and splice the view supplementary features and the extracted features based on the attention weights to obtain the multi-view features;

[0060] Step S42: Downsample the distance supplementary features at different distances, align them with the original feature map, that is, the corresponding extracted features, then perform convolution on the distance supplementary features and the extracted features to compress the channel dimension to obtain the distance compressed features, then input the distance compressed features into the multi-modal multiplication unit of three parallel tensors different from those in step S41 to calculate the distance attention weights, and splice the distance supplementary features and the extracted features based on the attention weights to obtain the multi-distance features;

[0061] Step S43: Use the multi-view features as supplementary semantics to supplement the multi-view features to obtain the multi-view collaborative features at different scales; specifically:

[0062] Successively process the multi-view features through channel attention and spatial attention to obtain the supplementary features, splice the supplementary features and the multi-distance features in the channel dimension, and then further fuse them through convolution to obtain the multi-view collaborative features at different scales;

[0063] Step S44: Perform channel fusion on the multi-view collaborative features at different scales to obtain the multi-view collaborative fusion features.

[0064] Further, the process of obtaining multi-view features in step S42 is expressed as:

[0065] ;

[0066] ;

[0067] ;

[0068] ;

[0069] ;

[0070] Among them, represents the view supplementary feature, represents the linear activation function, represents the concatenation operation, represents the convolution operation, represents the feature extraction, represents the view supplementary feature under the first view, represents the view supplementary feature under the second view, represents the view compression feature, represents the Sigmoid activation function, , and represent the weight matrix of the first multi-modal multiplication unit, represents the view attention weight of the first multi-modal multiplication unit, , and represent the weight matrix of the second multi-modal multiplication unit, represents the view attention weight of the second multi-modal multiplication unit, , and represent the weight matrix of the third multi-modal multiplication unit, represents the view attention weight of the third multi-modal multiplication unit, represents the multi-view feature, represents the element-wise product.

[0071] Further, step S44 is specifically:

[0072] Divide the multi-view collaborative features of different scales into n blocks along the channel dimension, and each block has the same number of channels; divide each block of multi-view collaborative features into two parts along the channel dimension again, one part is used for channel-based interaction with the multi-view collaborative features of adjacent blocks, and the other part is used for output; recombine and splice the outputs of all blocks to obtain the multi-view collaborative fusion feature.

[0073] Further, the backbone feature extraction module in step S2 includes a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fifth convolutional layer, a third C2f module, a fourth convolutional layer, a fourth C2f module, and an SPPF module connected in series in sequence; among them, the second C2f module, the third C2f module, and the SPPF module respectively output extracted features of different scales.

[0074] Further, the second C2f module, the third C2f module, and the fourth C2f module adopt deformable convolutions to perform dynamic sampling on the input.

[0075] As described above in the specific implementation manner, the purpose, technical solution, and beneficial effects of the present invention have been further described in detail. It should be understood that the above is only the specific implementation manner of the present invention and is not used to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A traffic sign detection method based on multi-scale feature fusion, characterized in that: The following steps are involved: Step S1: construct a data set, which includes several traffic sign images and corresponding labels; Step S2: construct a detection model based on the YOLOv8 network, which includes a color difference feature extraction module, a backbone feature extraction module, a multi-view collaborative fusion module and a detection head, and import the image of the data set in step S1 into the color difference feature extraction module to obtain the color difference feature; Step S3: importing the color difference feature in step S2 into the backbone feature extraction module to obtain extraction features of different scales; Step S4: importing the extracted features of different scales in step S3 into the multi-view collaborative fusion module to obtain multi-view assisted fusion features; In this process, the multi-view assistance module performs data enhancement on the extracted features to obtain multi-view features at different viewpoints and multi-distance features at different distances, and uses the multi-view features as supplementary semantics to supplement the multi-distance features to obtain multi-view collaborative features at different scales. The multi-view collaborative features at different scales are channel-fused to obtain multi-view collaborative fusion features. Step S5: performing dimension conversion on the multi-view collaborative fusion features in step S4 and the color difference features in step S2, adjusting the multi-view collaborative fusion features and the color difference features to a unified dimension, performing residual connection on the adjusted multi-view collaborative fusion features and the color difference features, and inputting them into the detection head to obtain a prediction output; Step S6: Construct a loss function and minimize the loss function to optimize the model parameters.

2. The traffic sign detection method based on multi-scale feature fusion according to claim 1, characterized in that: Step S2 is specifically as follows: Step S21: The color difference feature extraction module assigns different weight coefficients to the red color difference map and the blue color difference map of the input image based on the three-component color difference method, and calculates the red, blue, and yellow color difference maps of the image, which are expressed as: ; ; ; in, Indicates that the input image is at position The intensity value of the red component, Indicates that the input image is at position The intensity value of the blue component of Indicates that the input image is at position The intensity value of the green component of represents the red color difference map, represents the blue color difference diagram, represents the yellow color difference diagram; Step S22: Based on the image features of the input image Adaptively adjust the weights of the red color difference map, blue color difference map, and yellow color difference map based on image features To calculate the weight; picture features This includes the lighting, contrast, and color distribution of the input image; Step S23: weighted concatenation of the red color difference map, the blue color difference map, and the yellow color difference map is performed using adaptive weights to obtain color difference features.

3. The traffic sign detection method based on multi-scale feature fusion according to claim 1, characterized in that: Step S4 is specifically as follows: Step S41: the multi-view assisting module performs data enhancement on the extracted features to obtain supplementary features of different viewing angles and supplementary features of views at different distances; wherein the data enhancement includes scaling, enlarging, diagonal flipping and vertical flipping operations; Step S42: down-sampling the perspective supplementary features of different perspectives, aligning them with the original feature map, that is, the corresponding extracted features, and then convolving the perspective supplementary features and the extracted features to compress the channel dimension to obtain the perspective compression features, and then inputting the perspective compression features into the multivariate modular multiplication unit of three parallel tensors to calculate the perspective attention weights, and splicing the perspective supplementary features and the extracted features based on the attention weights to obtain multi-perspective features; Step S42: down-sample the distance supplementary features of different perspectives, align them with the original feature map, that is, the corresponding extracted features, and then convolve the distance supplementary features and the extracted features to compress the channel dimension to obtain the distance compression features, and then input the distance compression features into three multivariate modular multiplication units of parallel tensors different from step S41 to calculate the distance attention weights, and splice the distance supplementary features and the extracted features based on the attention weights to obtain multiple distance features; Step S43: using the multi-view feature as a supplementary semantic to supplement the multi-view feature so as to obtain multi-view collaborative features of different scales; specifically: The multi-view features are processed by channel attention and spatial attention in turn to obtain supplementary features, and the supplementary features are concatenated with the multi-distance features in the channel dimension, and then further fused by convolution to obtain multi-view collaborative features of different scales; Step S44: performing channel fusion on multi-view collaborative features of different scales to obtain multi-view collaborative fusion features.

4. The traffic sign detection method based on multi-scale feature fusion according to claim 3 is characterized in that: The process of obtaining multi-view features in step S42 is expressed as: ; ; ; ; ; in, represents the perspective supplement feature, represents the linear activation function, Represents a splicing operation, represents the convolution operation, represents the extracted features, represents the perspective supplement feature under the first perspective, represents the perspective supplement feature under the second perspective, represents the visual compression feature, represents the Sigmoid activation function, , and represents the weight matrix of the first multivariate modular multiplication unit, represents the visual attention weight of the first multivariate modular multiplication unit, , and represents the weight matrix of the second multivariate modular multiplication unit, represents the perspective attention weight of the second multivariate modular multiplication unit, , and represents the weight matrix of the third multivariate modular multiplication unit, represents the perspective attention weight of the third multivariate modular multiplication unit, Represents multi-view features, Represents element-wise product.

5. The traffic sign detection method based on multi-scale feature fusion according to claim 3 is characterized in that: Step S44 is specifically as follows: The multi-view collaborative features of different scales are divided into n blocks along the channel dimension, and the number of channels in each block is the same; each block of multi-view collaborative features is further divided into two parts along the channel dimension, one part is used for channel-based interaction with the multi-view collaborative features of adjacent blocks, and the other part is used for output; the outputs of all blocks are re-joined to obtain multi-view collaborative fusion features.

6. The traffic sign detection method based on multi-scale feature fusion according to claim 1, characterized in that: The backbone feature extraction module in step S2 includes a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a third C2f module, a fifth convolutional layer, a fourth C2f module and an SPPF module connected in series in sequence; wherein the second C2f module, the third C2f module and the SPPF module respectively output extracted features of different scales.

7. The traffic sign detection method based on multi-scale feature fusion according to claim 6 is characterized in that: The second C2f module, the third C2f module, and the fourth C2f module use deformable convolution to dynamically sample the input.

Citation Information

Patent Citations

  • High-resolution traffic sign rapid detection method based on adaptive region screening

    CN117132962A

  • Lightweight network-based traffic sign recognition method

    WO2022205685A1