A Road Traffic Sign Detection Method Based on Image Segmentation
Through the improved UNet segmentation network, road traffic sign detection of multi-scale feature fusion and pyramid split attention module is solved, and detection adaptability problems under complex weather and lighting conditions are achieved, and high-precision and real-time traffic sign detection are achieved.
Patent Information
- Application Number
- CN202411616278.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-11-13
AI Technical Summary
The existing traffic sign detection technology has poor adaptability in complex weather and lighting conditions, resulting in a significant decrease in detection effect.
The improved UNet segmentation network is used to detect road traffic signs. Through multi-scale feature fusion, pyramid split attention module and double-cubic interpolation algorithm, a road traffic sign detection model is constructed and fully supervised learning and training is carried out.
It realizes high-precision traffic sign detection, has strong real-time and adaptability, and is suitable for autonomous driving and intelligent traffic management.
Smart Images

Figure CN119580220B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic sign detection, and particularly to a road traffic sign detection method based on image segmentation. Background Art
[0002] In autonomous driving and intelligent transportation systems, the vehicle's perception ability of the surrounding road environment is crucial. In particular, the recognition and detection of traffic signs on the road (such as lane lines, traffic signs, arrows, stop lines, etc.) are directly related to the safety of autonomous driving vehicles and the accuracy of path planning. Currently, most traffic sign detection technologies rely on traditional image processing methods, such as edge detection, Hough transform, color segmentation, etc. However, these methods have poor adaptability to the environment. Especially in complex weather (such as rainy days, foggy days) and lighting conditions, the detection effect will be significantly reduced. Therefore, it is very necessary to design a road traffic sign detection method based on image segmentation. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a road traffic sign detection method based on image segmentation.
[0004] To achieve the above purpose, the present invention provides the following solutions:
[0005] The present invention provides a road traffic sign detection method based on image segmentation, including:
[0006] Collect road traffic sign images and preprocess them to obtain a road traffic sign data set;
[0007] Construct a road traffic sign detection model based on the improved UNet segmentation network;
[0008] Train the road traffic sign detection model based on the road traffic sign data set;
[0009] Input the collected road traffic sign images into the trained traffic sign detection model for effect testing.
[0010] Preferably, collecting road traffic sign images and preprocessing them to obtain a road traffic sign data set specifically includes:
[0011] Collect road traffic sign images and perform image normalization processing on them;
[0012] Perform data augmentation on the road traffic sign images after image normalization processing;
[0013] Perform manual annotation on the road traffic sign images after data augmentation.
[0014] Preferably, collect road traffic sign images and perform image normalization on them. Specifically:
[0015] Collect road traffic sign images and normalize the collected road traffic sign images to a fixed resolution to complete the image normalization process.
[0016] Preferably, perform data augmentation on the road traffic sign images after image normalization. Specifically:
[0017] Perform image enhancement on the road traffic sign images after image normalization based on the edge enhancement principle and the non-edge enhancement principle.
[0018] Preferably, construct a road traffic sign detection model based on the improved UNet segmentation network. Specifically:
[0019] Improve the UNet segmentation network. On the basis of the original UNet segmentation network, update the decoder in terms of feature fusion so that the decoder fuses feature maps from the encoder at the same scale, as well as feature maps at smaller and larger scales. Among them, decoder X 3 De Fuse through an asymmetric skip connection feature maps from encoder X 1 En and X 2 En Feature maps at a smaller scale, feature maps at the same scale, and decoder X obtained by upsampling through the bicubic interpolation algorithm 4 De Feature maps at a larger scale, decoder X 4 De Fuse in the same operation feature maps from encoder X 2 En and X 3 En Feature maps at a smaller scale, feature maps at the same scale, and decoder X obtained by the previous sampling operation 5 De Feature maps at a larger scale, decoder X 4 De and decoder X 5 De Fuse feature maps at different scales through asymmetric and symmetric skip connections to generate fine-grained and coarse-grained semantic information. Among them, encoder X 1 En and X 2 En The skip connections of are downsampled through a pooling operation combined with a residual module incorporating an attention mechanism to reduce the resolution and convey low-level semantic information at the bottom layer; in the structures of the encoder and decoder, a pyramid split attention module is integrated;
[0020] Build a road traffic sign detection model based on the improved UNet segmentation network.
[0021] Preferably, train the road traffic sign detection model based on the road traffic sign data set, specifically:
[0022] Train the road traffic sign detection model through the road traffic sign data set based on the full-supervised learning method to obtain the trained road traffic sign detection model.
[0023] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0024] The present invention provides a road traffic sign detection method based on image segmentation. The method includes collecting road traffic sign images, preprocessing them to obtain a road traffic sign data set, building a road traffic sign detection model based on the improved UNet segmentation network, training the road traffic sign detection model based on the road traffic sign data set, inputting the collected road traffic sign images into the trained traffic sign detection model for effect testing. The present invention can not only achieve high-precision traffic sign detection, but also has strong real-time performance and adaptability, is applicable to multiple fields such as autonomous driving and intelligent traffic management, and has a wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;
[0027] Figure 2 It is a schematic diagram of the UNet network structure;
[0028] Figure 3 It is a schematic diagram of an example of constructing a feature map by multi-scale skip connections;
[0029] Figure 4 It is a schematic diagram of a residual module integrating an attention mechanism;
[0030] Figure 5 It is a schematic diagram of the structure of a pyramid split attention module;
[0031] Figure 6 It is a schematic diagram of the SEWeight module;
[0032] Figure 7 It is a schematic diagram of the SPC structure;
[0033] Figure 8 It is a schematic diagram of the improved UNet network structure;
[0034] Figure 9 It is a schematic diagram of an example of the Ceymo dataset;
[0035] Figure 10 It is a schematic diagram for comparing the segmentation effects of the UNet structure and the algorithm of the present invention. Detailed implementation manners
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] The purpose of the present invention is to provide a road traffic sign detection method based on image segmentation, which can realize road traffic sign detection based on image segmentation and is convenient to use.
[0038] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0039] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a road traffic sign detection method based on image segmentation, including:
[0040] Step 100: Collect road traffic sign images and preprocess them to obtain a road traffic sign dataset;
[0041] Step 200: Build a road traffic sign detection model based on the improved UNet segmentation network;
[0042] Step 300: Train the road traffic sign detection model based on the road traffic sign dataset;
[0043] Step 400: Input the collected road traffic sign images into the trained traffic sign detection model for effect testing.
[0044] In step 100, collect road traffic sign images and preprocess them to obtain a road traffic sign dataset. Specifically:
[0045] Step 101: Collect road traffic sign images and perform image normalization processing on them;
[0046] Step 102: Perform data augmentation on the road traffic sign image after image normalization;
[0047] Step 103: Manually annotate the road traffic sign image after data augmentation.
[0048] In Step 101, collect the road traffic sign image and perform image normalization on it. Specifically:
[0049] Collect the road traffic sign image and normalize the collected road traffic sign image to a fixed resolution to complete the image normalization process;
[0050] Normalize images from different camera devices or of different sizes to a fixed resolution (such as 512×512 or 256×256) to meet the network input requirements and enhance the adaptability of the model.
[0051] The method for performing data augmentation on the road traffic sign image after image normalization is specifically as follows:
[0052] Perform image enhancement on the road traffic sign image after image normalization based on the edge enhancement principle and the non-edge enhancement principle. The details are as follows:
[0053] After obtaining the image, obtain the grayscale image of the image and equivalent it to a matrix l, the size of which is the number of rows and columns of pixels;
[0054] In the obtained non-planar component image, the regional clarity increases with the increase of the distance from the component area counterpart plane. The gradient of the blurred area is low, and the clarity difference between the edge area and the non-edge area is small, and it is easier to be smoothed in filtering;
[0055] Bilateral filtering takes into account the pixel spatial difference and intensity difference, aiming to preserve the edge as much as possible during filtering. Its formula is:
[0056]
[0057]
[0058] In the above formula, w(i, j, k, l) represents the weight coefficient, and its value depends on the spatial domain kernel d(I, j, k, l) and the range domain kernel r(i, j, k, l);
[0059] After bilateral filtering, since the clear area of each image is different, the retained area is also different. At this time, the edges of different images in the image group are extracted and assigned with the original image sequence number. Common edge extraction operators include Prewitt, Sobel, Laplacian operator and Roberts operator. The Roberts operator has a high accuracy for edge detection, so this operator is selected to extract the edge of the filtered image group. The operator is divided into horizontal and vertical directions, as shown in the formula:
[0060]
[0061] After edge extraction, the edge extracted from the image is marked, and the image sequence number of the edge is marked. After the operation is completed, the edge contour in the target image can be constructed based on the previous data;
[0062] After determining the clear edge area, in order to improve the clarity of the non-edge area, the window range is selected, and the fluctuation of all pixels within the window range is used to characterize the clarity of the central pixel. The variance is used to calculate the value, which is mathematically expressed as:
[0063]
[0064]
[0065] In the formula, I(x, y) represents the gray value of the pixel with coordinates (x, y) in the image, μ(x, y) represents the mathematical expectation of the pixel within a certain range, and V ar (x, y) represents the degree of fluctuation of the pixel with coordinates (x, y) in the image within a certain range. The richer the details, the higher the quality and the clearer the area, the more drastic the fluctuation around a single pixel. ar The larger the value, the more the sliding window traverses all valid images in the image group, and the V stored at the corresponding position can be obtained. ar The matrix of values M1-M k , where moment numbers 1-k correspond to images with the same numbers, the variance values of the same coordinate positions in M1-M are traversed to find the maximum value, and the pixel value of the position in the image with the highest value is assigned to the value of the pixel in the result image;
[0066] [m, address] = max (M1 (x, y), M2 (x, y)...M k (x, y)
[0067] Where m represents the matrix M1-M k The maximum value of the same coordinates in , address is the source image number of the maximum value point;
[0068] L(x, y) = Iaddress (x, y)
[0069] In the formula, the matrix L is the result image, and the value of the coordinate (x, y) in the image with the serial number address is the element value at this position in the result image. After image enhancement, each normalized road traffic sign image is subjected to image enhancement once.
[0070] In step 200, a road traffic sign detection model is constructed based on the improved UNet segmentation network, specifically:
[0071] Improve the UNet segmentation network. First, analyze the UNet model:
[0072] The structure of UNet contains two main stages: encoding and decoding, and the two achieve the transfer of feature information through skip connections. In the encoding stage, the network extracts deep abstract features from the input image using convolutional layers, activation functions, and pooling operations. After entering the decoding stage, the network gradually refines the feature map with the help of convolutional layers, activation functions, and upsampling operations to achieve pixel-level image segmentation. During this process, a copy of the feature map is transmitted from the encoding stage to the decoding stage through skip connections, ensuring the integrity and accuracy of the information. This design enables the decoding stage to utilize more color and texture information contained in the low-level feature map, thereby improving the segmentation accuracy. This structure enables the UNet network to consider both global and local features simultaneously, better capturing the shape and context information of the target. Through the collaborative effect of the encoding and decoding stages, UNet achieves accurate image segmentation while improving the accuracy of the segmentation results while retaining details. The UNet network structure is as Figure 2 shown.
[0073] The UNet network adopts the bilinear interpolation method during the upsampling process. Bilinear interpolation is a technique based on calculating the weighted average of pixel points in a local area, aiming to improve the image resolution. Specifically, during upsampling, UNet selects the four nearest pixel points adjacent to the target pixel point and constructs a square area. Subsequently, the algorithm performs linear interpolation calculations in the horizontal and vertical directions respectively to obtain the value of the target pixel point. Compared with bicubic interpolation, bilinear interpolation has a lower computational complexity, but it can also effectively reduce the generation of artifacts such as jagged edges and stripes. Therefore, bilinear interpolation is widely used in the field of semantic segmentation. By using bilinear interpolation as the upsampling method, the UNet network can maintain the smoothness of the image while ensuring that the reconstructed image details are clearer and more real. This enables UNet to perform excellently in image segmentation tasks. The bicubic interpolation algorithm, as another commonly used image upsampling technique, utilizes more neighboring pixel point information when calculating the value of the target pixel point. When performing bicubic interpolation, first, 16 neighboring pixel points are selected around the target pixel point, and they jointly form a 4×4 rectangular area. Then, two cubic spline interpolation operations are performed in the horizontal and vertical directions respectively. In this way, bicubic interpolation can more accurately estimate the value of the target pixel, thereby improving the quality and accuracy of image interpolation. Bicubic interpolation accurately calculates the value of the target pixel point by using the gray values of neighboring pixel points and their relative position information to the target pixel point. During this process, bicubic interpolation uses a cubic spline function to approximate the curve features, thereby maintaining the smoothness and detail information of the image. Since bicubic interpolation uses more neighboring pixel points and a higher-order interpolation method, compared with bilinear interpolation, it can more precisely retain the details of the image and significantly reduce the generation of artifacts such as jagged edges and stripes. The present invention aims to improve the segmentation accuracy and finally replaces the bilinear interpolation in the original network with the bicubic interpolation algorithm.
[0074] Improve the UNet segmentation network using techniques such as multi-scale feature fusion and pyramid split attention module, and introduce them respectively:
[0075] (1) Multi-scale feature fusion
[0076] The multi-scale feature fusion structure designed in the present invention uses an asymmetric connection different from the original network, which can achieve the fusion of feature maps of different scales between the encoding and decoding structures. During the construction process, a general convolutional layer structure is selected as the basic component unit of the encoder and decoder to efficiently extract feature information in the image. Each convolutional operation consists of two 3×3 convolutional kernels. After the convolutional operation, batch normalization and ReLU activation functions are further used to normalize and non-linearly process the features to enhance the generalization ability and expressive ability of the model. The design of the multi-scale skip connection structure enables the model to capture and fuse feature information of different scales. This structure not only helps the model capture the local details of the image but also enables in-depth understanding of its overall structure, thereby significantly improving the model's parsing ability for complex scenes. Compared with the UNet structure, the decoder designed in the present invention has been improved in terms of feature fusion.
[0077] As Figure 3 shown, the decoder not only fuses the feature maps from the same scale of the encoder but also fuses the feature maps of a smaller scale with the feature maps of a larger scale in the decoder, decoder X 3 De fuses the feature maps from encoder X through an asymmetric skip connection 1 En and X 2 En feature maps of a smaller scale, feature maps of the same scale, and the decoder X obtained by upsampling through the bicubic interpolation algorithm 4 De feature maps of a larger scale, decoder X 4 De fuses in the same operation the feature maps from encoder X 2 En and X 3 En feature maps of a smaller scale, feature maps of the same scale, and the decoder X obtained by the previous upsampling operation 5 De feature maps of a larger scale, decoder X 4 De and decoder X 5 De fuse the feature maps of different scales through asymmetric and symmetric skip connections to generate fine-grained and coarse-grained semantic information, where the skip connections of encoder X 1 En and X 2 En perform downsampling operations to reduce the resolution through a residual module incorporating an attention mechanism combined with a pooling operation to convey the underlying low-level semantic information, where the residual module incorporating the attention mechanism is as Figure 4 shown, as Figure 3and Figure 4 As shown, the feature Figure X 1 En and X 2 En have reduced sizes. During the following skip connection process, the bicubic interpolation algorithm is used to upsample the feature map, doubling the size of the decoder features Figure X 4 De and restoring some details and transmitting high-level semantic information. However, simply unifying the size of the feature map is not enough. It is also necessary to perform a convolution operation using 256 filters of size 3×3 to normalize its quantity in order to eliminate redundant information. After this step, features with 256 channels are obtained Figure X 3 De . The process of constructing the X 4 De feature map is similar. This multi-scale feature fusion method can combine the feature maps of the encoder and decoder, utilize feature information at different levels to improve the model performance, and make it more adaptable to various scenarios and scale changes. The multi-scale skip connection structure proposed in the present invention is an effective feature fusion method that can improve the model's performance at different scales.
[0078] (II) Pyramid Split Attention
[0079] In the field of image processing, the attention mechanism has become a research hotspot. However, traditional attention mechanisms mostly focus on channel attention, often neglecting the importance of spatial information or only limited to the information processing of local regions. The Pyramid Split Attention (PSA) module can extract fine-grained features of multi-scale spatial information while constructing channel dependencies spanning longer distances. By embedding PSA into the neural network, a significant improvement in network performance is achieved while effectively reducing the computational amount. This structure enables the neural network to exhibit higher efficiency and accuracy in image processing.
[0080] 1. PSA Module
[0081] In the present invention, a pyramid split attention module is proposed, and its processing flow is as follows: First, the SPC (Split and Concat) module is used to generate multi-scale feature maps to capture the key information of the image at different scales. Subsequently, the SE (Squeeze-and-Excitation) method is used to generate a channel-level attention vector to extract the channel key information in the multi-scale feature maps, laying a foundation for subsequent attention recalibration. Then, the generated channel attention vector is refined and calibrated through Softmax. Finally, the calibrated attention vector is applied to the multi-scale feature maps, and the processing result is output. Through this process, the PSA module can effectively extract multi-scale spatial information and establish dependencies between channels. Applying the PSA module to a neural network can significantly improve the network performance and accuracy. The detailed structure is as Figure 5 shown.
[0082] 2. SEWeight module
[0083] The channel attention mechanism can optimize the output of feature information mainly because it allows the network to selectively weigh the importance of each channel. A network input feature map is X, X ∈ R H×W×C , where C, H, and W represent the number of channels, height, and width of the feature map, respectively. The key component SE (squeeze and excitation) module that realizes the channel attention mechanism consists of two parts: squeezing and excitation, and its structure is as Figure 6 shown. In the Squeeze operation, a global pooling operation is performed on the input X to obtain the channel statistic z ∈ R C , as shown in the following formula, where X c ∈ R H×W , and based on the formula, the input of H×W×C is converted into an output of 1×1×C, thereby obtaining the global description feature;
[0084]
[0085] In the Excitation operation, a bottleneck structure composed of two fully connected layers is used to improve the generalization performance of the model and reduce the complexity. The number of channels is reduced to 1 / 16 of the original number of channels using the first fully connected layer; the ReLU activation function is used to activate the features to enhance the expression ability of the model; the feature dimension is restored to the initial size through the second fully connected layer to retain sufficient feature information. The Sigmoid function is used to normalize the weights, restricting them to the range of 0 to 1, thereby obtaining the normalized attention weight w c , as shown in the following formula:
[0086] w c= σ(W1δ(W0(z c )))
[0087] Perform a scaling operation to normalize the feature weights of each channel. Through this step, the model can focus more on the main channels, thereby achieving effective adjustment and transformation of features. The finally obtained output after re-adjustment transformation is shown as follows:
[0088] x c = F scale (X c , w c ) = X c w c
[0089] 3. SPC Module
[0090] In the pyramid segmentation attention module, the SPC (Split and Concat) module plays a crucial role in extracting multi-scale features. This module adopts a multi-branch structure, which can effectively extract spatial information from the input feature map at different scales. This design not only enriches the spatial information of the input vector, but also promotes the interaction between local channels, and can also perform multi-scale parallel processing operations. Its structure is as Figure 7 shown. Specifically, the SPC module can capture spatial information at different scales, mainly because it can utilize the multi-scale convolution operation of the pyramid structure. In addition, this module can also extract spatial information at different scales from each channel feature map, because this module also has the function of channel dimension compression. Therefore, the efficiency and accuracy of feature extraction can be significantly improved.
[0091] Assume that the input X is divided into S parts, namely [X0, X1,..., X S-1 . For each part, extract features at different scales to obtain more comprehensive information. After completing feature extraction, these multi-scale features are merged through the Concat operation, and each part maintains channels. In order to reduce the computational burden when processing inputs with different kernel scales, the SPC module adopts a grouped convolution strategy. The key to this method lies in selecting appropriate group size G and convolution kernel size K, which are matched according to a specific relationship. The advantage of doing this is that it can effectively process input vectors of various scales without increasing the model parameters. Through grouped convolution, the SPC module can efficiently extract multi-scale spatial information. These information not only enrich the input features, but also provide strong support for the subsequent channel attention mechanism, where the specific relationship is:
[0092]
[0093] The above process is briefly described as:
[0094] [X0, X1, …, X s-1 = Split(X)
[0095] F i = Conv(k i × k i , G i )(X i ), i = 0, 1, …, S - 1
[0096] F = Cat([F0, F1, …, F s -1])
[0097] Among them, the size of the i-th convolutional kernel is k i = 2×(i + 1) + 1, and the i-th group size is
[0098] The improved network structure is as Figure 8 shown. To achieve multi-scale feature fusion, a residual module with fusion attention is used in the newly introduced skip connection. The reason why the residual module can extract high-level features while retaining the original hierarchical information without changing the scale of the feature map is mainly that it combines convolution and skip connection. This design enables the network to obtain rich feature representations at different levels. At the same time, a channel attention mechanism is also introduced in the skip connection. Channel attention can suppress useless features according to the importance of feature channels and enhance useful features. By calculating the importance weights of feature channels, the network can automatically learn and select the most discriminative features for weighted fusion, thereby improving the accuracy of segmentation. The multi-scale skip connection combines the residual module with fused channel attention, which can make full use of feature information at different levels, improve the performance of the segmentation network, and make target segmentation more accurate.
[0099] In the encoding and decoding structures, a pyramid split attention module is integrated to effectively extract more refined multi-dimensional spatial information and build channel dependencies over long distances. This design enables the neural network containing this module to not only significantly reduce the computational load but also effectively improve the network performance. The bicubic interpolation algorithm can more accurately retain the details of the image compared to bilinear interpolation and significantly reduce the generation of artifacts such as jagged edges and stripes. The present invention aims to improve the segmentation accuracy and finally replaces the bilinear interpolation in the original network with the bicubic interpolation algorithm. To verify the improvement effect of multi-scale fusion, the pyramid attention split module, and the upsampling method on the road traffic sign detection network, the present invention will design multiple different sign detection networks for verification.
[0100] Construct a road traffic sign detection model based on the improved UNet segmentation network.
[0101] In step 300, the road traffic sign detection model is trained based on the road traffic sign data set, specifically as follows:
[0102] The road traffic sign detection model is trained based on the road traffic sign data set by using the fully supervised learning method to obtain the trained road traffic sign detection model.
[0103] The experimental results and result analysis of the present invention are as follows:
[0104] First, the evaluation index settings are introduced:
[0105] In semantic segmentation, the mean intersection over union (MIoU) and the F1 score are usually used for evaluation. The mean intersection over union refers to evaluating the segmentation accuracy of the model by calculating the ratio of the intersection and union between the prediction result and the ground truth. Specifically, the intersection between the prediction result and the ground truth is divided by their union to obtain a value between 0 and 1. Then, the average value is calculated for all images to obtain the mean intersection over union. The F1 score is an index that comprehensively considers the recall rate and the precision rate. It can measure that while the model correctly detects the positive samples, it can also minimize the number of misdetected negative samples as much as possible. Therefore, by comprehensively considering these two evaluation indexes, the semantic segmentation performance of the model can be evaluated more comprehensively and accurately, and it can provide guidance for further model improvement.
[0106] In semantic segmentation, the most important goal is to accurately assign each pixel in the image to the corresponding category. Assuming there is a data set with N + 1 categories (usually including the background category and N foreground categories), the following definitions can be made. p ij represents the total number of pixel points that are misclassified as the j-th category but actually belong to the i-th category, and p ii represents the total number of pixel points that are correctly classified as the i-th category, and p ji represents the total number of pixel points that are correctly classified as the j-th category but misclassified as the i-th category.
[0107] Mean intersection over union (MIoU): The mean intersection over union refers to calculating the intersection over union for each category and taking their average value, which is used to measure the classification accuracy of the model for the pixel points of each category. In this case, the higher the intersection over union, the more accurately the model classifies the pixel points of that category. Therefore, the higher the MIoU (mean intersection over union), the better the model classifies the pixel points of all categories. In other words, MIoU evaluates the average classification accuracy of the model for all categories, and the larger its value, the better the overall classification effect of the model. Therefore, for the semantic segmentation task, MIoU can be used as the main evaluation index to measure the performance of the model. The MIoU formula is shown as follows:
[0108]
[0109] F1 value: The F1 value comprehensively measures the classification performance of the model mainly by taking the harmonic mean of precision and recall. Precision is an indicator to measure the classification accuracy of the model, which represents the proportion of pixel points correctly classified by the model among all classified pixel points. A higher precision means that the model can less frequently misclassify pixel points in the background or other classes as a specific class. The formula is shown as follows.
[0110]
[0111] Recall measures the model's ability to recognize pixel points of the true class, that is, the proportion of pixel points correctly classified as a certain class among all pixel points of that class. The formula is shown as follows.
[0112]
[0113] The calculation method of the F1 value is shown as follows. When both precision and recall are high, the F1 value will be correspondingly high, indicating that the model has achieved good comprehensive performance in the classification task. Therefore, the F1 value is an indicator that comprehensively considers precision and recall and is used to evaluate the overall performance of the model in the pixel point classification task.
[0114]
[0115] In the comparative experiment with other models, F1 and Macro F1-score are selected as evaluation indicators, and the calculation formulas are shown as follows:
[0116]
[0117] Macro F1-score is an indicator used to evaluate the overall performance of the model in multi-class classification problems. It does not consider the frequency of each class in the dataset, but gives the same importance to each class. The method to calculate Macro F1-score is to first calculate the F1 value of each individual class. Once the F1 values of each class are calculated, the average of these values can be used as the Macro F1-score. This means that each class contributes equally to the final score, regardless of its frequency in the dataset. Therefore, if the model performs well on common classes but poorly on rare classes, the Macro F1-score may be low. Macro F1-score is very useful when wanting to evaluate the performance of the model on all classes without any specific class bias. It provides an overall measure of classification accuracy that is not affected by class imbalance. Through Macro F1-score, different road traffic sign detection algorithms can be directly compared according to the performance of the algorithm on each class and scenario, so as to evaluate different algorithms.
[0118] Next, the experimental environment and parameter settings are introduced:
[0119] (1) Dataset and experimental environment
[0120] In this experiment, the existing Ceymo dataset was used for training and testing. The Ceymo dataset contains a total of 2,887 images, Figure 9 showing some examples of the dataset;
[0121] The training set in the dataset has 2,099 images, and the test set has 788 images. The dataset annotates 4,706 road sign instances in 11 categories, and the image resolution is 1920×1080; Table 1 shows the instance distribution of the dataset.
[0122] Table 1 Data distribution table of Ceymo dataset
[0123]
[0124]
[0125] The algorithm development environment adopted in the present invention includes Python 3.8 and CUDA 11.3, and the main algorithm support libraries relied on are PyTorch 1.11.0 and OpenCV-Python. For the specific configuration of the training hardware environment, please refer to Table 2.
[0126] Table 2 Hardware environment configuration table
[0127]
[0128] (2) Parameter settings
[0129] The algorithm training parameters are shown in Table 3 below.
[0130] Table 3 Training parameter settings table
[0131] Parameter Size Epoch 100 Batchsize 4 Backbone Vgg num_classes 12 (number of classes + 1) Learning rate le-4 Optimizer Adam Training set 2309 Test set 578
[0132] Analysis of experimental results:
[0133] (1) Segmentation results of various traffic signs
[0134] The test set of the Ceymo dataset was tested using the trained model. The segmentation results of various traffic signs are shown in Table 5. At the same time, in order to compare the segmentation performance of the model longitudinally, Table 4 presents the result comparison between the algorithm of the present invention and the UNet algorithm. It can be seen from the table that the network model designed by the present invention is relatively stable. When performing segmentation tests on the test set, the MIoU, Precision, Recall, and F1 reached 80.00%, 89.78%, 87.27%, and 88.51% respectively, showing a better improvement compared to the segmentation accuracy of the original network.
[0135] Table 4 Algorithm Comparison Table
[0136] Network name MIoU Precision Recall F1 UNet 74.84 88.69 81.53 84.96 Ourmodel 80.00 89.78 87.27 88.51
[0137] Table 5 Segmentation Results Table of 11 Kinds of Traffic Signs
[0138]
[0139]
[0140] Figure 10 The comparison of the visualization results of the original UNet algorithm and the improved UNet algorithm on the test set is shown. Through the comparison of the visualization results of the two algorithms, the feasibility and effectiveness of the improved algorithm of the present invention can be more intuitively reflected. In this experiment, various situations such as sunny days, nights, and rainy days were selected from the input original images to observe the experimental results. It can be seen from the segmentation effect diagrams that the segmentation results of the model designed by the present invention are more regular and complete, indicating that the method of the present invention plays an obvious role in the process of image segmentation.
[0141] (2) Multi-scale Feature Fusion Experiment
[0142] The purpose of this experiment is to verify the impact of introducing multiscale skip connections (MSC) into the backbone structure of the original UNet network to achieve multi-scale feature fusion on the network performance. This experiment explores the impact of MSC on the UNet network performance by fusing features of different scales. The results are shown in Table 6 below.
[0143] Table 6 Experimental Results Table of Multiscale Skip Connections
[0144] Network name MIoU Precision Recall F1 UNet 74.84 88.69 81.53 84.96 UNet + MSC 77.91 89.10 84.94 86.97
[0145] It can be seen from the experimental results in Table 6 that in terms of the MIoU and F1 metrics, UNet + MSE has improved by 3.07% and 2.01% compared to UNet, reaching 77.91% and 86.97. The experimental results confirm that multi-scale feature fusion has a certain improvement effect on the semantic segmentation accuracy of road traffic signs.
[0146] (3) Efficient Pyramid Split Attention (EPSA) Experiment
[0147] The pyramid split attention experiment mainly verifies the performance impact of concatenating the pyramid split attention module in the middle of the encoding and decoding of the UNet+MSC network. The experimental results are shown in Table 7 below. It can be seen from the experimental results that the pyramid split attention module is effective and can improve the segmentation performance of the model to a certain extent, increasing the MIoU by 1.53% and the F1 by 1.25% on the basis of UNet+MSC for the UNet network.
[0148] Table 7 Results of Pyramid Split Attention Experiment
[0149]
[0150]
[0151] (4) Upsampling Experiment
[0152] The upsampling experiment is to verify the impact of the bicubic interpolation algorithm (Bicubic) of the upsampling method on the model performance. The experimental results are shown in Table 4 below. It can be seen from the experimental results that the change in the upsampling method also affects the network performance.
[0153] Table 8 Results of Upsampling Experiment
[0154] Network name MIoU Precision Recall F1 UNet + MSC + EPSA 79.44 88.92 87.54 88.22 UNet + MSC + EPSA + Bicubic 80.00 89.78 87.27 88.51
[0155] Based on the above experimental results, the MIoU and F1 metrics of the finally improved model reached 80.00% and 88.51%, respectively, with an overall improvement of 5.16% and 3.55% compared to the original network. It can be seen that the method proposed in the present invention is effective in improving the segmentation performance.
[0156] (5) Comparative Experiment
[0157] To verify the segmentation performance of the algorithm of the present invention, the algorithm of the present invention is compared with networks such as object detection and instance segmentation algorithms in the experiments of this section. Table 9 shows the comparison results of several methods on the Ceymo dataset. The F1 value and Macro F1-score of each detection model are listed in Table 9. Among them, SSD-MobileNet-v1 and SSD-MobileNetv2 are object detection methods, and Mask-RCNN-Inception-v2 is an instance segmentation algorithm. It can be seen from the experimental results that the performance of the algorithm of the present invention exceeds that of the former in various types of road traffic signs. Finally, the Macro F1-score reaches 87.53%, which is 6.6%, 4.65%, and 1.78% higher than SSD-MobileNet-v1, SSD-MobileNet-v2, and Mask-RCNN Inception-v2 respectively, achieving good segmentation performance.
[0158] Table 9 Comparison Table of Results of Multiple Algorithms on Ceymo Dataset
[0159]
[0160]
[0161] Through experimental comparison, the present invention demonstrates the effectiveness of the improved model of the present invention and improves the detection performance of road traffic signs.
[0162] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0163] Specific examples are used in the present invention to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A road traffic sign detection method based on image segmentation, characterized in that Including: Collect road traffic sign images, preprocess them, and obtain a road traffic sign dataset; Construct a road traffic sign detection model based on the improved UNet segmentation network, specifically: Improve the UNet segmentation network. On the basis of the original UNet segmentation network, update the decoder in terms of feature fusion, so that the decoder fuses the feature maps from the same scale of the encoder, as well as the feature maps of smaller and larger scales. Among them, the decoder X 3 De fuses the feature maps from the encoder through an asymmetric skip connection X 1 En and X 2 En the feature maps of smaller scales, the feature maps of the same scale, and the decoder obtained by upsampling through the bicubic interpolation algorithm X 4 De the feature maps of larger scales. The decoder X 4 De fuses the feature maps from the encoder in the same operation X 2 En and X 3 En the feature maps of smaller scales, the feature maps of the same scale, and the decoder obtained by the previous upsampling operation X 5 De the feature maps of larger scales. The decoder X 4 De and the decoder X 5 De fuse the feature maps of different scales through asymmetric and symmetric skip connections to generate fine-grained and coarse-grained semantic information. Among them, the encoder X 1 En and X 2 En perform downsampling operations to reduce the resolution through the residual module incorporating the attention mechanism combined with pooling operations in the skip connections, and convey the underlying low-level semantic information; in the structures of the encoder and decoder, the pyramid split attention module is integrated; Construct a road traffic sign detection model based on the improved UNet segmentation network; Train the road traffic sign detection model based on the road traffic sign dataset; Input the collected road traffic sign images into the trained traffic sign detection model for effect testing.
2. The method according to claim 1, wherein Collect road traffic sign images, preprocess them, and obtain a road traffic sign dataset, specifically: Collect road traffic sign images and perform image normalization on them; Perform data augmentation on the road traffic sign images after image normalization; Manually annotate the road traffic sign images after data augmentation.
3. The method according to claim 2, wherein Collect road traffic sign images and perform image normalization on them, specifically: Collect road traffic sign images and normalize the collected road traffic sign images to a fixed resolution to complete the image normalization process.
4. The method according to claim 3, wherein Perform data augmentation on the road traffic sign images after image normalization, specifically: Perform image enhancement on the road traffic sign images after image normalization based on the edge enhancement principle and the non-edge enhancement principle.
5. The method according to claim 1, wherein Train the road traffic sign detection model based on the road traffic sign dataset, specifically: Train the road traffic sign detection model based on the full-supervised learning method through the road traffic sign dataset to obtain the trained road traffic sign detection model.
Citation Information
Patent Citations
Remote sensing image segmentation method combining complete residual and multi-scale feature fusion
CN109447994A
High-precision crack detection method
CN111222580A