A Method for Detecting and Recognizing Traffic Sign Text for Driverless Systems
By building data sets and improving yolov5s, DBNet and CRNN networks, the detection and identification of Chinese traffic signs in complex environments are solved, and the perception and safety of unmanned driving systems are improved.
Patent Information
- Application Number
- CN202510353471.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In the prior art, the identification of Chinese traffic signs is limited, and the text detection accuracy is low in complex natural environments, lack of data sets and environmental adaptability, which affects the perception ability and safety of the unmanned driving system.
Build a data set of Chinese traffic sign text detection and identification, and conduct traffic sign detection, text area detection and identification through improved yolov5s, DBNet and CRNN networks. The network performance is improved by lightweight and attention mechanisms to adapt to complex environments.
It improves the accuracy and real-time nature of traffic sign text detection and recognition, enhances the robustness and perception capabilities of the unmanned driving system, and ensures driving safety.
Smart Images

Figure CN119863780B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision in intelligent transportation, and particularly relates to a method for detecting and recognizing traffic sign texts for an unmanned driving system. Background Art
[0002] In recent years, the construction of transportation infrastructure in China has achieved leapfrog development. At the same time, unmanned driving technology has also made rapid progress. Against this background, traffic signs with text, as a key information medium in the traffic system, carry a large amount of accurate and rich traffic semantic information. This information is crucial for the navigation, decision-making, and behavior planning of unmanned driving systems. Therefore, accurately detecting and recognizing the text information on traffic signs with text is of great significance for improving the perception ability of unmanned driving systems, ensuring driving safety, and optimizing traffic flow management.
[0003] With the rapid development of deep learning technology, significant progress has been made in the detection and recognition technology of text information in natural scenes. However, currently, only a small amount of work focuses on the detection and recognition of traffic signs with text, and the recognition work for Chinese traffic signs is even more limited. In addition, the publicly available Chinese traffic sign recognition datasets are relatively scarce, which restricts the development and application of related technologies. Existing research still has many problems, including the lack of distinction between traffic signs and other signs in complex natural road environments; and the low text detection accuracy under variable natural environments, such as the influence of weather changes and other factors. Summary of the Invention
[0004] Aiming at the problems of the existing technology, the present invention establishes a dataset for detecting and recognizing traffic sign texts in Chinese, and proposes a method for detecting and recognizing traffic sign texts in a complex environment for an unmanned driving system. The algorithm introduces a traffic sign detection module, which can accurately and efficiently identify traffic signs in natural scenes, and improves the text detection and recognition algorithms, which can improve the accuracy and real-time performance of text detection and recognition on traffic signs in complex natural environments, and effectively enhance the system robustness.
[0005] The purpose of the present invention is achieved through the following technical solutions: A method for detecting and recognizing traffic sign texts for an unmanned driving system, comprising the following steps:
[0006] (1) Obtain the to-be-detected images containing traffic signs in natural scenes, and construct a dataset for detecting and recognizing traffic sign texts in Chinese, and use the dataset to train a traffic sign detection network, a text detection network, and a text recognition network;
[0007] (2)Construct a traffic sign detection network improved based on yolov5s, and use the traffic sign detection network improved based on yolov5s to process the input images in natural scenes, and detect the areas of traffic signs therein;
[0008] (3)Construct a text detection network improved based on DBNet, and use the text detection network improved based on DBNet to perform text detection on the traffic sign images in the areas of traffic signs input, and extract the text areas in the traffic signs;
[0009] (4)Construct a text recognition network improved based on CRNN, and use the text recognition network improved based on CRNN to input the text areas after text detection into the recognition network to recognize the text content in the text areas.
[0010] Furthermore, in the step (1), constructing a Chinese traffic sign text detection and recognition dataset specifically includes: collecting and integrating traffic sign images with Chinese texts on the network, annotating the traffic sign images in the dataset, including annotating traffic signs, annotating text areas, and annotating text content, and dividing the annotated images into a training set and a test set.
[0011] Furthermore, the specific construction method of the traffic sign detection network improved based on yolov5s in the step (2) includes the following sub-steps:
[0012] (2.1)Perform lightweight processing on the backbone network in yolov5s, replace the backbone network CSPDarknet53 of yolov5s with the lightweight GhostNet network; replace the SE attention mechanism module in the model with the SRM module to obtain a traffic sign detection network improved based on yolov5s;
[0013] (2.2)Train the traffic sign detection network improved based on yolov5s, output a large number of pictures in the training set of the CTSTD dataset into the improved yolov5s network for training to obtain the optimal weights of the network model, output the pictures in the test set into the improved yolov5s network configured with the optimal weights for testing, and the evaluation indicators of the network performance adopt the average precision AP and the frames per second FPS.
[0014] Furthermore, the specific construction method of the text detection network improved based on DBNet in the step (3) includes the following sub-steps:
[0015] (3.1) Build a text detection network. In the text detection stage, an improved DBNet network is used as the detection model. The DBNet network includes feature extraction, feature fusion, and a Head network. Among them, the feature extraction part uses the ResNet18 residual network as the backbone network. In the conv3, conv4, and conv5 stages of the feature extraction network, all 3x3 convolutional layers are replaced with spatial and channel reconstruction convolutional layers (SCConv); for the feature fusion part, an improved spatial and channel hybrid attention mechanism is used to enhance feature expression, that is, after the last upsampling and feature fusion operations, an attention mechanism that combines spatial attention mechanism and channel attention mechanism (CBAM) is added.
[0016] (3.2) Train the traffic sign text detection network. The traffic sign region images obtained by extracting the images in the CTSTD dataset through the traffic sign detection network are uniformly sized and then input into the text detection network improved based on DBNet for training. After obtaining the optimal weights of the network model, the performance of the network is tested using the test set data. The evaluation metrics for network performance are the F1 value and frames per second (FPS).
[0017] Further, the implementation process of the spatial attention mechanism in step (3.1) is as follows: First, perform max-pooling and average-pooling operations on the feature map output by the channel attention mechanism along the channel dimension respectively to generate features with different context scales; then perform a concatenation operation along the channel dimension, specifically, a 7×7 convolutional layer with an input channel of 2 and an output channel of 1 is used to concatenate the two feature maps, and finally generate the spatial attention weights; finally, the channel attention weights are constrained between 0 and 1 through the Sigmoid activation function.
[0018] Further, the implementation process of the channel attention mechanism in step (3.1) is as follows: First, perform global average pooling and global max-pooling on the input feature map to generate two feature vectors containing the number of channels, and obtain the global features of each channel; then input the above two feature vectors containing the number of channels into a shared multi-layer perceptron to learn the attention weights of each channel. This operation is to construct a fully connected layer through 1×1 convolution to learn the weights of each channel; finally, the channel attention weights are constrained between 0 and 1 through the Sigmoid activation function.
[0019] Further, the specific construction of the text recognition network improved based on CRNN in step (4) includes the following sub-steps:
[0020] (4.1)Construct a text recognition network. The text recognition network uses an improved CRNN network as the recognition model. The CRNN network extracts features through a convolutional layer, learns and predicts through a recurrent layer, and decodes and classifies through a transcription layer. Improve the activation function of the convolutional layer, replacing the original activation function with a Softplus activation function. The Softplus activation function conforms to the biological model of neuron activation and avoids the gradient being 0 during the training process.
[0021] (4.2)Train the traffic sign text detection network. After uniformly sizing the text regions extracted in step (3), input them into the text recognition network improved based on CRNN for training. After obtaining the optimal weights of the network model, use the test set data to test the network performance. The evaluation index text recognition accuracy Accuracy is used to evaluate the recognition accuracy.
[0022] Specifically, in the step (4.1), the Softplus activation function is defined as Softplus( x ) = ln(1 + e x ), which is smooth and has a continuous derivative over the entire real number domain.
[0023] Furthermore, in the step (4.1), the recurrent layer uses a bidirectional long short-term memory network Bi-LSTM to capture the context information in the sequence with two networks in the forward and backward directions respectively, and then fuses the information in the two directions through addition and concatenation. After being processed by Bi-LSTM, a probability distribution will be output at each time step. This distribution contains the prediction probabilities of all characters. These probability distributions form a posterior probability matrix, where each row corresponds to a time step and each column corresponds to a character category. The transcription layer takes this posterior probability matrix as input and uses the CTC loss function to process it. In CTC, an output sequence corresponds to multiple paths. The task of the transcription layer is to find the label sequence with the highest probability combination according to the prediction of each frame, that is, the maximum probability path. After being translated by CTC, the sequence feature information learned by the network is transformed into the final recognized text.
[0024] Specifically, in the step (4.2), the evaluation index text recognition accuracy Accuracy is used to evaluate the recognition accuracy. Its calculation formula is , represents the number of correctly recognized samples, represents the total number of test set samples.
[0025] The beneficial effects of the present invention are as follows:
[0026] A method for detecting and recognizing traffic sign texts for an unmanned driving system is proposed. Through three stages of traffic sign detection, traffic sign text area detection, and traffic sign text recognition, it can effectively extract the text information on traffic signs from the pictures taken by in-vehicle cameras, which is of great significance for improving the perception ability of the unmanned driving system, ensuring driving safety, and optimizing traffic flow management. A target detection network improved based on yolov5s is proposed. By adding the detection of traffic signs before the traditional text detection and recognition algorithms, it can effectively eliminate the interference of other signs in complex natural scenes. At the same time, the network is lightweight processed and the attention mechanism is updated, which can effectively improve the real-time performance and accuracy of traffic sign detection. A text detection network improved based on DBNet is proposed. By performing convolutional reconstruction during the feature extraction process and adding an attention mechanism module during the feature fusion process, the text detection network can accurately extract the traffic sign text area even when facing weather changes in natural scenes. A text recognition network improved based on CRNN is proposed. By improving the activation function of the convolutional layer during the feature extraction process, the training and learning process is made more stable, which can effectively cope with the problem of sharp changes in the brightness of the traffic sign images obtained due to light changes in complex natural scenes and enhance the robustness of the system. Brief Description of the Drawings
[0027] Figure 1 is a schematic diagram of the working process of the present invention;
[0028] Figure 2 is a network structure diagram of GhostNet, the backbone network of the traffic sign detection network improved based on yolov5s;
[0029] Figure 3 is a structure diagram of the SRM attention module;
[0030] Figure 4 is a network structure diagram before and after the improvement of the feature extraction network ResNet18 of the traffic sign text detection network improved based on DBNet;
[0031] Figure 5 is a schematic diagram of the Softplus activation function replaced in the convolutional layer of the traffic sign text recognition network improved based on CRNN. Detailed Embodiments
[0032] The core technology of the present invention is to detect and recognize the text content of traffic signs to achieve accurate and rapid recognition of the text content of traffic signs.
[0033] The present invention proposes a method for detecting and recognizing traffic sign texts for an unmanned driving system, and its process is as Figure 1As shown in the figure, it includes the following steps:
[0034] (1) Obtain the image to be detected containing traffic signs in the natural scene; and construct a Chinese traffic sign text detection and recognition dataset;
[0035] (2) Input the image into the traffic sign detection network improved based on yolov5s that has been trained for detection, and output the traffic sign area;
[0036] (3) Input the obtained traffic sign image into the text detection network that has been trained, and output the text area;
[0037] (4) Input the obtained text area image into the text recognition network that has been trained, and output the text content.
[0038] For the three-stage deep learning networks of the traffic sign detection network, text detection network, and text recognition network improved based on yolov5s mentioned in steps (2)-(4), a dataset is required to train and test the network. Since the currently publicly available Chinese traffic sign text detection and recognition datasets are relatively scarce, a Chinese traffic sign text detection and recognition dataset is constructed by ourselves. The steps to construct the dataset are as follows: Collect and integrate traffic sign images with Chinese text on the Internet, annotate the traffic sign images in the dataset, and divide the annotated images into a training set and a test set. The specific construction method is as follows: a) The image sources of traffic signs mainly include the following parts, the TT100K dataset, the Chinese Traffic Sign Detection Dataset (CCTSDB), the publicly available dataset (CTSU Dataset), and the traffic sign dataset collected by personal shooting. Filter out the pictures containing Chinese text signs from the first two datasets and use the Chinese text traffic sign pictures taken personally as a supplement to construct a new traffic sign text dataset CTSTD (Chinese Traffic Sign Text Dataset). The traffic sign text dataset CTSTD has a total of 2000 original images, which are expanded to 8000 images after data augmentation. Among them, 7000 images are used as the training set and 1000 images are used as the test set.
[0039] b) Annotate the CTSTD dataset, including the annotation of traffic signs, the annotation of text areas, and the annotation of text content. Among them, the annotation of traffic signs is used for the training and testing of the traffic sign detection stage, the annotation of text areas is used for the training and testing of the traffic sign text detection stage, and the annotation of text content is used for the training and testing of the traffic sign text recognition stage.
[0040] Step (2) is the traffic sign detection stage. The detection network uses a traffic sign detection network improved based on yolov5s to process the input images in natural scenes and detect the areas of traffic signs therein.
[0041] The specific construction method of the traffic sign detection network improved based on yolov5s is as follows:
[0042] (2.1) The yolov5s object detection network mainly consists of three parts: the backbone network, the neck network, and the head network. In the present invention, the backbone network CSPDarknet53 of yolov5s is replaced with a more lightweight GhostNet network to reduce the number of network parameters and improve the feature extraction efficiency. The network structure of the GhostNet network is as Figure 2 shown (where Conv2d is the convolution operation with a two-dimensional 1×1 convolution kernel; G-bneck, that is, the Ghost bottleneck residual module). The input image is unified into the format of 224×224×3 and input into the GhostNet network. First, it is processed by a 16-channel ordinary convolution block and then processed through a series of Ghost Bottlenecks. Each Ghost Bottleneck consists of two stacked GhostModules. The first Ghost Module is used to increase the feature dimension, while the second is used to reduce the feature dimension, and finally a feature layer of 7×7×160 is obtained. Then, a 1×1 convolution layer is used to adjust the number of channels to obtain a feature layer of 7×7×960, and then a global average pooling operation is performed. After the pooling is completed, a 1×1 convolution layer is used to adjust the number of channels, and finally, after flattening, a full connection is performed to complete the feature extraction of the image.
[0043] (2.2) At the same time, the SE attention module in the model is replaced with the SRM attention module to improve the recognition accuracy in complex natural scenes. The main composition of the SRM attention module is as Figure 3 shown (where Figure 3The cross - multiplication symbol in it represents the multiplication of the generated weight vector and the original feature map for fusion to complete feature recalibration). This module mainly consists of two parts: Style Pooling and Style Integration. Style Pooling uses global average pooling and global standard deviation pooling to extract style information from each channel of the feature map, which can effectively extract the picture style information. After Style Pooling, SRM calculates the recalibration weights for each channel through Style Integration, uses the style features obtained from Style Pooling to generate specific style weights, and these weights will ultimately recalibrate the feature map to emphasize or suppress their information. The specific implementation of the Style Integration part is to perform channel - full - connection (CFC) on the result of Style Pooling through a 1x1 convolution, and then apply batch normalization (BN) and sigmoid activation function to generate the attention weights for each channel. The SRM attention module enhances the representational ability of CNN by adaptively recalibrating the intermediate feature map, introduces only a small number of parameters, has the characteristics of being lightweight, and its effect is better than SENet.
[0044] (2.3) The specific implementation of the training process of the traffic sign detection network improved based on yolov5s is to output a large number of pictures in the training set of the CTSTD dataset into the improved yolov5s network for iterative training to obtain the optimal weights of the network model, and use the test set to test the improved yolov5s network configured with the optimal weights. The evaluation metrics are the average precision AP and frames per second FPS.
[0045] Step (3) is traffic sign text detection. The detection network uses a text detection network improved based on DBNet to perform text detection on the input traffic sign image and extract the text area in the traffic sign.
[0046] The specific construction method of the traffic sign text detection network improved based on DBNet is as follows:
[0047] (3.1) The DBNet text detection network mainly consists of a feature extraction part, a feature fusion part, and a Head network. The present invention mainly improves the feature extraction part and the feature fusion part of the network.
[0048] (3.2) The feature extraction part uses the ResNet18 residual network as the backbone network. At the conv3, conv4, and conv5 stages of the feature extraction network, all 3x3 convolutional layers are replaced with Spatial and Channel reconstruction Convolution (SCConv) to reduce the redundant information in the feature map, reduce the computational cost and the number of model parameters. The improved ResNet18 network structure is asFigure 4 As shown. The Spatial and Channel Reconstruction Convolution (SCConv) consists of a Spatial Reconstruction Unit (SRU) and a Channel Reconstruction Unit (CRU). The purpose of the SRU is to reduce spatial redundancy. It uses a separation-reconstruction method to achieve this goal. Specifically, the SRU first uses the scaling factor of group normalization to evaluate the information content in different feature maps, and then separates the feature maps with large information content from those with small information content, thereby reducing redundant features in the spatial dimension. The purpose of the CRU is to reduce channel redundancy. It uses a split-transform-fuse method. The CRU first splits the input features into upper-layer and lower-layer features, then performs conversions of global convolution and point convolution on the upper-layer features and similar conversions on the lower-layer features. Finally, feature fusion is performed through adaptive average pooling and the softmax function to obtain the channel-refined features.
[0049] (3.3) The feature fusion part uses an attention mechanism that combines spatial attention mechanism and channel attention mechanism (CBAM) to enhance feature expression. A spatial and channel hybrid attention mechanism is added after the last upsampling and feature fusion operations in the network to improve the accuracy of the text detection network.
[0050] The specific implementation of the spatial attention mechanism is to first perform max-pooling and average-pooling operations on the feature map output by the channel attention mechanism along the channel dimension respectively to generate features with different context scales. Then, a concatenation operation is performed along the channel dimension. Specifically, a 7×7 convolutional layer with an input channel of 2 and an output channel of 1 is used to concatenate the two feature maps, and finally the generated spatial attention weights are obtained. Finally, the Sigmoid activation function is used to constrain the channel attention weights between 0 and 1.
[0051] The specific implementation of the channel attention mechanism is to first perform global average pooling and global max-pooling on the input feature map to generate two feature vectors containing the number of channels, obtaining the global features of each channel. Then, the above two feature vectors are input into a shared multi-layer perceptron to learn the attention weights of each channel. This operation mainly constructs a fully connected layer through 1×1 convolution to learn the weights of each channel. Finally, the Sigmoid activation function is also used to constrain the channel attention weights between 0 and 1.
[0052] (3.4)The specific implementation of the training process of the traffic sign text detection network improved based on DBNet is as follows: The traffic sign area images obtained by extracting the images in the CTSTD dataset through the traffic sign detection network are processed to have a unified size and then input into the text detection network improved based on DBNet for training. After obtaining the optimal weights of the network model, the network performance is tested using the test set data, and the evaluation metrics are the F1 value and the frames per second (FPS) to evaluate the detection accuracy.
[0053] Step (4) is traffic sign text recognition. The recognition network uses a text recognition network improved based on CRNN. The text area after text detection is input into the recognition network to recognize the text content in the text area.
[0054] The specific construction method of the text recognition network improved based on CRNN is as follows:
[0055] The text recognition network improved based on CRNN is divided into three stages: the convolutional layer for feature extraction, the recurrent layer for learning and prediction, and the transcription layer for decoding and classification. The present invention mainly improves the activation function of the convolutional layer for feature extraction. The activation function in the original network is the ReLU activation function, which makes it non-differentiable when the input is less than 0 during training, which may lead to unstable gradients during the training process. Therefore, the present invention chooses to replace it with the Softplus activation function, and the Softplus activation function is defined as Softplus( x ) = ln(1 + e x ), and its image is as Figure 5 shown. The Softplus activation function is smooth over the entire real number domain, that is, its derivative is continuous. This helps to provide more stable gradients during the training process and makes the algorithm of the training and learning process of the feature extraction part more stable. Subsequently, the recurrent layer uses a bidirectional long short-term memory network (Bi-LSTM) to capture the context information in the sequence with two networks in the forward and backward directions respectively, and then the information in the two directions is fused through addition and concatenation. After being processed by Bi-LSTM, a probability distribution will be output at each time step. This distribution contains the prediction probabilities of all possible characters, and these probability distributions form a posterior probability matrix, where each row corresponds to a time step and each column corresponds to a character category. Finally, the transcription layer takes this posterior probability matrix as input and uses the CTC loss function to process it. In CTC, an output sequence can correspond to multiple paths, and the task of the transcription layer is to find the label sequence with the highest probability combination according to the prediction of each frame, that is, the maximum probability path. After the translation of CTC, the sequence feature information learned by the network is transformed into the final recognized text.
[0056] The specific implementation of the training process of the traffic sign text recognition network improved based on CRNN is as follows: The traffic sign text region images extracted by the traffic sign text detection network improved based on DBNet are processed to a unified size and then input into the text recognition network improved based on CRNN for training. After obtaining the optimal weights of the network model, the network performance is tested with the test set data, and the evaluation index uses the text recognition accuracy Accuracy to evaluate the detection. The calculation formula is , represents the number of correctly recognized samples, represents the total number of test set samples.
[0057] Example 1
[0058] The implementation examples of the present invention are implemented on a machine equipped with an Intel Core i9 central processing unit, an NVidia RTX3060 graphics processing unit and 32GB of memory. Three implementation cases are carried out for three improved networks respectively.
[0059] 1. Train and test the traffic sign detection network before improvement and the traffic sign detection network improved based on yolov5s. The CTSTD dataset is selected as the dataset. The learning rate adopts a dynamically changing learning rate, initially set to 0.01, the batchsize is set to 32, and the epoch is 100. The experimental results are shown in Table 1 below:
[0060] Table 1
[0061] Model Original yolov5s network Improved yolov5s network AP (Average Precision) 79.5 83 FPS (Frames Per Second) 74.8 85.3
[0062] The experimental results show that the average precision AP of the model before improvement is 79.5%, and the frames per second FPS is 74.8. After improvement, the average precision has increased by 3.5 percentage points, reaching 83%, and the inference speed has also been improved, and the FPS has increased to 85.3.
[0063] 2. Train and test the traffic sign text detection model before improvement and the improved traffic sign text detection model. The CTSTD dataset is selected as the dataset. All traffic sign pictures are unified in size through preprocessing. The learning rate adopts a dynamically changing learning rate, initially set to 0.005, the batchsize is set to 16, and the epoch is 100. The experimental results are shown in Table 2 below:
[0064] Table 2
[0065] Model Original DBNet network Improved DBNet network F1 value (Harmonic mean of precision and recall) 83.4 86.1 FPS (Frames Per Second) 39.5 40.7
[0066] The F1 value refers to the harmonic mean of Precision and Recall, which is an important comprehensive indicator for evaluating the performance of a model in object detection tasks. Its definition is as follows:
[0067] ;
[0068] Precision reflects the proportion of correctly predicted positive samples, and its definition is as follows:
[0069] ;
[0070] Recall reflects the proportion of samples correctly detected among the actual positive samples, and its definition is as follows:
[0071] ;
[0072] Among them, TP represents the number of samples that are actually positive and predicted to be positive; FP represents the number of samples that are actually negative but predicted to be positive; FN represents the number of samples that are actually positive but predicted to be negative. The experimental results show that the F1 value has increased from 83.4% before improvement to 86.1%, and the FPS has increased from 39.5 before improvement to 40.7. The accuracy and real-time performance of the improved network have been effectively improved.
[0073] 3. Train and test the traffic sign text recognition model before improvement and the improved traffic sign text recognition model. The CTSTD dataset is selected as the dataset. The pictures of all traffic sign text regions are unified in size through preprocessing. The learning rate adopts a dynamically changing learning rate, which is initially set to 0.001, the batchsize is set to 16, and the epoch is 100. The experimental results are shown in Table 3 below:
[0074] Table 3
[0075] Model Original DBNet network Improved DBNet network (Accuracy) Accuracy 73.7% 78%
[0076] The experimental results show that the text recognition accuracy Accuracy has increased from 73.7% to 78%. The accuracy of the improved traffic sign text recognition network has been effectively improved.
[0077] After considering the specification and the content disclosed herein, those skilled in the art will easily think of other implementation schemes of this application. This application aims to cover any variations, uses, or adaptive changes of this application, and these variations, uses, or adaptive changes follow the general principles of this application and include the common general knowledge or conventional technical means in the technical field not disclosed in this application.
[0078] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for traffic sign text detection and recognition for an unmanned driving system, characterized in that It includes the following steps: (1) Obtain the image to be detected containing traffic signs in the natural scene, construct a Chinese traffic sign text detection and recognition dataset, and use the dataset to train a traffic sign detection network, a text detection network, and a text recognition network; (2) Construct a traffic sign detection network improved based on yolov5s, and use the traffic sign detection network improved based on yolov5s to process the input image in the natural scene to detect the area of traffic signs therein; (3) Construct a text detection network improved based on DBNet, and use the text detection network improved based on DBNet to perform text detection on the traffic sign image in the area of the traffic sign input, and extract the text area in the traffic sign; The specific construction method of the text detection network improved based on DBNet includes the following sub-steps: (3.1) Construct a text detection network. In the text detection stage, use the improved DBNet network as the detection model. The DBNet network includes feature extraction, feature fusion, and the Head network. Among them, the feature extraction part uses the ResNet18 residual network as the backbone network. In the conv3, conv4, and conv5 stages of the feature extraction network, all 3x3 convolutional layers are replaced with the spatial and channel reconstruction convolution SCConv; For the feature fusion part, use an improved spatial and channel hybrid attention mechanism to enhance feature expression, that is, add an attention mechanism CBAM that mixes a spatial attention mechanism and a channel attention mechanism after the last upsampling and feature fusion operations; (3.2) Train the traffic sign text detection network. The traffic sign area images extracted from the pictures in the CTSTD dataset by the traffic sign detection network are uniformly sized and then input into the text detection network improved based on DBNet for training. After obtaining the optimal weights of the network model, use the test set data to test the network performance. The evaluation indicators of the network performance are the F1 value and the frames per second FPS; (4) Construct a text recognition network improved based on CRNN, and use the text recognition network improved based on CRNN to input the text area after text detection into the recognition network to recognize the text content in the text area.
2. The traffic sign text detection and recognition method for an unmanned driving system according to claim 1, characterized in that In step (1), to construct the Chinese traffic sign text detection and recognition dataset, specifically: collect and integrate traffic sign images with Chinese texts on the network, annotate the traffic sign images in the dataset, including the annotation of traffic signs, the annotation of text areas, and the annotation of text content, and divide the annotated images into a training set and a test set.
3. A method for traffic sign text detection and recognition for an unmanned driving system according to claim 1, characterized in that, The specific construction method of the traffic sign detection network improved based on yolov5s in step (2) includes the following sub-steps: (2.1) Lightweight processing is performed on the backbone network in YOLOv5s. The backbone network CSPDarknet53 of YOLOv5s is replaced with the lightweight GhostNet network; the SE attention mechanism module in the model is replaced with the SRM module to obtain a traffic sign detection network improved based on YOLOv5s. (2.2) The traffic sign detection network improved based on YOLOv5s is trained. A large number of pictures in the training set of the CTSTD dataset are output to the improved YOLOv5s network for training to obtain the optimal weights of the network model. The pictures in the test set are output to the improved YOLOv5s network configured with the optimal weights for testing. The evaluation metrics for network performance are the average precision AP and the frames per second FPS.
4. A method for traffic sign text detection and recognition for an unmanned driving system according to claim 1, characterized in that, The implementation process of the spatial attention mechanism in step (3.1) is as follows: First, max-pooling and average-pooling operations are respectively performed on the feature map output by the channel attention along the channel dimension to generate features with different context scales. Then, a concatenation operation is performed along the channel dimension. Specifically, two feature maps are concatenated through a 7×7 convolutional layer with an input channel of 2 and an output channel of 1. Finally, the generated spatial attention weights are obtained. Finally, the channel attention weights are constrained between 0 and 1 through the Sigmoid activation function.
5. A traffic sign text detection and recognition method for an unmanned driving system according to claim 1, characterized in that, The implementation process of the channel attention mechanism in step (3.1) is as follows: First, global average pooling and global max-pooling are performed on the input feature map to generate two feature vectors containing the number of channels, and the global features of each channel are obtained. Then, the above two feature vectors containing the number of channels are input into a shared multi-layer perceptron to learn the attention weights of each channel. This operation is to construct a fully connected layer through 1×1 convolution to learn the weights of each channel. Finally, the channel attention weights are constrained between 0 and 1 through the Sigmoid activation function.
6. The traffic sign text detection and recognition method for an unmanned driving system according to claim 1, characterized in that The specific construction of the text recognition network improved based on CRNN in step (4) includes the following sub-steps: (4.1) A text recognition network is constructed. The text recognition network uses the improved CRNN network as the recognition model. The CRNN network performs feature extraction through the convolutional layer, learning and prediction through the recurrent layer, and decoding and classification through the transcription layer. The activation function of the convolutional layer is improved, and the original activation function is replaced with the Softplus activation function. The Softplus activation function conforms to the biological model of neuron activation and avoids the gradient being 0 during the training process. (4.2) The traffic sign text detection network is trained. The text regions extracted in step (3) after being processed to have a unified size are input into the text recognition network improved based on CRNN for training. After obtaining the optimal weights of the network model, the test set data is used to test the network performance. The evaluation metric text recognition accuracy Accuracy is used to evaluate the recognition accuracy.
7. A traffic sign text detection and recognition method for an unmanned driving system according to claim 6, characterized in that In the step (4.1), the Softplus activation function is defined as Softplus( x ) = ln(1 + e x ), which is smooth and has a continuous derivative over the entire real number domain.
8. A traffic sign text detection and recognition method for an unmanned driving system according to claim 6, characterized in that In the step (4.1), the recurrent layer uses a bidirectional long short-term memory network (Bi-LSTM) to capture the context information in the sequence with two networks, namely the forward network and the backward network respectively, and then fuses the information in the two directions through addition and concatenation. After being processed by the Bi-LSTM, a probability distribution is output at each time step. This distribution contains the prediction probabilities of all characters, and these probability distributions form a posterior probability matrix, where each row corresponds to a time step and each column corresponds to a character category. The transcription layer takes this posterior probability matrix as input and uses the CTC loss function to process it. In CTC, one output sequence corresponds to multiple paths. The task of the transcription layer is to find the label sequence with the highest probability combination according to the prediction of each frame, that is, the maximum probability path. After being translated by CTC, the sequence feature information learned by the network is transformed into the final recognized text.
9. A traffic sign text detection and recognition method for an unmanned driving system according to claim 6, characterized in that, In the step (4.2), the evaluation index text recognition accuracy Accuracy is used to evaluate the recognition accuracy, and its calculation formula is , represents the number of correctly recognized samples, represents the total number of test set samples.
Citation Information
Patent Citations
Character recognition method and device for traffic sign board, equipment and storage medium
CN113971792A
Method and system for quickly identifying and positioning mature apples and apple picking equipment
CN116704353A