A traffic sign target detection and recognition method and system

By improving the adaptive feature fusion module and character coordinate clustering algorithm of the YOLO model, combined with feature mapping and two-time feature vector fusion, the problem of low accuracy in detecting and recognizing small targets and road signs in extreme weather conditions is solved, and the detection and recognition effect of the autonomous driving system is improved.

CN119380311BActive Publication Date: 2025-10-17INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310926610.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-26
Publication Date
2025-10-17
Estimated Expiration
2043-07-26

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in detecting and recognizing traffic signs for small targets and in extreme weather conditions. In particular, in complex and changing road environments, it is difficult to accurately detect the road sign information in front of the vehicle.

Method used

A road sign character detection model is built with an improved YOLO model and an adaptive feature fusion module. The character coordinate clustering algorithm and feature mapping are combined to improve the detection and recognition accuracy by fusing feature vectors at two moments.

Benefits of technology

The accuracy of road sign detection and recognition for small targets and in extreme weather conditions has been enhanced, improving the stability and accuracy of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380311B_ABST
    Figure CN119380311B_ABST
Patent Text Reader

Abstract

The application provides a traffic sign target detection and recognition method and system, the method comprising: S1, detecting a pretreated image at time t by using a detection network model to obtain a sign image block at time t; S2, extracting character information in the sign image block at time t by using a sign character detection model; S3, clustering the character information at time t and sorting the character information according to semantics based on the clustering result to obtain character semantic information at time t; S4, finding the same traffic sign as the character semantic information at time t in an image at time t+1 by using a feature mapping method; S5, detecting a pretreated image at time t+1; S6, finding the same traffic sign as the character semantic information at time t+1 in an image at time t+2 by using the method as described in S2-S4; S7, calculating the position and size of a traffic sign bounding box to obtain traffic sign images at two times; S8, extracting feature vectors of the traffic sign images and fusing, and classifying and recognizing the fused feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of automatic driving, in particular to the field of detection data processing technology in the field of automatic driving, and more particularly to a traffic sign target detection and recognition method and system. BACKGROUND

[0002] In recent years, the automatic driving technology has developed rapidly, and traffic sign detection and recognition is a key technology in the intelligent traffic automatic driving system, which has many practical applications in the fields of advanced auxiliary driving, automatic driving, human-computer interaction, etc. With the rapid development of deep convolutional neural networks, the vehicle-mounted visual detection and recognition technology has also developed rapidly.

[0003] At present, most traffic sign detection and recognition methods have achieved good results in the case of close distance detection, but in the case of small size of road signs in the vehicle driving road environment or extremely blurred video shooting due to extreme weather, the detection and recognition accuracy of these algorithms is poor, resulting in low traffic sign detection and recognition rate. Moreover, due to the complex and changeable traffic scene, the interference in the road is changeable, such as the light intensity caused by weather, road congestion or shielding, etc. These problems will also cause the automatic driving system to be unable to accurately detect the road sign information in front of the current vehicle.

[0004] Therefore, how to improve the detection and recognition accuracy of traffic signs under small targets and extreme weather is a problem to be solved. SUMMARY

[0005] Therefore, the purpose of the present application is to overcome the defects of the prior art, and to provide a traffic sign target detection and recognition method and system.

[0006] The purpose of the present application is achieved by the following technical solutions:

[0007] According to a first aspect of the present application, a traffic sign target detection and recognition method is provided, the method comprising the following steps: step S1, acquiring a video image, extracting a video image at time t for pre-processing to obtain a pre-processed image at time t; and using a trained detection network model to detect the pre-processed image at time t to obtain a sign image block at time t; step S2, using a trained sign character detection model to extract character information in the sign image block at time t, wherein the trained sign character detection model is obtained by improving an original YOLO model, and an adaptive feature fusion module for fusing a sign image feature map at time t is added in the original YOLO model; step S3, clustering the character information in the sign image block at time t to obtain an initial clustering center, and sorting the character information in the sign image block at time t according to semantics based on the initial clustering center to obtain character semantic information of the sign image block at time t; step S4, using a feature mapping method to find a traffic sign with the same character semantic information as the sign image block at time t in an image at time t+1; step S5, in the case of finding the same traffic sign, using the trained detection network model to detect the pre-processed image at time t+1 to obtain a sign image block at time t+1; step S6, based on the obtained sign image block at time t+1, using the method as described in steps S2-S4 to find a traffic sign with the same character semantic information as the sign image block at time t+1 in an image at time t+2; step S7, in the case of finding the same traffic sign, using a preset method to calculate the position and size of the traffic sign bounding box to obtain traffic sign images at times t and t+1; step S8, using a feature extractor in a recognition network to extract feature vectors of the traffic sign images at times t and t+1, and performing feature fusion on the feature vectors of the traffic sign images at times t and t+1 to obtain a fused feature vector, and using a classifier in the recognition network to classify and recognize the fused feature vector to obtain a final recognition result.

[0008] In some embodiments of the present application, in step S1, the following steps are included: S11, acquiring a video image, extracting a video image at time t, and using an image compression algorithm to pre-process the video image at time t to obtain a pre-processed image at time t; S12, using a trained detection network model to detect the pre-processed image at time t to obtain a sign image block at time t.

[0009] In some embodiments of the present application, the trained detection network model is trained in the following manner: a first training set is obtained, which includes a plurality of first training samples and a first label corresponding to each first training sample, the first training sample is a pre-processed video image, and the first label indicates the bounding box information labeled for the corresponding first training sample; the detection network model is trained one or more times using the first training set to obtain the trained detection network model.

[0010] In some embodiments of the present application, in step S2, the following steps are included: S21, inputting the t-time bus board image block into the trained bus board character detection model, performing feature extraction on the t-time bus board image block using a YOLO model to obtain a feature map of the t-time bus board image block; S22, fusing the feature map of the t-time bus board image block using an adaptive fusion module to obtain character information in the t-time bus board image block, the character information including character position information.

[0011] In some embodiments of the present application, the trained bus board character detection network model is trained in the following manner: a second training set is obtained, which includes a plurality of second training samples and a second label corresponding to each second training sample, the second training sample is a pre-detected bus board image block, and the second label indicates the character information labeled for the corresponding second training sample; the bus board character detection model is trained one or more times using the second training set to obtain the trained bus board character detection model.

[0012] In some embodiments of the present application, the adaptive fusion module is configured with a fusion mode for fusing the t-time bus board image feature map, and the fusion mode is as follows:

[0013]

[0014] wherein, represents a tensor product, x1, x2, x3 represent feature maps that need to detect bus boards of different sizes or dimensions, y1, y2, y3 represent fused feature maps, w a1 , w b1 , w c1 are fusion weight values of the feature map y1 fused by x1, x2, x3, w a2 , w b2 , w c2 are fusion weight values of the feature map y2 fused by x1, x2, x3, w a3 , w b3 , w c3 are fusion weight values of the feature map y3 fused by x1, x2, x3.

[0015] In some embodiments of the present application, the step S3 comprises: S31, dividing the character information in the t time point sign image block into a plurality of feature regions, constructing a data set with the plurality of feature regions, calculating the average value of the vertical direction of the first data in the data set, and selecting the data with the minimum average value as the first clustering center point; calculating the distance of the remaining data to the first distance center point, selecting the second data with the maximum distance to the first clustering center point as the second clustering center point, and storing the first clustering center point and the second clustering center point in the clustering center data set; S32, constructing a vector set corresponding to the first data based on the first data, calculating the sum of the outer product module of the vector set, and selecting the maximum value as the third clustering center point and storing it in the clustering center data set, so that the data in the clustering center data set is updated; S33, when the data in the clustering center data set reaches a first threshold value, the clustering center data set is divided by using the nearest neighbor algorithm; S34, an arbitrary data is selected from the plurality of feature regions divided from the character information as a clustering center point, and the bandwidth of the feature region to which the data belongs is calculated; S35, the mean deviation between the clustering center point in each feature region and other data is calculated based on a preset manner, when the mean deviation is greater than a second threshold value, the mean value of the data in the bandwidth of the feature region to which the current clustering center point belongs is calculated to obtain a new center point coordinate; S36, repeating steps S34 and S35 until the character information in the t time point sign image block is classified, obtaining the initial clustering center point, and obtaining the character corresponding to each initial clustering center point according to the query dictionary method, then connecting the characters in the corresponding order to form a correct character sequence to obtain the character semantic information of the t time point sign image block.

[0016] In some embodiments of the present application, the bandwidth of the feature region to which the data value belongs is calculated in the following manner:

[0017]

[0018] Wherein, z ij is all data in the feature region, c i is the clustering center point in the feature region, N i is the number of all data in the feature region, and η is the coefficient corresponding to the feature region.

[0019] In some embodiments of the present application, the mean deviation between the clustering center point data and other data in each region is calculated based on a preset manner, which is the mean deviation vector obtained by dividing the sum of the difference values of the clustering center point data and other data by the total number of other data.

[0020] In some embodiments of the present application, the identification network is trained in the following manner:

[0021] A third training set is acquired, which includes a plurality of third training samples and a third label corresponding to each third training sample, the third training sample is a pre-computed traffic sign image, and the third label indicates the recognition result of the corresponding third training sample; a feature extractor is used to extract features of the third training sample and fuse the extracted features to obtain fused features; a classifier is used to classify based on the fused features; a loss function is used to calculate a loss value based on the classification result and the sample label, and the parameters of the feature extractor and the classifier are updated according to the loss value.

[0022] In some embodiments of the present application, the preset method comprises an RGB component fixed difference value and an HIS hue threshold method.

[0023] According to a second aspect of the present application, a traffic sign target detection and recognition system is provided, which comprises a detection network model, a sign character detection model, a character calculation module, a feature mapping module and a recognition module, wherein: the detection network model is used for sign detection on a pre-processed video image to obtain a sign image block; the sign character detection model is used to obtain character information in the sign image block based on a YOLO target detection method with an added adaptive feature fusion module for fusing the feature map of the sign image block; the character calculation module is used to calculate an initial clustering center of the character information in the sign image block based on a character coordinate clustering algorithm, and sort the character information in the sign image block based on the initial clustering center to obtain character semantic information of the sign image block; the feature mapping module is used to find the same traffic sign in the t+1 time image as the character semantic information of the t time sign image block by using a feature mapping method; and the recognition module is used to process a final feature vector obtained by fusing the feature vectors of the t time traffic sign and the t+1 time traffic sign by using a recognition network to obtain a final recognition result.

[0024] Compared with the prior art, the present application has the following advantages:

[0025] The traffic sign target detection and recognition method and system provided by the present application can use a sign character detection model with an added adaptive feature fusion module to extract a sign image block, and also uses a clustering algorithm more suitable for character information recognition, thereby enhancing the classification precision, and further using fused feature vectors of two time points to enhance the stability and accuracy of the detection and recognition result, and improving the accuracy of sign detection and recognition in an automatic driving system under small targets and extreme weather conditions. BRIEF DESCRIPTION OF DRAWINGS

[0026] The embodiments of the present application will be further described below with reference to the accompanying drawings, in which:

[0027] Figure 1 A traffic sign target detection and recognition method flowchart according to an embodiment of the present application is shown.

[0028] Figure 2 An example schematic diagram of a traffic sign according to an embodiment of the present application.

[0029] Figure 3 An example schematic diagram of a signboard image block according to an embodiment of the present application.

[0030] Figure 4 An example schematic diagram of character information in a signboard image block according to an embodiment of the present application.

[0031] Figure 5 An example schematic diagram of a traffic signboard target detection and recognition method according to an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0033] As described in the background, in the prior art, most signboard detection and recognition methods have achieved good results in the case of close distance detection, but in the case of small size of road signs in the vehicle driving road environment or extremely blurred video shooting due to extreme weather, the detection and recognition accuracy of these algorithms is poor, resulting in low signboard detection and recognition rate. Moreover, due to the complex and variable traffic scene, the interference in the road is variable, such as the light intensity caused by weather, road congestion or obstruction, etc. These problems will also cause the automatic driving system to be unable to accurately detect the signboard information in front of the current vehicle.

[0034] In order to solve the above problems, the present application focuses on the detection and recognition of traffic signboards in the case of small targets and extreme weather, and proposes a traffic signboard target detection and recognition scheme that can improve the accuracy of signboard detection and recognition in the case of small targets and extreme weather in the automatic driving system. In summary, as Figure 1As shown, the present application provides a traffic sign target detection and recognition method, which comprises the following steps: S1, acquiring a video image, extracting a video image at time t for preprocessing to obtain a preprocessed image at time t, and using a trained detection network model to detect the preprocessed image at time t to obtain a sign image block at time t; S2, using a trained sign character detection model to extract character information in the sign image block at time t, wherein the trained sign character detection model is improved based on an original YOLO model, and an adaptive feature fusion module for fusing a sign image feature map at time t is added in the original YOLO model; S3, clustering the character information in the sign image block at time t to obtain an initial clustering center, and sorting the character information in the sign image block at time t according to semantics based on the initial clustering center to obtain character semantic information of the sign image block at time t; S4, using a feature mapping method to find a traffic sign with the same character semantic information as the sign image block at time t in an image at time t+1; S5, in the case of finding the same traffic sign, using the trained detection network model to detect a preprocessed image at time t+1 to obtain a sign image block at time t+1; S6, based on the obtained sign image block at time t+1, using the method as described in steps S2-S4 to find a traffic sign with the same character semantic information as the sign image block at time t+1 in an image at time t+2; S7, in the case of finding the same traffic sign, using a preset method to calculate the position and size of the traffic sign bounding box to obtain traffic sign images at times t and t+1; S8, using a feature extractor in a recognition network to extract feature vectors of the traffic sign images at times t and t+1, and performing feature fusion on the feature vectors of the traffic sign images at times t and t+1 to obtain a fused feature vector, and using a classifier in the recognition network to classify and recognize the fused feature vector to obtain a final recognition result. In the traffic sign target detection and recognition scheme of the present application, the sign character detection model with the added adaptive feature fusion module is used to extract the sign image block, and a clustering algorithm more suitable for character information recognition is also used, which enhances the classification precision. Moreover, the fused feature vectors of two times are used, which enhances the stability and accuracy of the detection and recognition result, and improves the accuracy of sign detection and recognition in an automatic driving system under small targets and extreme weather. In order to better understand the present application, the scheme of the present application will be described in detail below in combination with specific embodiments and examples.

[0035] In the step S1, a video image is acquired, a video image at time t is extracted for preprocessing to obtain a preprocessed image at time t, and a trained detection network model is used to detect the preprocessed image at time t to obtain a sign image block at time t.

[0036] According to an embodiment of the present application, the step s1 comprises:

[0037] S11, acquiring a video image, extracting a video image at time t, and preprocessing the video image at time t using an image compression algorithm to obtain a preprocessed image at time t;

[0038] S12, detecting a road sign image block at time t using the trained detection network model.

[0039] According to an embodiment of the present application, in the step S11, the video image in the high-speed camera is acquired, and the video image at time t is extracted. The extracted video image at time t can be scaled to a fixed size using a discrete cosine transformation (DCT) lossless image compression algorithm. In an example, the fixed size is 416x416 pixels and the number of channels is 3.

[0040] According to an embodiment of the present application, the detection network model used can be a YOLOv3 model (or an RCNN model). The training process is as follows: a first training set is obtained, which includes a plurality of first training samples and a first label corresponding to each first training sample. The first training sample is a preprocessed video image, and the first label indicates the bounding box information labeled for the corresponding first training sample. The detection network model is trained one or more times using the first training set to obtain the trained detection network model. According to an example of the present application, the training process can also be as follows: a plurality of data sets are obtained, each of which includes a scaled video image as a training sample and a bounding box information labeled for the scaled video image as a label. The YOLOv3 model is trained using the data set.

[0041] According to an embodiment of the present application, when labeling the scaled video image, image labeling is used to provide the bounding box information of the video image.

[0042] According to an embodiment of the present application, in step S12, the trained detection network model is used to detect the scaled video image at time t to obtain a road sign image block at time t. It should be noted that traffic signs are various, as shown in the following table: Figure 2 For example: 1. Speed limit sign: including different speed limit signs such as 30 km / h, 60 km / h, 100 km / h, etc. 2. Stop sign: including red octagonal stop signs to stop traffic. 3. Turn sign: including different types of turn signs such as left turn, right turn and U-turn. 4. Warning sign: including signs to warn drivers to pay attention, such as attention to children, attention to animals, etc. 5. Indication sign: including signs indicating a specific direction or location, such as road name, exit sign, etc. Therefore, the road sign image block at time t obtained by using the trained detection network model is also different.

[0043] According to an example of the present invention, Figure 3 As shown, the trained detection network model is used to detect the scaled video image at time t (such as the turn sign, main road sign, speed limit sign, and stop sign in the figure), and the road sign image block at time t is a triangular (or diamond, circular, or octagonal) road sign image block.

[0044] In step S2, a road sign character detection model improved based on the original YOLO model is used to extract character information in the road sign image block at time t.

[0045] According to one embodiment of the present invention, step S2 includes:

[0046] S21, inputting the road sign image block at time t into the trained road sign character detection model, and using the YOLO model to perform feature extraction on the road sign image block at time t to obtain a feature map of the road sign image block at time t;

[0047] S22 , using an adaptive fusion module to fuse the feature map of the road sign image block at time t to obtain character information in the road sign image block at time t, wherein the character information includes character position information.

[0048] According to one embodiment of the present invention, the trained road sign character detection model is improved based on the original YOLO model. The improvement principle is: an adaptive feature fusion module for fusing the feature map of the road sign image at time t is added to the original YOLO model.

[0049] According to one embodiment of the present invention, the trained road sign character detection model can be trained in the following manner: obtaining a second training set, which includes multiple second training samples and each second training sample corresponds to a second label, the second training sample is a pre-detected road sign image block, and the second label indicates the character information marked corresponding to the second training sample; using the second training set to train the road sign character detection model once or multiple times to obtain a trained road sign character detection model.

[0050] According to one embodiment of the present invention, in step S22, the adaptive feature fusion module is configured with a fusion method for fusing the feature map of the road sign image at time t, and the obtained multiple feature maps are fused to obtain the character information of the road sign image block. The fusion method is specifically:

[0051]

[0052] in, represents a tensor product, x1, x2, x3 represent feature maps of different sizes or sizes of signs to be detected, y1, y2, y3 represent feature maps obtained by fusion, w a1 , w b1 , w c1 are fusion weight values of feature map y1 obtained by fusion of x1, x2, x3, w a2 , w b2 , w c2 are fusion weight values of feature map y2 obtained by fusion of x1, x2, x3, w a3 , w b3 , w c3 are fusion weight values of feature map y3 obtained by fusion of x1, x2, x3. According to an example of the present application, the fusion weight values can be learned to optimal values by back propagation or genetic algorithm.

[0053] According to an example of the present application, as shown in Figure 4 , the character information in the sign image block at time t obtained by using the trained sign character detection model is "4" and "0", the character information includes character position information, and the character position information is used for subsequent clustering operation.

[0054] In the step S3, the character information in the sign image block at time t is clustered to obtain an initial clustering center, and the character information in the sign image block at time t is sorted according to semantics based on the initial clustering center to obtain character semantic information of the sign image block at time t.

[0055] According to an example of the present application, the step S3 includes:

[0056] S31, the character information in the sign image block at time t is divided into a plurality of feature regions, a data set is constructed with the plurality of feature regions, the average value of the vertical direction of the first data in the data set is calculated, and the data with the smallest average value is selected as the first clustering center point; the distance of the remaining data to the first distance center point is calculated, the second data with the largest distance from the first clustering center point is selected as the second clustering center point, and the first clustering center point and the second clustering center point are stored in the clustering center data set;

[0057] S32, a vector set corresponding to the first data is constructed based on the first data, the sum of the outer product module of the vector set is calculated, and the maximum value is selected as the third clustering center point and stored in the clustering center data set, so that the data in the clustering center data set is updated;

[0058] S33, when the data in the clustering center data set reaches a first threshold, the clustering center data set is divided by using the nearest neighbor algorithm;

[0059] S34, randomly selecting a data point from the multiple feature regions divided by the character information as a cluster center point, and calculating the bandwidth of the feature region to which the data point belongs;

[0060] S35. Calculate the mean offset between the cluster center point and other data in each feature region based on a preset method. When the mean offset is greater than a second threshold, obtain the new center point coordinates by calculating the mean value of the data within the bandwidth of the feature region to which the current cluster center point belongs.

[0061] S36. Repeat steps S34 and S35 until all the character information in the road sign image block at time t is classified, obtain the initial cluster center points, and obtain the characters corresponding to each initial cluster center point based on the query dictionary method. Then, connect the characters into a correct character sequence in the corresponding order to obtain the character semantic information of the road sign image block at time t.

[0062] According to an embodiment of the present invention, in step S31, the character information in the road sign image block at time t is divided into n feature regions, and a data set Z={z1, z2, z3, ..., z n}, calculate the first data z in data set Z i ={z i1 , z i2 , z i3 ,…,z im}, select the data with the smallest average value as the first cluster center c1, calculate the distance of the remaining data to the first cluster center c1, select the second data with the largest distance from the first cluster center c1 as the second cluster center c2, and store c1 and c2 in the cluster center data set H. According to an example of the present invention, the center point in the cluster center data set is the center position obtained by the average value of the coordinates of all data points.

[0063] According to one embodiment of the present invention, in step S32, based on the first data z i , construct the corresponding vector set Z i , calculate the vector set Z i The sum of the outer product moduli of , select the maximum value as the third cluster center c3, and store the third cluster center c3 in the cluster center dataset H to update the data in the cluster center dataset H.

[0064] According to one embodiment of the present application, in the step S33, when the number of data in the cluster center data set H reaches a set threshold, the cluster center data set H is divided into m classes by using the nearest neighbor algorithm. If the number of data in the cluster center data set H does not reach the set threshold, the step S32 is returned to, and the data in the cluster center data set H is continuously updated until the data in the cluster center data set reaches the set threshold.

[0065] According to one embodiment of the present application, in the step S34, one data is selected as a cluster center point from any one of the n feature regions divided from the character information, and the bandwidth of the feature region to which the data belongs is calculated, the bandwidth is used to represent the width of the region in which the character is recognized, that is, the width (size / range) of the region to which the cluster center point belongs, and the bandwidth of the region to which the data value belongs is calculated in the following way:

[0066]

[0067] wherein z ij is all the data in the feature region, c i is the cluster center point in the feature region, N i is the number of all the data in the feature region, and η is the coefficient corresponding to the feature region.

[0068] According to one embodiment of the present application, in the step S35, the mean shift vector between the cluster center point and other data in each feature region is calculated by dividing the sum of the difference between the cluster center point data and other data by the total number of other data. The cluster center point data is moved in the direction of the probability density gradient along the mean shift vector, and when the distance of the moved mean shift vector is greater than a set threshold, a new cluster center point is obtained by calculating the mean value of the data in the bandwidth of the feature region to which the current cluster center point belongs; if the distance of the moved mean shift vector is not greater than the set threshold, the mean shift vector is recalculated. The probability density gradient can be calculated by the meanshif algorithm. According to one example of the present application, the set threshold is any constant, which can be optimized, and the set threshold is the termination condition of the iteration of the clustering algorithm for the character information in the sign image block at time t.

[0069] According to one embodiment of the present application, in the step S36, the step S34 and the step S35 are repeated, the class with the largest frequency is selected as the class of the current data based on the access frequency of each data, and until the character information in the sign image block at time t is classified, m data are obtained as the initial cluster center points H={c1, c2, …, cm}. m} and according to the query dictionary method, the characters corresponding to each initial clustering center point are obtained, and then the characters are connected in the corresponding order to obtain the correct character sequence, thereby obtaining the character semantic information of the bus stop image block at time t. According to an example of the present application, still taking the character information "4" and "0" shown in Figure 4 Figure 4 the figure as an example, the semantic order after clustering can be "40" and "04", and according to the common sense of the traffic sign, "04" is in reverse order, so the character semantic information of the bus stop image block at time t is "40". It should be noted that the access frequency of each data can be obtained through the iteration process in step S35.

[0070] It should be noted that the above method of clustering the character information in the bus stop image block at time t is a character coordinate clustering algorithm designed by the present application. This algorithm improves the traditional k-means algorithm by adding the maximum value as the initial clustering center point and using the offset to optimize the clustering center iteration, so that the initial clustering center is more accurate, the convergence speed is faster, the iteration times are reduced, and the classification accuracy is improved.

[0071] In the step S4, the SIFT feature mapping method is used to find the same traffic sign in the image at time t+1 as the character semantic information of the bus stop image block at time t.

[0072] In the step S5, in the case of finding the same traffic sign, the trained detection network model is used to detect the bus stop in the preprocessed image at time t+1, thereby obtaining the bus stop image block at time t+1.

[0073] According to an example of the present application, when the same traffic sign as the character semantic information of the bus stop image block at time t is found in the image at time t+1, it indicates that the character semantic information of the bus stop image block at time t is accurate.

[0074] According to an example of the present application, if the same traffic sign as the character semantic information of the bus stop image block at time t is not found in the image at time t+1, then the detection is restarted from step S1. Because in the case of not finding the same traffic sign, the bus stop information at time t has the possibility of being unreliable, the detection result at time t needs to be abandoned and the detection is restarted.

[0075] In the step S6, based on the obtained bus stop image block at time t+1, the method described in steps S2-S4 is used to find the same traffic sign in the image at time t+2 as the character semantic information of the bus stop image block at time t+1.

[0076] ​In this step, the method for finding the same traffic sign as the character semantic information of the sign image block at t+1 time in the image at t+2 time is the method in steps S2-S4 as described above.

[0077] In the step S7, in the case of finding the same traffic sign, the position and size of the traffic sign bounding box are calculated by using a preset method to obtain the traffic sign images at t time and t+1 time.

[0078] According to one embodiment of the present application, since the meaning represented by the shape, color, number, letter and other information of the traffic sign is different, in the case of finding the same traffic sign as the character semantic information of the sign image block at t+1 time in the image at t+2 time, the position and size of the traffic sign bounding box can be calculated by using the RGB component fixed difference value and HIS hue threshold method to obtain the traffic sign images at t time and t+1 time.

[0079] According to one embodiment of the present application, in the case of not finding the same traffic sign as the character semantic information of the sign image block at t+1 time in the image at t+2 time, it is necessary to start from step S1 to re-detect.

[0080] In the step S8, the feature vectors of the traffic sign images at t time and t+1 time are extracted by using the feature extractor in the recognition network, the feature vectors of the traffic sign images at t time and t+1 time are fused to obtain a fused feature vector, and the fused feature vector is classified and recognized by using the classifier in the recognition network to obtain the final recognition result.

[0081] According to one embodiment of the present application, before the feature extraction and classification recognition of the traffic sign images at t time and t+1 time, the recognition network needs to be trained in the following manner: a third training set is obtained, which includes a plurality of third training samples and a third label corresponding to each third training sample, the third training sample is a pre-calculated traffic sign image, and the third label indicates the recognition result of the corresponding third training sample; the feature extractor is used to extract the features of the third training sample and fuse the extracted features to obtain a fused feature; the classifier is used to classify based on the fused feature; based on the classification result and the sample label, the loss value is calculated by using the loss function, and the parameters of the feature extractor and the classifier are updated according to the loss value.

[0082] According to one embodiment of the present application, the recognition network can use a VGG network (more specifically, a VGG19 network), extract feature vectors of traffic sign images at time t and time t+1 using the VGG network, perform feature fusion on the feature vectors of the traffic sign images at time t and time t+1 to obtain a fused feature vector, and perform classification recognition on the fused feature vector using a classifier in the VGG network to obtain a final recognition result. The use of feature vectors at two times for final recognition enhances the stability and accuracy of the detection and recognition result.

[0083] It should be noted that time t, time t+1 and time t+2 can be three consecutive times, or three times with the same time interval (for example, a time interval of 0.5s), or three times with different time intervals (for example, a time interval of 0.1s between time t and time t+1, and a time interval of 0.2s between time t+1 and time t+2). Of course, time t can be the current time, time t+1 can be the next time, and time t+2 can be the time after the next time.

[0084] In order to better illustrate the implementation process of the traffic sign target detection and recognition method of the present application, the following will be specifically described in conjunction with an example. Figure 5 The entire completion process includes the following steps:

[0085] Step T1, import video data of a vehicle-mounted camera;

[0086] Step T2, process video image data;

[0087] Step T3, make and label road signs and backgrounds, complete video image road sign detection data set preparation and detection network model training, and apply the detection network to detect video images at time k;

[0088] Step T4, video image road sign character detection: prepare a video image road sign character detection data set, complete video image road sign character detection network model training; input the road sign detection result into the trained video road sign character detection network to obtain character information of a road sign image block at time k;

[0089] Step T5, detect and recognize character information in the road sign image block at time k based on a character coordinate clustering algorithm;

[0090] Step T6, use SIFT feature mapping to find the same traffic sign in the image at time k+1 as the detection and recognition result at time k;

[0091] Step T7, determine whether the same traffic sign is found; if not, return to step T3; if so, perform step T8;

[0092] Step T8, the video image at the k+1 moment is detected by applying a detection network;

[0093] Step T9, video image signboard character detection, the signboard detection result is input into a trained video signboard character detection network to obtain character information of the signboard image block at the k+1 moment;

[0094] Step T10, the character information in the signboard image block at the k+1 moment is detected and recognized based on a character coordinate clustering algorithm;

[0095] Step T11, the same traffic sign as the detection and recognition result at the k+1 moment is found in the image at the k+2 moment by using SIFT feature mapping;

[0096] Step T12, whether the same traffic sign is found is judged; if not, the step T3 is returned; if yes, the step T13 is executed;

[0097] Step T13, in the case that the same traffic sign is found, the position and size of the traffic sign boundary box are calculated to obtain the traffic sign image at the k and k+1 moments;

[0098] Step T14, the feature vectors of the traffic sign images at the k and k+1 moments are extracted by using the feature extractor in the recognition network;

[0099] Step T15, the extracted feature vectors are fused to obtain a fused feature vector;

[0100] Step T16, the fused feature vector is classified and recognized by using the classifier of the recognition network to obtain a final recognition result;

[0101] Step T17, whether the image is the last one is judged; if not, the step T3 is returned; if yes, the process is ended.

[0102] In order to better understand the implementation process of the traffic signboard target detection and recognition method of the application, a traffic signboard target detection and recognition system is used for illustration.

[0103] The traffic signboard target detection and recognition system for realizing the traffic signboard target detection and recognition scheme of the application includes a detection network model, a signboard character detection model, a character calculation module, a feature mapping module and a recognition module.

[0104] The detection network model is used for signboard detection on the preprocessed video image to obtain a signboard image block;

[0105] The signboard character detection model is used for obtaining character information in the signboard image block based on the YOLO target detection mode with an added self-adaptive feature fusion module for fusing the feature map of the signboard image block;

[0106] The character calculation module is configured to calculate an initial clustering center of character information in the sign image block based on a character coordinate clustering algorithm, and sort the character information in the sign image block based on the initial clustering center to obtain character semantic information of the sign image block.

[0107] The feature mapping module is configured to find the same traffic sign as the character semantic information of the sign image block at the t time in the image at the t+1 time by using a feature mapping method.

[0108] The recognition module is configured to process a final feature vector obtained by fusing feature vectors of the traffic sign at the t time and the traffic sign at the t+1 time by using a recognition network to obtain a final recognition result.

[0109] To sum up, the traffic sign target detection and recognition method and system provided by the application can use the sign character detection model with the adaptive feature fusion module to extract the sign image block, and also uses the clustering algorithm more suitable for character information recognition, thereby enhancing the classification precision, and the fusion feature vectors of two times are also used, thereby enhancing the stability and accuracy of the detection and recognition result, and improving the accuracy of the sign detection and recognition in the automatic driving system under small targets and extreme weather.

[0110] It should be noted that although the above describes the steps in a specific order, it does not mean that the steps must be performed in the above specific order, in fact, some of the steps can be executed concurrently, or even changed in order, as long as the required function can be realized.

[0111] The application can be a system, a method and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for causing a processor to implement various aspects of the application.

[0112] The computer readable storage medium can be a tangible device that maintains and stores instructions for use by an instruction execution device. The computer readable storage medium can include, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch cards or punched tape, and any suitable combination of the foregoing. A non-transitory, computer-readable storage medium does not include a signal.

[0113] Having described various embodiments of the application, it is to be understood that the above description is meant not to limit and not to encompass all of the possible embodiments covered by the claims. Many modifications and variations of this application can be apparent to those of ordinary skill in the art without departing from the spirit and scope of the described embodiments. It is intended that the scope of the application should only be limited by the appended claims.

Claims

1. A method for detecting and recognizing traffic signs, characterized in that: The method comprises the following steps: Step S1: Acquire a video image, extract the video image at time t, perform preprocessing on the video image, and obtain a preprocessed image at time t; and use the trained detection network model to perform road sign detection on the preprocessed image at time t to obtain a road sign image block at time t; Step S2: extracting character information from the road sign image block at time t using a trained road sign character detection model, wherein the trained road sign character detection model is improved based on the original YOLO model, and an adaptive feature fusion module is added to the original YOLO model for fusing the feature map of the road sign image at time t; Step S3: clustering the character information in the road sign image block at time t to obtain initial cluster centers, and sorting the character information in the road sign image block at time t according to semantics based on the initial cluster centers to obtain character semantic information of the road sign image block at time t; Step S4: using a feature mapping method to find a traffic sign with the same character semantic information as the road sign image block at time t in the image at time t+1; Step S5: When the same traffic sign is found, the trained detection network model is used to perform road sign detection on the pre-processed image at time t+1 to obtain a road sign image block at time t+1; Step S6: Based on the obtained road sign image block at time t+1, use the method described in steps S2-S4 to find a traffic sign with the same character semantic information as the road sign image block at time t+1 in the image at time t+2; Step S7: When the same traffic sign is found, the position and size of the traffic sign bounding box are calculated using a preset method to obtain traffic sign images at time t and time t+1; Step S8: Use the feature extractor in the recognition network to extract the feature vectors of the traffic sign images at time t and time t+1, and perform feature fusion on the feature vectors of the traffic sign images at time t and time t+1 to obtain a fused feature vector. Use the classifier in the recognition network to classify and identify the fused feature vector to obtain the final recognition result.

2. The method according to claim 1, wherein In step S1, the following steps are included: S11, acquiring a video image, extracting the video image at time t, and preprocessing the video image at time t using an image compression algorithm to obtain a preprocessed image at time t; S12. Use the trained detection network model to perform road sign detection on the pre-processed image at time t to obtain a road sign image block at time t.

3. The method according to claim 2, characterized in that The trained detection network model is trained as follows: Obtaining a first training set, which includes a plurality of first training samples and a first label corresponding to each first training sample, wherein the first training sample is a pre-processed video image, and the first label indicates bounding box information marked corresponding to the first training sample; The detection network model is trained once or multiple times using the first training set to obtain a trained detection network model.

4. The method according to claim 1, wherein In step S2, the following steps are included: S21, inputting the road sign image block at time t into the trained road sign character detection model, and using the YOLO model to perform feature extraction on the road sign image block at time t to obtain a feature map of the road sign image block at time t; S22 , using an adaptive fusion module to fuse the feature map of the road sign image block at time t to obtain character information in the road sign image block at time t, wherein the character information includes character position information.

5. The method according to claim 4, characterized in that The trained street sign character detection network model is trained as follows: Obtaining a second training set, which includes a plurality of second training samples and a second label corresponding to each second training sample, wherein the second training sample is a pre-detected road sign image block, and the second label indicates character information labeled corresponding to the second training sample; The road sign character detection model is trained once or multiple times using the second training set to obtain a trained road sign character detection model.

6. The method according to claim 4, characterized in that The adaptive fusion module is configured with a fusion method for fusing the feature map of the road sign image at time t, and the fusion method is as follows: in, It represents the tensor product, x1, x2, x3 represent the feature maps of road signs of different sizes or dimensions, y1, y2, y3 represent the fused feature maps, and w a1 、w b1 、w c1 The fusion weight value of feature map y1 is obtained by fusing x1, x2, and x3 respectively, w a2 、w b2 、w c2 The fusion weight value of feature map y2 is obtained by fusing x1, x2, and x3 respectively, w a3 、w b3 、w c3 Fuse x1, x2, and x3 respectively to obtain the fusion weight value of the feature map y3.

7. The method according to claim 1, wherein: The step S3 includes: S31, dividing the character information in the road sign image block at time t into multiple feature regions, constructing a data set based on the multiple feature regions, calculating the vertical average value of the first data in the data set, and selecting the data with the smallest average value as the first cluster center point; calculating the distance from the remaining data to the first distance center point, selecting the second data with the largest distance from the first cluster center point as the second cluster center point, and storing the first cluster center point and the second cluster center point in the cluster center data set; S32: constructing a vector set corresponding to the first data based on the first data, calculating the sum of the outer product moduli of the vector set, selecting the maximum value as the third cluster center point and storing it in the cluster center data set, so that the data in the cluster center data set is updated; S33, when the data in the cluster center data set reaches a first threshold, using a nearest neighbor algorithm to divide the cluster center data set; S34, randomly selecting a data point from the multiple feature regions divided by the character information as a cluster center point, and calculating the bandwidth of the feature region to which the data point belongs; S35. Calculate the mean offset between the cluster center point and other data in each feature region based on a preset method. When the mean offset is greater than a second threshold, obtain the new center point coordinates by calculating the mean value of the data within the bandwidth of the feature region to which the current cluster center point belongs. S36. Repeat steps S34 and S35 until all the character information in the road sign image block at time t is classified, obtain the initial cluster center points, and obtain the characters corresponding to each initial cluster center point based on the query dictionary method. Then, connect the characters into a correct character sequence in the corresponding order to obtain the character semantic information of the road sign image block at time t.

8. The method according to claim 7, characterized in that The bandwidth of the characteristic region to which the data value belongs is calculated in the following manner: Among them, z ij is all the data in the feature area, c i is the cluster center point in the feature area, N i is the number of all data in the feature area, and η is the coefficient corresponding to the feature area.

9. The method according to claim 7, characterized in that The method of calculating the mean offset between the cluster center point data and other data in each feature area based on a preset method is to obtain a mean offset vector by dividing the sum of the differences between the cluster center point data and the other data by the total number of the other data.

10. The method according to claim 1, wherein The recognition network is trained as follows: Obtaining a third training set, which includes a plurality of third training samples and a third label corresponding to each third training sample, wherein the third training sample is a pre-calculated traffic sign image, and the third label indicates a recognition result corresponding to the third training sample; Extracting features from the third training sample using a feature extractor and fusing the extracted features to obtain fused features; Classification is performed using a classifier based on fusion features; Based on the classification results and sample labels, the loss function is used to calculate the loss value, and the parameters of the feature extractor and classifier are updated according to the loss value.

11. The method according to claim 1, wherein The preset methods include: RGB component fixed difference and HIS hue threshold method.

12. A traffic sign target detection and recognition system, characterized in that: The system includes: a detection network model, a road sign character detection model, a character calculation module, a feature mapping module and a recognition module, wherein: The detection network model is used to perform road sign detection on the preprocessed video image to obtain a road sign image block; The road sign character detection model is used to obtain character information in the road sign image block based on the YOLO target detection method with an additional adaptive feature fusion module for fusing feature maps of the road sign image block; The character calculation module is used to calculate the initial cluster center of the character information in the road sign image block based on the character coordinate clustering algorithm, and sort the character information in the road sign image block based on the initial cluster center to obtain the character semantic information of the road sign image block; The feature mapping module is used to find the traffic sign with the same character semantic information as the road sign image block at time t in the image at time t+1 by using the feature mapping method; The recognition module is used to use the recognition network to process the final feature vector obtained by fusing the feature vectors of the traffic sign at time t and the traffic sign at time t+1 to obtain the final recognition result.

13. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of any one of the methods of claims 1 to 11.

14. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Viewpoint tracking method based on geometrical reconstruction and semantic integration

    CN104408158A

  • Method and device for extracting key scene in video

    CN109525892A