A road traffic sign recognition and classification method and device
By adding the OCR text detection and recognition algorithm after the ByteTrack target tracking algorithm, the identification and classification of road traffic signs is solved, and the problems of low detection efficiency of small and medium-sized targets and occlusion targets in the prior art are solved, and efficient target deduplication and detection effects are achieved.
Patent Information
- Application Number
- CN202510173758.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
When detecting road traffic signs, the prior art is prone to lose small targets and occlusion targets, and the tracking algorithm will leave a large number of repeated targets when deduplication, resulting in low detection efficiency and frequent repeated targets.
After the target tracking algorithm ByteTrack, the target area text detection and recognition algorithm is added, and the secondary deduplication is performed through OCR text detection and recognition, reducing the number of repeated targets in the detection results, and improving the target detection efficiency.
It effectively reduces the number of repeated targets in the detection results, improves the efficiency of target detection, and reduces the workload of manual review of data in the later stage.
Smart Images

Figure CN119649341B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence model detection, and in particular to a method and device for identifying and classifying road traffic signs. Background Art
[0002] As the number of cars increases, road traffic becomes more and more congested, and traffic accidents become more and more frequent. Autonomous driving is considered to be a way to optimize this situation. Traffic signs, as an important part of road safety facilities, can effectively regulate traffic behavior and play an important role in indicating road conditions, guiding pedestrians and driving regulations. And as more and more highway infrastructure is built, the regular maintenance and asset statistics of road traffic signs are particularly important.
[0003] Among the existing technologies, the recognition of road traffic signs mainly uses a detection and recognition model to annotate pictures containing targets, and then trains a deep neural network model to obtain the target position coordinates of the signs in the pictures. The targets are then deduplicated through a tracking algorithm, and finally the target area is classified into disease categories through a classification network to obtain a single, non-repetitive target disease category.
[0004] Although current deep neural networks can detect targets from images very well, they are prone to losing targets when they are small or obscured, and the tracking algorithm will still leave a large number of duplicate targets when deduplicating the target area. Therefore, how to improve the efficiency of target detection and effectively reduce the frequency of duplicate targets is a series of technical problems that need to be solved urgently. Summary of the invention
[0005] The purpose of the present invention is to provide a method and device for identifying and classifying road traffic signs, which improves the deduplication rate of detected targets by the network structure by adding a target area text detection and recognition algorithm behind the target tracking algorithm ByteTrack, effectively reduces the number of repeated targets in the detection results, improves the detection efficiency of the target and effectively reduces the frequency of repeated targets.
[0006] To achieve the above objectives, the specific scheme of this application is:
[0007] In one aspect, the present invention provides a method for identifying and classifying road traffic signs, which specifically comprises the following steps:
[0008] S1. Obtain a video stream of road traffic signs, sample the video stream, obtain a number of initial images and store them in a first image set in sequence;
[0009] S2. Build a YOLOv8 target detection optimization model, input the first image set into the YOLOv8 target detection optimization model, and obtain a number of images to be identified, each of which is marked with the coordinate position and label information of each traffic sign;
[0010] S3, according to the coordinate position and label information of each traffic sign, use the target tracking network ByteTrack to track each traffic sign as a target to obtain the target ID of each traffic sign;
[0011] S4, performing a first deduplication process on the image to be identified according to the target ID of each traffic sign in each image to be identified, and storing the image to be identified after the first deduplication process in a second data set;
[0012] S5, extracting the target area image where each traffic sign is located from the image to be recognized after the first deduplication process, performing text recognition on the target area image, and obtaining a text recognition result for each traffic sign;
[0013] S6. Perform a second deduplication process on the second data set according to the target ID of each traffic sign and its text recognition result, and save the image to be recognized after the second deduplication process into a third data set;
[0014] S7. Input the images in the third data set into the YOLOv8 classification network for disease recognition to obtain the detection result of each traffic sign.
[0015] In some specific implementation schemes, the specific process of the first deduplication process in step S4 is:
[0016] S41, comparing the target ID of each traffic sign in each image to be identified, if one of the images to be identified includes the target IDs of all traffic signs in another image to be identified, deleting the other image to be identified;
[0017] S42: If the target ID of some traffic signs in one of the images to be identified is the same as that in the other image to be identified, then the annotations of the traffic signs with the same target ID are deleted from the other image to be identified, where the annotations include the coordinate positions and label information of the traffic signs.
[0018] In some specific implementation schemes, the text recognition method uses OCR text detection and recognition. For each target area image where a traffic sign is located, the specific recognition process is as follows:
[0019] S51, preprocessing the target area image to obtain a preprocessed target area image;
[0020] S52, using the CRAFT algorithm to detect the position of the text area on the preprocessed target area image to obtain a polygonal text area image for text detection;
[0021] S53, inputting the polygonal text area image into the ResNet+LSTM+CTC network structure to recognize the text content, and obtaining the text content of the text area;
[0022] S54. Use a greedy decoder to perform text meaning analysis on the text content, output the word with the highest probability, and obtain the text recognition result of each traffic sign.
[0023] In some specific implementation schemes, the second deduplication process in step S6 is as follows:
[0024] S61, obtaining a plurality of consecutive images to be identified in the second data set, comparing the target ID of each traffic sign in each image to be identified and its text recognition result, and if the text recognition results contained in each image to be identified are the same, retaining only the first image to be identified;
[0025] S62. When some traffic signs in the images to be identified have different target IDs and text recognition results, and some traffic signs have different target IDs but the same text recognition results, the markings of the traffic signs with different target IDs but the same text recognition results in the images to be identified are deleted.
[0026] In some specific implementation schemes, when the image after the second deduplication process is saved in a third data set, the following steps are also included:
[0027] According to the coordinate position and label information of each traffic sign, the recognition area image of each traffic sign is cut out from the image to be recognized after the second deduplication process, and the cut recognition area image is saved in the third data set.
[0028] In some specific implementation schemes, the backbone network of the YOLOv8 target detection optimization model includes a CSPDarknet structure, a C2f structure, and two C3STR structures.
[0029] In some specific implementation schemes, each C3STR structure includes a PatchPartition module for performing a block operation on an image and four cascaded Stage modules for converting the block-operated image into feature maps of different sizes, and each Stage module includes a Swin TransformerBlock network.
[0030] In some specific implementation schemes, the four Stage modules are Stage_1, Stage_2, Stage_3 and Stage4, wherein Stage_1 includes a cascaded Linear Embedding layer and a Swin Transformer Block network, and Stage_2, Stage_3 and Stage4 each include a Patch Merging layer for downsampling and a Swin Transformer Block network.
[0031] In some specific implementations, each Swin Transformer Block network includes a W-MSA structure and a SW-MSA structure used in pairs.
[0032] In a second aspect, the present application provides a road traffic sign recognition and classification device, comprising:
[0033] one or more processors;
[0034] A storage unit is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors can implement the road traffic sign recognition and classification method described in the first aspect.
[0035] The present invention has the beneficial effects:
[0036] 1. During the target detection process, the YOLOv8 target detection network is optimized, two C2f structures in its backbone network are replaced with C3STR structures, and the offset fusion of windows is used to solve the problem that windows cannot communicate with each other, thereby achieving better detection effects for occluded targets and small targets;
[0037] 2. After the images with traffic signs are subjected to preliminary target detection through the improved YOLOv8 target detection network, the location and type of the traffic signs can be marked in the image. In order to reduce the amount of image data for the subsequent recognition network, the target tracking network ByteTrack is used to achieve deduplication for images without text content; for images with text content, the target tracking network ByteTrack is first used for the first deduplication process, and then combined with OCR recognition for the second target deduplication, achieving a higher target detection rate, greatly improving the deduplication effect, effectively reducing the number of duplicate targets in the detection results, and greatly reducing the workload of manual data review in the later stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A flow chart of a method for identifying and classifying road traffic signs provided by an embodiment of the present invention;
[0039] Figure 2 A YOLOv8 model structure diagram provided by an embodiment of the present invention;
[0040] Figure 3 A schematic diagram of the C3STR network structure provided in an embodiment of the present invention;
[0041] Figure 4 This is a network structure diagram of Swin_Transformer_Blocks provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] The relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.
[0044] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0045] Additionally, descriptions of well-known structures, functions, and configurations may be omitted for clarity and conciseness.One of ordinary skill in the art will recognize that various changes and modifications may be made to the examples described herein without departing from the spirit and scope of the present disclosure.
[0046] Technologies, methods, and apparatus known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, such technologies, methods, and apparatus should be considered part of the authorization specification.
[0047] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0048] Example 1
[0049] like Figure 1 As shown, this embodiment provides a method for identifying and classifying road traffic signs, which specifically includes the following steps:
[0050] S1. Obtain a video stream of road traffic signs, sample the video stream, obtain a number of initial images and store them in a first image set in sequence;
[0051] In this embodiment, when sampling the video stream, since the information contained in adjacent frames is almost unchanged, a frame skipping saving method can be adopted. For example, every ten frames of the image are saved once, and the images directly captured from the video stream are stored in the first image set in sequence according to the acquisition order.
[0052] S2. Build a YOLOv8 target detection optimization model, input the first image set into the YOLOv8 target detection optimization model, and obtain a number of images to be identified, each of which is marked with the coordinate position and label information of each traffic sign;
[0053] Specifically, in order to improve the detection rate of the network model for small targets and occluded targets and realize the detection and recognition of small target detection tasks, the YOLOv8 target detection model is optimized in this embodiment, and two of the last three C2f structures in the backbone network are replaced with two C3STR structures, and the data dimension output by the C3STR structure is ensured to be consistent.
[0054] S3, according to the coordinate position and label information of each traffic sign, use the target tracking network ByteTrack to track each traffic sign as a target to obtain the target ID of each traffic sign;
[0055] S4, performing a first deduplication process on the image to be identified according to the target ID of each traffic sign in each image to be identified, and storing the image to be identified after the first deduplication process in a second data set;
[0056] Specifically, the specific process of the first deduplication process is as follows:
[0057] S41, comparing the target ID of each traffic sign in each image to be identified, if one of the images to be identified includes the target IDs of all traffic signs in another image to be identified, deleting the other image to be identified;
[0058] S42: If the target ID of some traffic signs in one of the images to be identified is the same as that in the other image to be identified, then the annotations of the traffic signs with the same target ID are deleted from the other image to be identified, where the annotations include the coordinate positions and label information of the traffic signs.
[0059] S5, extracting the target area image where each traffic sign is located from the image to be recognized after the first deduplication process, performing text recognition on the target area image, and obtaining a text recognition result for each traffic sign;
[0060] In this embodiment, the text recognition method adopts OCR text detection and recognition. For each target area image where a traffic sign is located, the specific recognition process is as follows:
[0061] S51, preprocessing the target area image to obtain a preprocessed target area image;
[0062] S52, using the CRAFT algorithm to detect the position of the text area on the preprocessed target area image to obtain a polygonal text area image for text detection;
[0063] S53, inputting the polygonal text area image into the ResNet+LSTM+CTC network structure to recognize the text content, and obtaining the text content of the text area;
[0064] S54. Use a greedy decoder to perform text meaning analysis on the text content, output the word with the highest probability, and obtain the text recognition result of each traffic sign.
[0065] S6. Perform a second deduplication process on the second data set according to the target ID of each traffic sign and the text recognition result thereof. For the image to be recognized after the second deduplication process, cut out the recognition area image of each traffic sign according to the coordinate position and label information of each traffic sign, and save the cut out recognition area image into the third data set;
[0066] Specifically, the second deduplication process in step S6 is as follows:
[0067] S61, obtaining a plurality of consecutive images to be identified in the second data set, comparing the target ID of each traffic sign in each image to be identified and its text recognition result, and if the text recognition results contained in each image to be identified are the same, retaining only the first image to be identified;
[0068] S62. When some traffic signs in the images to be identified have different target IDs and text recognition results, and some traffic signs have different target IDs but the same text recognition results, the markings of the traffic signs with different target IDs but the same text recognition results in the images to be identified are deleted.
[0069] S7. Input the images in the third data set into the YOLOv8 classification network for disease recognition to obtain the detection result of each traffic sign.
[0070] In order to better understand the specific process in this embodiment, the following is a detailed description:
[0071] 1. Build a YOLOv8 target detection optimization model
[0072] like Figure 2As shown in the figure, the constructed YOLOv8 target detection optimization model includes the backbone network (Backbone), the neck network (Neck) and the head network (Head). The backbone network is the basis of the model and is mainly responsible for extracting features from the input image. These features are the basis for subsequent network layers to perform target detection. In the constructed YOLOv8 target detection optimization model, the backbone network adopts a structure similar to CSPDarknet, and C2f in the local structure is replaced by C3STR structure. The head network is mainly responsible for the decision-making part of the target detection model and produces the final detection result. The neck network is located between the backbone network and the head network, and its function is to perform feature fusion and enhancement. Introduction to the various parts involved in the network structure:
[0073] ConvModule: It consists of a convolutional layer, BN (batch normalization), and an activation function (such as SiLU), and its function is to extract features.
[0074] DarknetBottleneck: Using residual connections to increase network depth while maintaining efficiency.
[0075] CSPLayer_2Conv: As a variant of the CSP structure, it uses partial connections to improve the training efficiency of the model.
[0076] SPPF: Reduces the amount of computation by optimizing the pooling operation while maintaining the model's detection performance for multi-scale targets;
[0077] Upsample: The main function is to enlarge the low-resolution feature map through upsampling so that it can be fused with the high-resolution feature. Upsampling enables the model to process features of different scales, thereby improving the detection ability of small objects. YOLOv8 uses multi-scale feature maps to enhance the detection performance of targets of different sizes.
[0078] concate: The concatenation layer is mainly used to concatenate feature maps from different layers through the above-mentioned upsampling operation. By concatenating features from different layers, the model can obtain more comprehensive contextual information, thereby improving the positioning and classification of the target. The concatenation operation enables the model to learn more complex feature representations, thereby improving the overall performance, which is especially important when dealing with complex scenes or small object detection.
[0079] Bbox_Loss (bounding box regression loss) calculates the difference between the predicted bounding box and the true bounding box. The mean square error (MSE) is a commonly used loss function that gives higher penalties when large errors occur. This feature helps the model quickly correct large prediction errors. Therefore, Bbox_Loss calculates the sum of the squares of the differences between the predicted and actual coordinates. The calculation formula is as follows:
[0080]
[0081] Among them, xi represents the coordinates of the true bounding box, and represents the coordinates of the predicted bounding box. This loss function is used as an optimization target to guide the model to reduce the gap between the predicted box and the true box during training.
[0082] Cls_Loss (classification loss) measures the difference between the class distribution predicted by the model and the true label. In classification tasks, the cross entropy loss function is a commonly used loss function that imposes a large penalty on wrong predictions, especially when the predicted probability is very different from the actual label. Therefore, Cls_Loss helps the model optimize its predictions in classification problems so that the predicted probability distribution is as close as possible to the true label distribution. Its calculation formula is:
[0083]
[0084] Among them, y o,c is an indicator. If sample o belongs to category c, it is 1, otherwise it is 0. o is the probability that the model predicts that sample o belongs to category c.
[0085] like Figure 3-Figure 4 As shown in the figure, the C3STR structure that replaces C2f uses a hierarchical construction method similar to that in convolutional neural networks. For example, the feature map size has downsampling of 4 times, 8 times, and 16 times. Such a backbone helps to build tasks such as target detection and instance segmentation on this basis, and has a good effect on the detection and recognition of small target detection tasks, such as Figure 3As shown in the figure, each C3STR structure includes a Patch Partition module for performing block operations on the image and four cascaded Stage modules for converting the block-operated image into feature maps of different sizes. Each Stage module includes a Swin Transformer Block network. The four Stage modules are Stage_1, Stage_2, Stage_3 and Stage4, where Stage_1 includes a cascaded LinearEmbeding layer and a Swin Transformer Block network, and Stage_2, Stage_3 and Stage4 each include a Patch Merging layer for downsampling and a Swin Transformer Block network. Figure 4 As shown, each SwinTransformer Block network includes a W-MSA structure and a SW-MSA structure used in pairs. The working process of the C3STR structure is:
[0086] First, the image is input into the Patch Partition module for block operation. Specifically, each 4×4 adjacent pixels is used as a patch, and then flattened in the channel direction. If the input is an RGB three-channel image, then each patch will contain 4×4 = 16 pixels, and each pixel has three values of R, G, and B, so after flattening it is 16×3 = 48. Therefore, after passing through the Patch Partition module, the shape of the image changes from [H, W, 3] to [H / 4, W / 4, 48]. Next, four stages are used to construct feature maps of different sizes. In Stage_1, the Linear Embedding layer performs a linear transformation on the channel data of each pixel, transforming it from 48 to C. After passing through a Swin Transformer Block network, the shape of the image changes from [H / 4, W / 4, 48] to [H / 4, W / 4, C]. In fact, the PatchPartition module and the LinearEmbeding layer are directly implemented with the help of a convolutional layer. Then in Stage_2, downsampling is first performed through a PatchMerging layer, and each 2×2 adjacent pixels are divided into a patch, and then the pixels at the same position in each patch are spliced together, so that 4 feature maps are obtained. Then through a SwinTransformer Block network, these four feature maps are spliced in the depth direction, and the LayerNorm layer and the fully connected layer MLP in the SwinTransformer Block network are used to implement linear changes in the depth direction of the feature map, so that the depth of the feature map changes from C to C / 2. It can be seen that after the image passes through the PatchMerging layer, the height and width of the feature map will be reduced by half, but the depth will double. W-MSA (Windows Multi-head Self-Attention) divides the feature map into Windows of size M×M, and then performs Self-Attention on each Windows separately.SW-MSA (Windows Multi-head Self-Attention) is an offset W-MSA. The window is offset from the upper left corner to the right and bottom by [] pixels respectively. This can solve the problem that the four independent windows cannot communicate with each other. Finally, after Stage_2, the shape of the image is further changed from [H / 4, W / 4, C] to [H / 8, W / 8, 2C]. The structure and processing process of Stage_3 and Stage_4 are the same as Stage_2. After Stage_3, the shape of the image changes from [H / 8, W / 8, 2C] to [H / 16, W / 16, 4C], and after Stage_4, the shape of the image changes from [H / 16, W / 16, 4C] to [H / 32, W / 32, 8C].
[0087] 2. ByteTrack target tracking algorithm
[0088] When the YOLOv8 target detection optimization model constructed above is used to obtain the position coordinates and label information of the traffic signs in the initial image, the target detection results are input into the target tracking algorithm ByteTrack for deduplication processing. Specifically, the ByteTrack algorithm is a tracking algorithm based on target detection. Like other non-ReID algorithms, it only relies on the bbox obtained by target tracking for tracking. This tracking algorithm uses Kalman filtering to predict the bounding box and uses the Hungarian algorithm to match the target with the trajectory. The biggest innovation of the ByteTrack algorithm lies in the use of low-resolution frames. Low-resolution frames may be frames generated when objects are occluded. If low-resolution frames are directly discarded, performance will be affected. Therefore, low-resolution frames can be used to perform secondary matching on the tracking algorithm, which effectively optimizes the problem of ID change caused by occlusion during the tracking process. This algorithm does not calculate the appearance similarity through ReID features, and is a non-deep method that does not require training. It effectively solves the occlusion problem by distinguishing and matching between high-resolution frames and low-resolution frames. Algorithm framework process: The main idea is to create tracking trajectories, and then use these tracking trajectories to match the target of each frame, complete the target matching operation frame by frame, and then form a complete trajectory. Two concepts need to be understood. The first is the tracking trajectory, which is created from the first image in the image set, which covers all trajectories of continuous tracking and interrupted tracking. The second is the bounding box of the current image. The bounding box of the current image is only the bounding box obtained by the current image and does not contain any information of previous images.
[0089] The tracking status includes the following parts:
[0090] ① Activation status: Activate the target frame that has tracked more than two frames of images (including the newly created track of the target frame of the first frame of the image. The target frame refers to the target frame obtained according to the coordinate position of the traffic sign);
[0091] ② Inactive state: A new track appears in the middle of the video, and the second point of the track has not yet been matched
[0092] ③New trajectory: Newly generated trajectory
[0093] ④ Tracked trajectory: The trajectory successfully tracked in the previous frame
[0094] ⑤ Lost tracking track: The track that lost tracking in the previous n frames (n<=10, in this application, the video stream is saved by skipping frames, so the frame number is set to 10 frames)
[0095] ⑥ Deleted tracks: tracks that lost tracking in the previous n frames (n>10)
[0096] When the algorithm is testing the first frame, since no track has appeared before, the algorithm will store all track objects created by the target box as tracked tracks. When the second frame is tested, the following steps are performed:
[0097] (1) Classify tracking trajectories and bounding boxes: All tracking trajectories are divided into activated and inactivated categories (activated trajectories have tracked the target box for more than two frames (including the newly created trajectories of the target box in the first frame)), and all bounding boxes in the current frame are divided into high-scoring and low-scoring categories.
[0098] (2) Perform the first tracking of the trajectory (only for the high-score matching of the activated trajectory): merge all the tracked trajectories and lost trajectories, which is called the preliminary tracking trajectory, and then predict the possible position and size of the next frame bounding box of the preliminary tracking trajectory, and calculate the IoU (intersection over union) value between the next frame bounding box predicted by the preliminary tracking trajectory and the high-score bounding box of the current frame, and then obtain a pairwise IoU relationship loss matrix (the smaller the IoU value, the greater the degree of association, and the maximum IoU value is 1, which means that there is no intersection between the bounding boxes). Finally, based on the IoU loss matrix, the Hungarian algorithm is used to match the preliminary tracking trajectory and the high-score bounding box of the current frame. Three results can be obtained: one is the matched trajectory and bounding box, the second is the unsuccessfully matched trajectory, and the third is the unsuccessfully matched current frame bounding box, and the successfully matched current frame bounding box is used to update the preliminary tracking trajectory.
[0099] (3) Track the trajectory for the second time (only for the low-score matching of the trajectory in the activated state): First, find the trajectories that were not matched in the first match, and filter out the tracked trajectories. Since these trajectories have already predicted the bounding box of the next frame, they do not need to be predicted here. Then calculate the IoU between the above trajectory and the low-score bounding box of the current frame, and use the Hungarian algorithm to match the above tracked trajectory with the low-score bounding box of the current frame. Update the above tracked trajectory with the successfully matched current frame bounding box, and mark the trajectory that has not been successfully tracked at this time as a lost trajectory.
[0100] (4) Tracking the inactivated trajectory state: First, find the current frame bounding box and the inactivated trajectory that were not successfully matched in the first step, calculate the IoU value between the above trajectory and the current frame bounding box, and then use the Hungarian algorithm to match the above trajectory and the bounding box, and use the successfully matched current frame bounding box to update the above tracking trajectory. At this time, the inactivated trajectory that was not successfully tracked is directly marked as a deleted trajectory.
[0101] (5) New track: If there is no successfully matched high-scoring bounding box, it is considered a new target and a new ID is assigned to it. The tracking task is then completed and all tracks will be returned with a unique target ID. This result is used as the tracking result for each target.
[0102] Through the target tracking algorithm ByteTrack, the label result of the unique target ID of each traffic sign in each image can be obtained. The traffic sign targets corresponding to the duplicate target IDs are removed and only one traffic sign target is retained to achieve the purpose of primary deduplication. The deduplicated traffic sign targets are then sent to the OCR text detection and recognition network for text recognition and secondary deduplication.
[0103] 3. OCR text detection and recognition
[0104] The target area image after deduplication by the ByteTrack target tracking algorithm is cropped and sent to the OCR text detection and recognition network for text recognition. First, the target area is pre-processed (Pre-Process), and then sent to the CRAFT (Character Region Awareness for Text Detection) algorithm to detect the position of the text area. After obtaining the polygonal text area for text detection, it is sent to the ResNet+LSTM+CTC network structure to identify the content of the text area, and then a greedy decoder is used to output the most likely word to obtain the text content of the target area.
[0105] The OCR text recognition content of the target area in the image is combined with the target ID of the traffic sign. When the target IDs in five consecutive frames are different but the OCR recognition content is exactly the same, they are considered to be the same traffic sign and secondary deduplication is performed, which can further reduce the number of duplicate targets.
[0106] 4. YOLOv8 classification and data transmission
[0107] The final deduplicated target area image is input into the YOLOv8 classification network to classify the defects of the traffic signs in the target area image (such as deformation, damage, and dislocation, etc.), and then the target detection results, target tracking results, OCR recognition results, and classification results of the image are transmitted to the web end for display.
[0108] The sign detection part of the present application is to improve the network of the YOLOv8 model by replacing two of the three C2f structures in the back with C3STR structures in the backbone of the network of the YOLOv8 model to improve the detection rate of the network model for small targets and occluded targets, and then obtain the coordinate position of the detection target sign through training with a pre-made traffic sign data set, and then send the coordinate information of the coordinate position and the label category information to the tracking algorithm ByteTrack to track each traffic sign as a target, and obtain the target ID of each target. Through the same target ID value, we can remove duplicate targets in the image. Since there are still a large number of duplicate targets that cannot be removed in the final result when the tracking algorithm is deduplicated, after analyzing the sign data features, the target area processed by the ByteTrack tracking algorithm is added with OCR recognition to identify the sign symbol content of the target area for secondary deduplication, and the targets with different tracking IDs but consistent OCR recognition content in the pictures of several frames before and after are considered to be the same target, so as to improve the probability of target deduplication, and finally the target area of secondary deduplication is cut off and sent to the classification network of YOLOv8 for disease category classification. Through the above algorithm and network improvements, the detection rate of marked small targets and occluded targets can be improved, as well as the efficiency of target deduplication. It can effectively reduce the workload of later manual compound data and achieve the effect of reducing costs and increasing efficiency.
[0109] Example 2
[0110] This embodiment provides a road traffic sign recognition and classification device, including:
[0111] one or more processors;
[0112] A storage unit is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors can implement a road traffic sign recognition and classification method in Example 1.
[0113] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.
Claims
1. A method for identifying and classifying road traffic signs, characterized in that: The specific steps include: S1. Obtain a video stream of road traffic signs, sample the video stream, obtain a number of initial images and store them in a first image set in sequence; S2. Build a YOLOv8 target detection optimization model, input the first image set into the YOLOv8 target detection optimization model, and obtain a number of images to be identified, each of which is marked with the coordinate position and label information of each traffic sign; The backbone network of the YOLOv8 target detection optimization model includes CSPDarknet structure, C2f structure and two C3STR structures; S3, according to the coordinate position and label information of each traffic sign, use the target tracking network ByteTrack to track each traffic sign as a target to obtain the target ID of each traffic sign; S4, performing a first deduplication process on the image to be identified according to the target ID of each traffic sign in each image to be identified, and storing the image to be identified after the first deduplication process in a second data set; S5, extracting the target area image where each traffic sign is located from the image to be recognized after the first deduplication process, performing text recognition on the target area image, and obtaining a text recognition result for each traffic sign; S6. Perform a second deduplication process on the second data set according to the target ID of each traffic sign and its text recognition result, and save the image to be recognized after the second deduplication process into a third data set; S7. Input the images in the third data set into the YOLOv8 classification network for disease recognition to obtain the detection result of each traffic sign.
2. A method for identifying and classifying road traffic signs according to claim 1, characterized in that: The specific process of the first deduplication process in step S4 is as follows: S41, comparing the target ID of each traffic sign in each image to be identified, if one of the images to be identified includes the target IDs of all traffic signs in another image to be identified, deleting the other image to be identified; S42. If the target ID of some traffic signs in one of the images to be identified is the same as that in the other image to be identified, then the annotations of the traffic signs with the same target ID are deleted from the other image to be identified, where the annotations include the coordinate positions and label information of the traffic signs.
3. A method for identifying and classifying road traffic signs according to claim 2, characterized in that: The text recognition method uses OCR text detection and recognition. For each target area image where a traffic sign is located, the specific recognition process is as follows: S51, preprocessing the target area image to obtain a preprocessed target area image; S52, using the CRAFT algorithm to detect the position of the text area on the preprocessed target area image to obtain a polygonal text area image for text detection; S53, inputting the polygonal text area image into the ResNet+LSTM+CTC network structure to recognize the text content, and obtaining the text content of the text area; S54. Use a greedy decoder to perform text meaning analysis on the text content, output the word with the highest probability, and obtain the text recognition result of each traffic sign.
4. A method for identifying and classifying road traffic signs according to claim 2, characterized in that: The second deduplication process in step S6 is as follows: S61, obtaining a plurality of consecutive images to be identified in the second data set, comparing the target ID of each traffic sign in each image to be identified and its text recognition result, and if the text recognition results contained in each image to be identified are the same, retaining only the first image to be identified; S62. When some traffic signs in the images to be identified have different target IDs and text recognition results, and some traffic signs have different target IDs but the same text recognition results, the markings of the traffic signs with different target IDs but the same text recognition results in the images to be identified are deleted.
5. The method for identifying and classifying road traffic signs according to claim 3, characterized in that: When the image after the second deduplication process is saved in the third data set, the following steps are also included: According to the coordinate position and label information of each traffic sign, the recognition area image of each traffic sign is cut out from the image to be recognized after the second deduplication process, and the cut recognition area image is saved in the third data set.
6. A method for identifying and classifying road traffic signs according to claim 1, characterized in that: Each C3STR structure includes a Patch Partition module for performing a block operation on the image and four cascaded Stage modules for converting the block-operated image into feature maps of different sizes. Each Stage module includes a Swin TransformerBlock network.
7. A method for identifying and classifying road traffic signs according to claim 6, characterized in that: The four Stage modules are Stage_1, Stage_2, Stage_3 and Stage4, where Stage_1 includes a cascaded LinearEmbeding layer and a Swin Transformer Block network, and Stage_2, Stage_3 and Stage4 each include a Patch Merging layer for downsampling and a Swin Transformer Block network.
8. A method for identifying and classifying road traffic signs according to claim 7, characterized in that: Each SwinTransformer Block network includes a W-MSA structure and a SW-MSA structure used in pairs.
9. A road traffic sign recognition and classification device, characterized in that: include: one or more processors; A storage unit, used to store one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement a road traffic sign recognition and classification method as described in any one of claims 1-8.
Citation Information
Patent Citations
Vehicle and pedestrian online detection and tracking method based on improved ByteTrack
CN116682078A
Highway pavement dynamic small target tracking detection method and system based on improved YOLOv5 and ByteTrack
CN119091394A