Traffic sign damage detection method and system, storage medium, and computer system

By enhancing and splicing the traffic sign images, forming an enhanced data set and training the YOLOv11 model, the misjudgment and misjudgment problems of traffic sign damage detection in the prior art are solved, and higher detection accuracy and robustness are achieved.

CN119723528BActive Publication Date: 2025-05-13CHONGQING URBAN CONSTR ECONOMIC (GRP) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510231877.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art has problems of misjudgment or misjudgment in traffic sign damage detection, especially in the case of complex backgrounds or slight damage to the sign, and further optimization of image processing and algorithms is needed to improve robustness.

Method used

By enhancing the data on road traffic sign images, including calculating the area proportion and debris occlusion ratio of traffic signs, performing image stitching and data perturbation, forming an enhanced data set, and using this data set to train the YOLOv11 model to improve the recognition accuracy of the model.

Benefits of technology

It significantly improves the accuracy and robustness of traffic sign damage detection, especially in the face of complex backgrounds and slight damage, and can more accurately identify and detect the status of traffic signs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723528B_ABST
    Figure CN119723528B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic sign damage detection method and system, storage medium, and computer system. The method comprises: collecting a number of road traffic sign images as a data set; enhancing the data set to form an enhanced data set; training a YOLOv11 model using the enhanced data set; and inputting the road traffic sign image to be identified into the trained YOLOv11 model for identification. In the present invention, the enhanced data set is used to train the YOLOv11 model, so that the model has excellent small target detection capabilities, can accurately deal with complex occlusion and contamination, and is highly adaptable to the diversity of real scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a traffic sign damage detection method and system, a storage medium, and a computer system. Background Art

[0002] With the continuous advancement of smart city construction, the intelligent and automated management of traffic facilities has gradually become a hot topic of research. As an important part of urban traffic management, there are roughly two methods for monitoring and damage detection of road traffic signs. One is the traditional traffic sign damage detection method, which mainly relies on manual inspections. Staff check the integrity of traffic signs through regular inspections. However, this method has many limitations. First, manual inspections are inefficient and have limited coverage, making it impossible to achieve full-time and full-domain monitoring. In addition, the time lag of manual inspections means that some traffic signs cannot be repaired in time for a long time after being damaged, increasing the safety risks of road traffic.

[0003] The other is a traffic sign damage detection method based on deep learning and computer vision technology. This method relies on image acquisition and analysis technology, and realizes rapid assessment of the status of traffic signs by automatically processing image data. Especially with the application of target detection algorithms such as YOLO (You Only Look Once), deep learning technology can efficiently and accurately identify the type, location and status of traffic signs, and detect whether the signs are damaged or blurred. With its high speed and high precision, the YOLO algorithm can process images in real time and accurately detect the damage of traffic signs, greatly improving the management efficiency of traffic signs, especially in large-scale urban traffic management. It has important application value.

[0004] At present, the application of YOLO and other deep learning technologies in traffic sign damage detection has made significant progress and has become an important means to solve the problems of low efficiency and insufficient coverage of traditional manual inspection methods. However, in some complex backgrounds or when the signs are slightly damaged, the algorithm may make misjudgments or miss judgments, and further optimization of image processing and algorithms is still needed to improve its robustness.

[0005] Prior art Chinese patent application CN202310560548.0 discloses a method for detecting insulator targets by drone aerial photography based on an improved YOLO algorithm, including: (1) drone aerial photography of the insulator scene to be applied, and extracting frames from the aerial video into image data; (2) annotating the image data, and dividing the collected and annotated data into a training set and a validation set in proportion; (3) hybrid enhancement of the training set data by the Mosaic+Mixup method; (4) using the Ghostnetv2 network to replace the YOLO v7 version backbone feature extraction network; (5) using a self-made data set to train the network; (6) exporting the trained model file as an ONNX file, and using TensorRT deployment to generate an Engine model; (7) using the Engine model to deploy and perform drone detection. This patent application constructs a target detection model based on the YOLO model, which has a better recognition effect for insulators with higher resolution images and smaller target areas, as well as occluded insulators, with higher recognition accuracy and more precise recognition positions.

[0006] Another prior art CN202011468618.2 discloses a deep learning image enhancement method based on annotation frame splicing, which also uses the Mosaic image enhancement method for image enhancement. Specifically, it includes the following steps: Step 1, select N pictures, scale the N pictures to the same size, and prepare a blackboard picture of the same size; Step 2, randomly sort several pictures, and randomly determine a splicing point in the picture; Step 3, crop the N pictures and the corresponding parts of each picture according to this ratio; Step 4, filter the annotation frame; Step 5, scale, transform, and rotate the filtered cropped area; Step 6, repeat steps 1 to 5.

[0007] It can be seen that enhancing the images in the dataset by the Mosaic image enhancement method and then using it to train the YOLO model can improve the image recognition efficiency. However, in practical applications, relying solely on the existing Mosaic image enhancement method to process images still has the problems of limited data diversity and insufficient feature mining, which makes the trained YOLO model need to improve its accuracy in traffic sign damage detection. Summary of the invention

[0008] The object of the present invention is to provide a traffic sign damage detection method and system, a storage medium, and a computer system, which partially solve or alleviate the above-mentioned deficiencies in the prior art, and can optimize the existing YOLO model to improve the model's recognition accuracy of traffic sign damage.

[0009] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions:

[0010] A traffic sign damage detection method, comprising:

[0011] Collecting a number of road traffic sign images as a data set, wherein the traffic signs in the road traffic sign images include road traffic signs in abnormal states and road traffic signs in normal states;

[0012] The data set is enhanced to form an enhanced data set, specifically comprising the following steps:

[0013] Calculate the area ratio of the traffic sign in each road traffic sign image;

[0014] The road traffic sign images whose traffic sign area ratio is less than the area threshold ratio are stitched and included in the enhanced dataset;

[0015] If the area ratio of the traffic sign in the road traffic sign image is greater than or equal to the area threshold ratio, the contamination and occlusion ratio of the traffic sign is identified;

[0016] When the contamination-occlusion ratio is greater than the first contamination-occlusion threshold, the corresponding road traffic sign images are spliced ​​and included in the enhanced data set;

[0017] When the contamination occlusion ratio is less than or equal to the first contamination occlusion threshold and greater than the second contamination occlusion threshold, the corresponding road traffic sign image is spliced ​​twice and then included in the enhanced data set;

[0018] When the contamination occlusion ratio is less than or equal to the second contamination occlusion threshold, a part of the contamination occlusion ratio is included in the enhanced data set, and the remaining part is perturbed and then included in the enhanced data set; the data perturbation includes at least one of flipping, scaling, and color gamut adjustment;

[0019] Train the YOLOv11 model using the enhanced dataset;

[0020] The road traffic sign image to be recognized is input into the trained YOLOv11 model for recognition.

[0021] As an improvement, the method for splicing road traffic sign images includes:

[0022] Select n road traffic sign images, and adjust the selected n road traffic sign images to have the same area;

[0023] Randomly select stitching points in a preset stitching image range;

[0024] Divide the preset stitching image range into n filling areas according to the stitching points;

[0025] The selected n road traffic sign images are randomly filled into n filling areas, and the areas of the road traffic sign images are adjusted to match the areas of the corresponding filling areas.

[0026] As an improvement, a stitching area is set within a preset stitching image range, and the stitching point is located within the stitching area; and the center point of the stitching area coincides with the center point of the preset stitching image range.

[0027] As an improvement, when performing the first stitching, 4 road traffic sign images are selected for stitching and the area is adjusted to 640 pixels * 640 pixels; the preset stitching image range is 1280 pixels * 1280 pixels; the area of ​​the stitching area is 640 pixels * 640 pixels;

[0028] When performing the second stitching, the areas of the four road traffic sign images after the first stitching are adjusted to 640 pixels*640 pixels before being stitched.

[0029] As an improvement, after the first splicing and the second splicing, data perturbation is performed on the spliced ​​road traffic sign image.

[0030] As an improvement, the step of adjusting the area of ​​the road traffic sign image to match the area of ​​the corresponding filling area includes:

[0031] If the area of ​​the road traffic sign image is larger than the area of ​​the corresponding filling area, the road traffic sign image is scaled down and then cropped;

[0032] If the area of ​​the road traffic sign image is smaller than the area of ​​the corresponding filling region, the road traffic sign image is enlarged in proportion and then cropped.

[0033] As an improvement, the step of selecting n road traffic sign images includes:

[0034] Select a road traffic sign image in an abnormal state, and then randomly select n-1 road traffic sign images.

[0035] The present invention also provides a traffic sign damage detection system, comprising:

[0036] A data set construction module, used for collecting a number of road traffic sign images as a data set, wherein the traffic signs in the road traffic sign images include road traffic signs in abnormal states and normal road traffic signs;

[0037] The data enhancement module is used to enhance the data set to form an enhanced data set, specifically comprising the following steps:

[0038] Calculate the area ratio of the traffic sign in each road traffic sign image;

[0039] The road traffic sign images whose traffic sign area ratio is less than the area threshold ratio are stitched and included in the enhanced dataset;

[0040] If the area ratio of the traffic sign in the road traffic sign image is greater than or equal to the area threshold ratio, the contamination and occlusion ratio of the traffic sign is identified;

[0041] When the contamination-occlusion ratio is greater than the first contamination-occlusion threshold, the corresponding road traffic sign images are spliced ​​and included in the enhanced data set;

[0042] When the contamination occlusion ratio is less than or equal to the first contamination occlusion threshold and greater than the second contamination occlusion threshold, the corresponding road traffic sign image is spliced ​​twice and then included in the enhanced data set;

[0043] When the contamination occlusion ratio is less than or equal to the second contamination occlusion threshold, a part of the contamination occlusion ratio is included in the enhanced data set, and the remaining part is included in the enhanced data set after data perturbation; the data perturbation includes at least one of flipping, scaling, and color gamut adjustment.

[0044] Training module, used to train the YOLOv11 model using the augmented dataset;

[0045] The recognition module is used to input the road traffic sign image to be recognized into the trained YOLOv11 model for recognition.

[0046] The present invention also provides a computer-readable storage medium, in which a computer program is stored; when the computer program is executed, the above-mentioned traffic sign damage detection method based on deep learning can be implemented.

[0047] The present invention also provides a computer system, comprising a processor and a memory; a computer program is stored in the memory; when the computer program is executed by the processor, the above-mentioned traffic sign damage detection method based on deep learning can be implemented.

[0048] Beneficial effects:

[0049] The present invention uses an enhanced data set to train the YOLOv11 model, which shows many significant advantages over the prior art.

[0050] First, excellent small target detection capability. By grading traffic sign images by area ratio, a splicing operation is performed specifically for low-level small targets, and multiple small target images are combined to focus the model on learning the faint but critical features of small targets. For example, in long-distance shooting scenarios, small speed limit and warning signs are often difficult to identify, but the YOLOv11 model can accurately capture their blurred outlines and key patterns remaining under partial occlusion after this training. This is what many existing technologies lack because they do not specifically strengthen small target learning. It greatly reduces the missed detection of small target traffic signs and improves the comprehensiveness of detection.

[0051] Second, it can accurately deal with complex occlusion and contamination. For small targets with a high percentage, the processing is subdivided according to the contamination-occlusion ratio. When the degree of occlusion or damage is large, for example, greater than 50%, one stitching allows the model to learn the special features of severely damaged signs after fusion with other images; in the range of 20%-50%, one stitching plus two stitchings are used to deepen the model's understanding of moderately damaged signs under diverse backgrounds and complex transformations. Compared with the single processing method of existing technologies, the YOLOv11 model can more delicately grasp the characteristics of signs with different degrees of damage. When faced with signs on actual roads, such as those blocked by large billboards or tree branches, or covered with graffiti or faded, it can accurately judge the type and status of the sign, and the detection accuracy has greatly increased.

[0052] Third, it is highly adaptable to the diversity of real-world scenarios. Starting from image grouping, it takes into account the similarity of sign features, the actual scene background, and the distribution of targets of different levels to ensure that the training samples fully simulate the real road conditions. The adaptive starting point optimization in the secondary stitching allows key small targets and moderately damaged high-level targets to be located in the central field of view of the image; dynamic light and shadow simulation reproduces the changes in light and shadow at different times and weather conditions; multi-scale cropping and combination enable the model to take into account both overall and detailed recognition. With this training, the YOLOv11 model can switch freely in complex scenes such as under urban neon lights, dusty rural environments, and strong light on highways, and can stably recognize traffic signs, with adaptability far exceeding traditional methods.

[0053] Fourth, efficient training and resource utilization. The entire enhanced data set construction process satisfies the model's demand for diverse data while reasonably balancing computing resources. For example, for relatively complete high-level targets, some are directly included in the training set to avoid unnecessary splicing; for low-proportion targets, they are spliced ​​at one time to prevent waste of resources for overly complex processing. Training the YOLOv11 model in this carefully designed data environment not only improves performance, but also accelerates convergence, reduces training time and hardware consumption, and is more feasible in actual deployment, enabling the rapid implementation of intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without paying creative labor.

[0055] Figure 1 This is a flow chart of Embodiment 1;

[0056] Figure 2 A schematic diagram of selecting images for stitching in Embodiment 1;

[0057] Figure 3 This is a schematic diagram of setting splicing points for splicing in Embodiment 1;

[0058] Figure 4 This is a schematic diagram of the detection results of road traffic sign damage detection using the optimized YOLOv11 model;

[0059] Figure 5 This is a structural diagram of the second embodiment. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0061] Herein, suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present invention, and have no specific meanings by themselves. Therefore, "module", "component" or "unit" can be used mixedly.

[0062] In this document, the terms "upper", "lower", "inner", "outer", "front", "back", "one end", "the other end" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.

[0063] In this document, unless otherwise clearly specified and limited, the terms "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0064] Herein "and / or" includes any and all combinations of one or more of the associated listed items.

[0065] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.

[0066] Embodiment 1:

[0067] like Figure 1 As shown, the present invention provides a traffic sign damage detection method, comprising:

[0068] S1 collects a number of road traffic sign images as a data set, where the traffic signs in the road traffic sign images include road traffic signs in abnormal states and road traffic signs in normal states.

[0069] In order to improve the generalization ability of the model, the road traffic sign images in the data set should cover as many traffic signs as possible on the road in various environments (including their corresponding backgrounds). That is, through various means, such as using cameras to shoot road traffic scenes in different sections, different time periods, and different weather conditions, a large number of road traffic sign images are obtained to form the original data set. These images cover the normal state of traffic signs, such as clear and complete, normal display of information, and abnormal states, such as damage, stains, and obstructions that result in incomplete or difficult to identify information, etc., providing comprehensive materials for subsequent model learning.

[0070] S2 enhances the data set to form an enhanced data set.

[0071] The purpose of enhancing the dataset is to generate more representative and diverse training samples, effectively simulating different damage scenarios and background changes. The specific steps include:

[0072] S21 calculates the area ratio of the traffic sign in each road traffic sign image.

[0073] The purpose of this step is to obtain the area ratio of the traffic sign in the road traffic sign image, so as to facilitate the subsequent steps to match different enhancement methods for each road traffic sign image according to the area ratio. In this embodiment, the specific steps of calculating the area ratio of the traffic sign in each road traffic sign image include:

[0074] S211 image preprocessing.

[0075] Use image processing libraries (such as OpenCV, etc.) to read each road traffic sign image in the dataset and convert the image format to a format that is easy to process, such as the common RGB format. This step ensures that all subsequent calculations and operations are performed on the same image data to avoid errors caused by format differences.

[0076] Then, according to actual needs, in order to simplify the calculation process and highlight the main body of the traffic sign, the color image can be converted into a grayscale image. Grayscale processing uses a specific weighted average algorithm, such as the common one that assigns different weights according to the sensitivity of the human eye to different colors, converts the values ​​of the three RGB channels into a grayscale value, reduces the data dimension while retaining the basic contour information of the image, which is conducive to the subsequent recognition and area calculation of traffic signs.

[0077] S212 Traffic sign recognition.

[0078] Pre-trained deep learning models for object detection are used to identify and locate traffic signs. These models have been trained on a large number of traffic sign samples and can accurately select the area where the traffic sign is located in a complex background image and output the bounding box coordinate information of the traffic sign, usually in the form of the coordinates of the upper left and lower right corners, such as (x1, y1) and (x2, y2), which provides a key basis for accurately calculating its area.

[0079] On the basis of model recognition, for some complex scenes where model recognition is inaccurate or ambiguous, images can be calibrated with the help of manual annotation. Professionals use image annotation tools to carefully outline the accurate contours of traffic signs to ensure the accuracy of the bounding box and further improve the reliability of area calculation. However, this step is relatively time-consuming and is usually used as an auxiliary supplement to intervene when model recognition deviates.

[0080] S213 area calculation and proportion solution.

[0081] Once the coordinates of the bounding box of the traffic sign are obtained, its area is calculated using a simple geometric formula. For example, for a rectangular bounding box, the area of ​​the traffic sign can be calculated using the coordinates of the diagonal points. At the same time, the pixel area of ​​the entire image is obtained, which can be easily obtained through the resolution information of the image (such as width w and height h, the pixel area is w*h). Finally, the pixel area of ​​the traffic sign is divided by the pixel area S of the entire image, and then multiplied by 100% to obtain the area ratio of the traffic sign in the image. The formula is expressed as: ratio = ((w * h) / S)*100%, which accurately quantifies the prominence of the traffic sign in the image and provides a key numerical basis for the subsequent implementation of the data enhancement strategy based on the ratio.

[0082] S22 stitches the road traffic sign images whose traffic sign area ratio is less than the area threshold ratio and incorporates them into the enhanced data set.

[0083] The purpose of image stitching is to expand the data and enhance the diversity of the data. Appropriate stitching greatly enriches the types of samples that the model is exposed to, effectively preventing the model from falling into the dilemma of overfitting due to a single learning sample, so that it can calmly cope with the traffic sign recognition needs under various real road conditions. However, if the stitching is excessive, the characteristics of the traffic sign will be diluted, which is not conducive to model capture. Therefore, in this embodiment, the classification is performed according to the area ratio of the traffic sign. For those with an area ratio less than a certain threshold, such as 25%, in this embodiment, only one stitching is used and then included in the enhanced data set for model training.

[0084] Specifically, the method for splicing road traffic sign images includes:

[0085] S221 selects n road traffic sign images, and adjusts the selected n road traffic sign images to have the same area.

[0086] like Figure 2 As shown, in this embodiment, each road traffic sign image has an index. When selecting images for splicing, first select an abnormal road traffic sign image 27 according to the index, and then randomly select n-1 road traffic sign images, such as Figure 2 Three images are randomly selected: image 100, image 308, and image 4005. The traffic signs in the randomly selected n-1 road traffic sign images may be in abnormal state or normal state. Usually, a traffic sign image includes at least one traffic sign, and the state of each traffic sign (i.e., small target) may be different, and the abnormal situation is also different.

[0087] The purpose of this selection is to ensure that the spliced ​​image includes at least one traffic sign in an abnormal state, which is convenient for the model to learn.

[0088] More specifically, in order to facilitate stitching, in this embodiment, four images are selected for stitching, that is, one image of a road traffic sign in an abnormal state is selected according to the index, and then three images of road traffic signs are randomly selected. Of course, other numbers of images may be selected and other arbitrary methods may be used for image stitching, which is not limited in the present invention.

[0089] For a solution to stitch 4 images, see Figure 3In one example, the preset stitching range is 1280 pixels * 1280 pixels. Before stitching, all the image areas in the data set need to be adjusted to 640 pixels * 640 pixels, so that the four images can be stitched into a large image of 1280 pixels * 1280 pixels. Of course, the stitching range can also be set according to actual needs.

[0090] S222 randomly selects a stitching point in a preset stitching image range.

[0091] In order to improve randomness, in this embodiment, the stitching point is not directly set in the middle of the stitching range, but a stitching area is set within a preset stitching image range (for example Figure 3 In the stitching range, O (0, 0) is the origin, the stitching area is a rectangular frame, the upper left corner vertex A (320, 320) of the rectangular frame; the lower right corner vertex B (960, -320), the stitching point C (590, 160) is located in the stitching area; and the center point D of the stitching area coincides with the preset center point of the stitching image range. Figure 3 As shown, the stitching area is a square area of ​​640 pixels*640 pixels, and the stitching points C are randomly set in the square area.

[0092] S223 divides the preset stitching image range into n filling areas according to the stitching points.

[0093] After the stitching point is determined, the entire stitching range can be divided into four filling areas using a horizontal line and a vertical line that pass through the stitching point, which are used to place four images for stitching.

[0094] S224 randomly fills the selected n road traffic sign images into n filling regions, and adjusts the area of ​​the road traffic sign image to match the area of ​​the corresponding filling region.

[0095] It is understandable that since the stitching point is not necessarily in the middle of the preset stitching range, the areas of the four filling areas are not necessarily equal. Some may be larger than 640 pixels * 640 pixels, while some may be smaller than 640 pixels * 640 pixels. And each filling area is not necessarily a square. The images in the data set are all 640 pixels * 640 pixels square. In order to completely fill the filling area, the image needs to be scaled and / or cropped.

[0096] Specifically, if the area of ​​the road traffic sign image is larger than the area of ​​the corresponding filling area, the road traffic sign image is scaled down and then cropped. If the area of ​​the road traffic sign image is smaller than the area of ​​the corresponding filling area, the road traffic sign image is scaled up and then cropped.

[0097] For example, if the image slightly exceeds the boundary of the filling area in a certain direction, it is scaled down and cropped according to the rules so that it just fills the filling area completely; if the image is smaller than the filling area, it is scaled up and the excess part is cropped to make it fit. Of course, if the filling area itself is a square and the size is corresponding, it only needs to be scaled without cropping.

[0098] Through these steps, the first image stitching is completed, and a 1280 pixel * 1280 pixel stitching image that integrates multiple traffic signs, different scenes and status information is obtained.

[0099] The above is an enhancement method for road traffic sign images whose area ratio is less than the area threshold ratio. For road traffic sign images whose area ratio is greater than or equal to the area threshold ratio, the enhancement method is as follows:

[0100] S23: If the area ratio of the traffic sign in the road traffic sign image is greater than or equal to the area threshold ratio, then identifying the contamination and occlusion ratio of the traffic sign.

[0101] For areas where the percentage is greater than or equal to the threshold percentage, the defacement and occlusion ratio of the traffic sign must first be calculated. The defacement and occlusion ratio refers to the ratio of the area of ​​the defaced area and the obscured part of the traffic sign to the area of ​​the entire traffic sign. For example, if there is a rectangular traffic sign, its complete area is calculated to be 100 square pixels, and an area of ​​20 square pixels is damaged due to graffiti in the image, and 10 square pixels are obscured by branches, then the defacement and occlusion ratio is (20+10) ÷ 100 = 30%. This is an example of an image with only one traffic sign.

[0102] In addition, in this embodiment, a first contamination occlusion threshold and a second contamination occlusion threshold are set, and the road traffic sign image is classified again, specifically including:

[0103] In step S231, when the contamination and occlusion ratio is greater than the first contamination and occlusion threshold, the corresponding road traffic sign images are spliced ​​and incorporated into the enhanced data set.

[0104] If the degree of occlusion or damage is greater than the set first contamination threshold (such as 50%), a stitching is also performed. The reason is that the characteristics of severely occluded or damaged high-level targets are quite different from those of complete signs. Through a single stitching and fusion with other images, it helps the model learn such special cases.

[0105] The specific method of stitching images is shown in step S22 and will not be described in detail in this step.

[0106] S232: When the contamination occlusion ratio is less than or equal to the first contamination occlusion threshold and greater than the second contamination occlusion threshold, the corresponding road traffic sign image is spliced ​​twice and then included in the enhanced data set.

[0107] If the degree of occlusion or damage is greater than the second contamination occlusion threshold (such as 20%), but less than the first contamination occlusion threshold, then perform primary and secondary stitching. That is, first perform primary stitching according to the process of step S22, so that the target is initially integrated with the other three images to obtain more background and related sign information; then incorporate the image after primary stitching into the secondary stitching process, and use the rich transformations and combinations of secondary stitching to further enhance the model's recognition ability for such moderately occluded or damaged high-level targets. It is worth noting that when performing the second stitching, the area of ​​the four road traffic sign images after the first stitching is adjusted to 640 pixels * 640 pixels before stitching.

[0108] Including the image after the first stitching into the second stitching process specifically means: stitching the stitched image obtained after the first stitching again. Specifically, randomly select four stitched images after the first stitching. Since the image size of the four stitched images is larger after the first stitching, first perform the same reduction operation on the four stitched images (for example, since the size of each stitched image after the first stitching is 1280 pixels * 1280 pixels, it is reduced to 640 pixels * 640 pixels here); then randomly select stitching points in the preset stitching image range, and divide the preset stitching image range into 4 filling areas according to the stitching points, and randomly fill the selected four stitching images into the 4 filling areas, and adjust their areas to match the corresponding filling area areas, that is, execute steps S222-S224 to obtain the second stitching image.

[0109] Of course, in other embodiments, the same calculation method as above may be used to calculate the area ratio of all traffic signs in the image after the first stitching, and identify the number of traffic signs whose area ratio is greater than or equal to the preset area threshold. If the number is greater than the preset number threshold, it is used as an alternative image for secondary stitching, and then the contamination and occlusion ratio of each traffic sign in the alternative image is calculated, and four alternative images whose sum of contamination and occlusion ratios (that is, the sum of the contamination and occlusion ratios of all traffic signs in each alternative image) has a difference within the preset contamination and occlusion ratio difference threshold range (for example, 0%-15%) are selected from all the alternative images for secondary stitching.

[0110] Specifically, since each stitched image obtained after the first stitching includes multiple traffic signs, the area ratio between each traffic sign in each stitched image and the corresponding stitched image is first calculated. If the number of traffic signs whose area ratio is greater than a preset area threshold (such as greater than 20%) is greater than a first preset number threshold (for example, greater than 3) (or the sum of the total area ratios of all traffic signs in each stitched image is greater than 50%), the stitched image is used as a candidate image for secondary stitching; then the defacement occlusion ratio of each traffic sign in the candidate image is obtained, and when the number of traffic signs whose defacement occlusion ratio is greater than the first defacement occlusion threshold is greater than a second preset number threshold (for example, 2), the corresponding candidate image is not used as a secondary stitching object; and when the number of traffic signs whose defacement occlusion ratio is less than or equal to the first defacement occlusion threshold and greater than the second defacement occlusion threshold is greater than a third preset number threshold (for example, 3), the corresponding candidate image is used as a secondary stitching object for secondary stitching and then included in the enhanced data set. Furthermore, when performing secondary stitching, four images whose difference in the sum of the pollution occlusion ratios (i.e., the sum of the pollution occlusion ratios of all traffic signs in the image) is within a preset pollution occlusion ratio difference threshold range (e.g., 0%-15%, 0% means no difference) are selected from the secondary stitching objects for secondary stitching. Of course, if the number of candidate images whose difference in the sum of the pollution occlusion ratios between two of them is within the preset pollution occlusion ratio difference threshold range exceeds four, four images are randomly selected for secondary stitching. For example, if the candidate images include images I-VIII, the difference in the sum of the stain occlusion ratios between any two of the five images: image I, image II, image V, image VI, and image VIII (such as the difference between the sum of the stain occlusion ratios of image I and the sum of the stain occlusion ratios of image II, the difference between the sum of the stain occlusion ratios of image I and the sum of the stain occlusion ratios of image V, the difference between the sum of the stain occlusion ratios of image I and the sum of the stain occlusion ratios of image VI, the difference between the sum of the stain occlusion ratios of image I and the sum of the stain occlusion ratios of image VIII; and the sum of the stain occlusion ratios of image II and image V is greater than or equal to the sum of the stain occlusion ratios of image I and image VI; The difference between the total stain occlusion ratios of image II and image VI, the difference between the total stain occlusion ratios of image II and image VIII; and the difference between the total stain occlusion ratios of image V and image VI, the difference between the total stain occlusion ratios of image V and image VIII; and the difference between the total stain occlusion ratios of image VI and image VIII) are all within the preset stain occlusion ratio difference threshold range, and four images are randomly selected from the five images for secondary stitching.

[0111] After the first stitching, some images may be enlarged and some images may be reduced, that is, the area ratio of some traffic signs in the original image (i.e., the image that has not been stitched) has undergone substantial changes. For example, the area ratio of the originally larger area ratio becomes smaller because the image is reduced, and the area ratio of the originally smaller area ratio becomes larger because the image is scaled, etc. Therefore, it is necessary to screen the secondary stitching images according to the area ratio.

[0112] For traffic signs with a large area ratio and less contamination and occlusion, the two-time stitching strategy has significant advantages. On the one hand, traffic signs with a large area ratio are often the main visual elements in the picture and are easier to capture clearly, but in reality they face more complex and diverse contamination and occlusion. In the first stitching, it is combined with three other images with different backgrounds and states to initially enrich the surrounding information environment and allow the model to learn the feature changes of the sign in common scenes. For example, a large, slightly contaminated road sign located at a busy intersection in the city is spliced ​​with a clear warning sign on a rural road and a speed limit sign under the changing light and shadow of a highway. The model begins to understand the presentation characteristics of large-area signs in different scenarios.

[0113] The second stitching is to further deepen the model's understanding of the complex situations of such signs on this basis. By integrating more diversified image information again and simulating extremely complex scenes such as multiple occlusions and multiple refractions and reflections of light, the model can accurately grasp the characteristic evolution of large-area signs under various harsh conditions. In actual detection, when encountering large traffic signs such as those partially blocked by large billboards and affected by strong backlight, the model can still accurately judge the type and status of the sign based on the learned complex features, greatly improving the accuracy and reliability of detection and meeting the high requirements of road traffic sign damage detection for dealing with complex scenes.

[0114] S233: When the contamination occlusion ratio is less than or equal to the second contamination occlusion threshold, a part of the contamination occlusion ratio is included in the enhanced data set, and the remaining part is perturbed and then included in the enhanced data set; the data perturbation includes at least one of flipping, scaling, and color gamut adjustment.

[0115] In this embodiment, for such relatively complete high-level targets, no stitching operation is performed. Instead, their original images are directly included in the training set with a certain probability (e.g., 60%), so that the model can prioritize learning complete and clear high-level target features and enhance its ability to recognize standard traffic signs.

[0116] The remaining 40% of the data set is slightly perturbed. For example, the image is flipped horizontally or vertically, scaled by a certain ratio, or only a small amount of color gamut adjustment is performed to simulate the visual difference caused by slight changes in light or slight color difference of the shooting equipment in reality; or extremely slight scaling is performed with the range of change controlled within ±5% to simulate the size change caused by a slight change in shooting distance, so that the model can adapt to slight environmental changes based on learning standard features, while avoiding excessive data processing to destroy the original clear features.

[0117] It is understandable that after the first splicing and the second splicing, the spliced ​​road traffic sign images can also be perturbed to increase the sample size. The specific method of data perturbation is as described above and will not be repeated here.

[0118] S3 uses the augmented dataset to train the YOLOv11 model.

[0119] The YOLOv11 model consists of a backbone network, a neck network, and a detection head. It uses an optimized CSPDarknet53 as the backbone network and generates feature maps of different scales through five downsamplings. In order to enhance the feature extraction capability, YOLOv11 uses the C3K2 module to replace the original C2f module, and introduces the spatial pyramid pooling fast module (SPFF) to improve the diversity of feature expression. In the neck network part, YOLOv11 uses the PAN-FPN structure to fuse shallow position information and deep semantic information through a bottom-up path, thereby making up for the lack of positioning information in the FPN structure. In general, through these innovative optimization designs, YOLOv11 has improved the accuracy and efficiency of real-time target detection and demonstrated its strong potential in computer vision tasks.

[0120] In this embodiment, the enhanced data set obtained in step S2 is used to train the YOLOv11 model, which effectively enhances the model's ability to handle various changes that may be encountered in actual scenes, such as different angles, sizes, and lighting. By training on diverse training data, the model can improve its performance in the test set, especially when faced with unseen images or complex backgrounds.

[0121] See also Figure 4, the YOLOv11 model optimized by the enhanced data set in the present invention is used to detect the abnormal state of traffic signs, and the occlusion of the road traffic signs in one image is 0.93 (i.e., the contamination occlusion ratio is 0.93), and the loss of the road traffic signs in another image is 0.76 (i.e., the contamination occlusion ratio is 0.76). In order to verify the specific effect of the present invention, SSD, Faster-RCNN, and unoptimized YOLOv11 models are used for detection, and the results of detecting the abnormal state of traffic signs are compared with the YOLOv11 model optimized by the enhanced data set in the present invention. The comparison results are shown in Table 1.

[0122] Table 1 Comparison of detection results of various models

[0123]

[0124] It can be seen that the YOLOv11 model optimized with the enhanced dataset has a detection accuracy of 92.34% in traffic sign anomaly detection, which is higher than the detection accuracy of SSD, Faster RCNN, and YOLOv11 algorithms.

[0125] S4 inputs the road traffic sign image to be recognized into the trained YOLOv11 model for recognition.

[0126] By collecting traffic sign images on the road through surveillance cameras, road-mounted collection vehicles, drones, etc., and inputting the images into the YOLOv11 model optimized by the enhanced data set in the present invention, the result of whether the road traffic sign image is damaged can be obtained.

[0127] Embodiment 2:

[0128] like Figure 5 As shown, the present invention also provides a traffic sign damage detection system based on deep learning, comprising:

[0129] A data set construction module, used for collecting a number of road traffic sign images as a data set, wherein the traffic signs in the road traffic sign images include road traffic signs in abnormal states and normal road traffic signs;

[0130] The data enhancement module is used to enhance the data set to form an enhanced data set, specifically comprising the following steps:

[0131] Calculate the area ratio of the traffic sign in each road traffic sign image;

[0132] The road traffic sign images whose traffic sign area ratio is less than the area threshold ratio are stitched and included in the enhanced dataset;

[0133] If the area ratio of the traffic sign in the road traffic sign image is greater than or equal to the area threshold ratio, the contamination and occlusion ratio of the traffic sign is identified;

[0134] When the contamination-occlusion ratio is greater than the first contamination-occlusion threshold, the corresponding road traffic sign images are spliced ​​and included in the enhanced data set;

[0135] When the contamination occlusion ratio is less than or equal to the first contamination occlusion threshold and greater than the second contamination occlusion threshold, the corresponding road traffic sign image is spliced ​​twice and then included in the enhanced data set;

[0136] When the contamination occlusion ratio is less than or equal to the second contamination occlusion threshold, a part of the contamination occlusion ratio is included in the enhanced data set, and the remaining part is included in the enhanced data set after data perturbation; the data perturbation includes at least one of flipping, scaling, and color gamut adjustment.

[0137] Training module, used to train the YOLOv11 model using the augmented dataset;

[0138] The recognition module is used to input the road traffic sign image to be recognized into the trained YOLOv11 model for recognition.

[0139] Embodiment three:

[0140] The present invention also provides a storage medium, in which a computer program is stored; when the computer program is executed, the above-mentioned traffic sign damage detection method based on deep learning can be implemented.

[0141] Embodiment 4:

[0142] The present invention also provides a computer system, comprising a processor and a memory; a computer program is stored in the memory; when the computer program is executed by the processor, the above-mentioned traffic sign damage detection method based on deep learning can be implemented.

[0143] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0144] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a computer terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0145] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A traffic sign damage detection method, characterized in that include: Collecting a number of road traffic sign images as a data set, wherein the traffic signs in the road traffic sign images include road traffic signs in abnormal states and road traffic signs in normal states; The data set is enhanced to form an enhanced data set, specifically comprising the following steps: Calculate the area ratio of the traffic sign in each road traffic sign image; The road traffic sign images whose traffic sign area ratio is less than the area threshold ratio are stitched and included in the enhanced dataset; If the area ratio of the traffic sign in the road traffic sign image is greater than or equal to the area threshold ratio, the contamination and occlusion ratio of the traffic sign is identified; When the contamination-occlusion ratio is greater than the first contamination-occlusion threshold, the corresponding road traffic sign images are spliced ​​and included in the enhanced data set; When the contamination occlusion ratio is less than or equal to the first contamination occlusion threshold and greater than the second contamination occlusion threshold, the corresponding road traffic sign image is spliced ​​twice and then included in the enhanced data set; When the contamination occlusion ratio is less than or equal to the second contamination occlusion threshold, a part of the contamination occlusion ratio is included in the enhanced data set, and the remaining part is perturbed and then included in the enhanced data set; the data perturbation includes at least one of flipping, scaling, and color gamut adjustment; Train the YOLOv11 model using the enhanced dataset; The road traffic sign image to be recognized is input into the trained YOLOv11 model for recognition.

2. A traffic sign damage detection method according to claim 1, characterized in that Methods for splicing road traffic sign images include: Select n road traffic sign images, and adjust the selected n road traffic sign images to have the same area; Randomly select stitching points in a preset stitching image range; Divide the preset stitching image range into n filling areas according to the stitching points; The selected n road traffic sign images are randomly filled into n filling areas, and the areas of the road traffic sign images are adjusted to match the areas of the corresponding filling areas.

3. A traffic sign damage detection method according to claim 2, characterized in that: A stitching area is set within a preset stitching image range, and the stitching point is located within the stitching area; and a center point of the stitching area coincides with a center point of the preset stitching image range.

4. A traffic sign damage detection method according to claim 3, characterized in that: When performing the first stitching, 4 road traffic sign images are selected for stitching and the area is adjusted to 640 pixels * 640 pixels; the preset stitching image range is 1280 pixels * 1280 pixels; the area of ​​the stitching area is 640 pixels * 640 pixels; When performing the second stitching, the areas of the four road traffic sign images after the first stitching are adjusted to 640 pixels*640 pixels before being stitched.

5. A traffic sign damage detection method according to claim 4, characterized in that: After the first splicing and the second splicing, data perturbation is performed on the spliced ​​road traffic sign image.

6. A traffic sign damage detection method according to claim 2, characterized in that The step of adjusting the area of ​​the road traffic sign image to match the area of ​​the corresponding filling area includes: If the area of ​​the road traffic sign image is larger than the area of ​​the corresponding filling area, the road traffic sign image is scaled down and then cropped; If the area of ​​the road traffic sign image is smaller than the area of ​​the corresponding filling region, the road traffic sign image is enlarged in proportion and then cropped.

7. A traffic sign damage detection method according to claim 2, characterized in that The steps of selecting n road traffic sign images include: Select a road traffic sign image in an abnormal state, and then randomly select n-1 road traffic sign images.

8. A traffic sign damage detection system, characterized in that include: A data set construction module, used to collect a number of road traffic sign images as a data set, wherein the traffic signs in the road traffic sign images include road traffic signs in abnormal states and normal road traffic signs; The data enhancement module is used to enhance the data set to form an enhanced data set, specifically comprising the following steps: Calculate the area ratio of the traffic sign in each road traffic sign image; The road traffic sign images whose traffic sign area ratio is less than the area threshold ratio are stitched and included in the enhanced dataset; If the area ratio of the traffic sign in the road traffic sign image is greater than or equal to the area threshold ratio, the contamination and occlusion ratio of the traffic sign is identified; When the contamination-occlusion ratio is greater than the first contamination-occlusion threshold, the corresponding road traffic sign images are spliced ​​and included in the enhanced data set; When the contamination occlusion ratio is less than or equal to the first contamination occlusion threshold and greater than the second contamination occlusion threshold, the corresponding road traffic sign image is spliced ​​twice and then included in the enhanced data set; When the contamination occlusion ratio is less than or equal to the second contamination occlusion threshold, a part of the contamination occlusion ratio is included in the enhanced data set, and the remaining part is perturbed and then included in the enhanced data set; the data perturbation includes at least one of flipping, scaling, and color gamut adjustment; The training module is used to train the YOLOv11 model using the enhanced dataset; The recognition module is used to input the road traffic sign image to be recognized into the trained YOLOv11 model for recognition.

9. A storage medium, characterized in that: The storage medium stores a computer program; when the computer program is executed, the traffic sign damage detection method described in any one of claims 1 to 7 is implemented.

10. A computer system, characterized in that: The invention comprises a processor and a memory; the memory stores a computer program; when the computer program is executed by the processor, the traffic sign damage detection method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A deep learning image enhancement method based on annotation box splicing

    CN112508836B

  • Unmanned aerial vehicle aerial photography insulator target detection method based on improved Yov algorithm

    CN116563736A

  • Damaged traffic sign detection method and device

    CN116580378A

  • Lake surface floating object small target detection method based on deep learning

    CN117237614A