Traffic sign identification method for automatic driving

By building a multi-layer perceptron model to optimize gamma factors and improve the YOLOv8 network model, the problem of low visibility and low accuracy of traffic sign recognition is solved, efficient recognition under low visibility conditions is achieved, and the safety of autonomous driving is improved.

CN120388351APending Publication Date: 2025-07-29LIAONING UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510459383.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Under low visibility, existing autonomous driving systems are difficult to accurately identify traffic signs, affecting driving safety.

Method used

A multi-layer perceptron model is built to optimize gamma factors, combined with the YOLOv8 network model for traffic sign image enhancement and detection, optimize model parameters through genetic algorithms, and improve the YOLOv8 model to improve recognition accuracy and efficiency.

Benefits of technology

In the case of low visibility, the accuracy and efficiency of traffic sign recognition are significantly improved and the safety of autonomous driving is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388351A_ABST
    Figure CN120388351A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving-oriented traffic sign recognition method, which comprises the following steps of: constructing a multi-layer perceptron model by taking a brightness average value and a saturation average value of all pixel points in a traffic sign image as input parameters and taking a gamma factor for performing power law transformation on the traffic sign image as an output parameter; optimizing parameters of the multi-layer perceptron model by taking the maximum objective function as an optimization objective to obtain an optimized multi-layer perceptron model; wherein the objective function is F = a < SSIM > + b < PSNR >; in the vehicle driving process, after a traffic sign image is obtained in real time through a vehicle-mounted camera, the brightness mean value and the saturation mean value of all pixel points in the traffic sign image are calculated and input into the optimized multi-layer perceptron model, and the optimal gamma factor of power law transformation is obtained; performing power law transformation on the traffic sign image by using the optimal gamma factor to obtain an enhanced traffic sign image; and performing target detection on the enhanced traffic sign image based on a YOLOv8 network model, and identifying the traffic sign in the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to a traffic sign recognition method for autonomous driving. Background Art

[0002] For an automotive autonomous driving system, the automatic detection and recognition of traffic signs is a very important part. A lot of important information is contained in traffic signs, and correctly recognizing traffic signs can reduce many unnecessary traffic accidents and protect the life and property safety of passengers. Currently, autonomous driving vehicles all collect through an in-vehicle vision system and then recognize through image detection. However, during actual driving, weather conditions will affect the accurate recognition of traffic signs. Especially in weather conditions with low visibility such as rain, snow, and haze, it often leads to errors in traffic sign recognition and affects the safety of autonomous driving. Summary of the Invention

[0003] The purpose of the present invention is to overcome the defects of the prior art and provide a traffic sign recognition method for autonomous driving, which can improve the accuracy of traffic sign recognition under low visibility conditions and also improve the recognition efficiency of traffic sign recognition.

[0004] The technical solution provided by the present invention is as follows:

[0005] A traffic sign recognition method for autonomous driving includes the following steps:

[0006] Step 1: Using the average brightness value and average saturation value of all pixel points in the traffic sign image as input parameters, and the gamma factor for power-law transformation of the traffic sign image as the output parameter, construct a multi-layer perceptron model;

[0007] Step 2: Optimize the parameters of the multi-layer perceptron model with the maximum of the objective function as the optimization goal to obtain an optimized multi-layer perceptron model;

[0008] Wherein, the objective function is: F = aSSIM + bPSNR;

[0009] In the formula, SSIM and PSNR are respectively the SSIM evaluation index value and PSNR evaluation index value of the traffic sign image after power-law transformation, and both a and b are weight coefficients;

[0010] Step 3: During the vehicle driving process, after obtaining the traffic sign image in real time through an in-vehicle camera, calculate the average brightness value and average saturation value of all pixel points in the traffic sign image, and input them into the optimized multi-layer perceptron model; the optimized multi-layer perceptron model outputs the optimal gamma factor for power-law transformation of the traffic sign image;

[0011] Step 4: Perform power-law transformation on the traffic sign image using the optimal gamma factor to obtain an enhanced traffic sign image;

[0012] Step 5: Based on the YOLOv8 network model, perform object detection on the enhanced traffic sign image to identify the traffic signs in the image.

[0013] Preferably, in Step 2, the genetic algorithm is used to optimize the weights and biases of the multi-layer perceptron model, including the following steps:

[0014] Step 1: Initialize the weights and biases in each layer of the multi-layer perceptron model to obtain multiple arrays composed of the weights and biases of each layer as the initial population;

[0015] Among them, each array corresponds to a multi-layer perceptron model;

[0016] Step 2: Calculate the fitness of each individual in the current population;

[0017] f i = aSSIM i + bPSNR i ;

[0018] Among them, f i represents the fitness of the i-th individual in the population, SSIM i represents the SSIM evaluation index value corresponding to the i-th individual in the population, and PSNR i represents the PSNR evaluation index value corresponding to the i-th individual in the population;

[0019] Step 3: Select the individuals with large fitness in the population as the parents, and use the parents to generate the next generation population through crossover and mutation;

[0020] Step 4: Loop through Steps 2 - 3 until the maximum number of iterations is reached; select the array corresponding to the individual with the largest fitness as the weights and biases of the multi-layer perceptron model to obtain an optimized multi-layer perceptron model.

[0021] Preferably, the hidden layer of the multi-layer perceptron model is 2 layers.

[0022] Preferably, the number of hidden layer nodes is determined according to the following formula:

[0023]

[0024] Among them, N hid is the number of hidden layer nodes; N in represents the number of input layer nodes; N out represents the number of output layer nodes; m represents any integer between 2 and 6, and [] represents rounding up.

[0025] Preferably, in the objective function, the values of the weight parameters satisfy: a + b = 1, a, b ∈ (0, 1);

[0026] Among them, if the difference between the brightness mean and the saturation mean of all pixel points is greater than the set threshold, a > b, otherwise a ≤ b.

[0027] Preferably, in the fifth step, it further includes improving the YOLOv8 model and performing object detection based on the improved YOLOv8 model;

[0028] The method for improving the YOLOv8 model is: replacing the 2nd to 4th C2f modules in the Backbone part with Ghost modules; in the Neck part, adding CBAM modules after each C2f module and before each Upsample structure.

[0029] The beneficial effects of the present invention are:

[0030] The traffic sign recognition method for autonomous driving provided by the present invention can improve the accuracy of traffic sign recognition under low visibility conditions and enhance the recognition efficiency of traffic sign recognition, thereby improving the safety of autonomous driving. Description of the Drawings

[0031] Figure 1 It is a framework diagram of the improved YOLOv8 network model described in the present invention. Detailed Embodiments

[0032] The following further elaborates the present invention with reference to the drawings, so that those skilled in the art can implement it according to the description in the specification.

[0033] The present invention provides a traffic sign recognition method for autonomous driving, and the specific implementation process is as follows.

[0034] I. Construct a multi-layer perceptron model

[0035] The present invention mainly aims to recognize traffic signs under visibility conditions. Under low visibility conditions, due to the low clarity of traffic sign pictures obtained by in-vehicle cameras, if directly using a picture recognition model for recognition, it will lead to low recognition accuracy or even inability to recognize. Therefore, it is necessary to enhance the traffic sign pictures before recognition. The present invention uses the method of power-law transformation for image enhancement processing. The expression of power-law transformation is: S = cI γ ; where I represents the input image, S represents the output image, c is the gray-scale scaling coefficient, usually taken as 1; γ is the gamma factor, which controls the scaling degree of the entire transformation. γ is usually manually set according to experience.

[0036] However, during the vehicle driving process, the visibility is not always constant. That is to say, the clarity of the traffic sign pictures (images) collected is also constantly changing. If a fixed gamma factor is used for power-law transformation, it may lead to insufficient image enhancement and inaccurate recognition; or over-enhanced images, affecting the recognition efficiency. Therefore, it is necessary to adjust the gamma factor according to the visibility to achieve better image recognition effects.

[0037] The measurement of visibility needs to be achieved through complex instruments or computer methods such as deep learning. These methods are too complex and not suitable for in-vehicle systems. Through research, it is found that there is a strong correlation between the brightness and saturation of pixel points in the image and the visibility. Therefore, in this application, the brightness and saturation of pixel points and the visibility are used to characterize the visibility. However, the relationship between the brightness and saturation of pixel points and the gamma factor cannot be expressed by a conventional linear relationship. Therefore, in the present invention, a multi-layer perceptron is used to realize the mapping between the brightness and saturation of pixel points and the gamma factor.

[0038] Taking the average brightness and average saturation of all pixel points in the traffic sign image as input parameters, and the gamma factor for power-law transformation of the traffic sign image as the output parameter, a multi-layer perceptron model is constructed.

[0039] In this embodiment, the hidden layer of the multi-layer perceptron model is set to 2 layers.

[0040] As a preference, the number of hidden layer nodes of the multi-layer perceptron model is determined according to the following formula:

[0041]

[0042] Where N hid is the number of hidden layer nodes; N in represents the number of input layer nodes; N out represents the number of output layer nodes; m represents any integer between 2 and 6, and [] represents rounding up.

[0043] Second, optimize the parameters of the multi-layer perceptron model with the maximum of the objective function as the optimization goal to obtain an optimized multi-layer perceptron model.

[0044] Since the gamma factor output by the multi-layer perceptron model plays an important role in the recognition effect of the image. Therefore, in the present invention, the effect evaluation index after image enhancement is used as the optimization goal for multi-layer perceptron training.

[0045] Among them, the objective function is set as: F = aSSIM + bPSNR;

[0046] Wherein, SSIM and PSNR are respectively the SSIM evaluation index value and the PSNR evaluation index value of the traffic sign image after power-law transformation, and both a and b are weight coefficients.

[0047] The calculation formulas of SSIM and PSNR are as follows:

[0048]

[0049] Wherein, x and y respectively represent the pixel values of the images before and after power-law transformation, μ x and μ y respectively represent the means of x and y, σ xy is the covariance of x and y, σ x 2 and σ y 2 respectively represent the variances of x and y, and c1 and c2 are constants to avoid division by zero.

[0050]

[0051] Wherein, x and y respectively represent the pixel values of the images before and after power-law transformation, MSE represents the mean square error of x and y, H and W are respectively the height and width of the image; n represents the maximum pixel value of the transformed image.

[0052] The basic idea of PSNR is to measure the image quality by comparing the mean square error between the images before and after transformation; while SSIM aims to measure the structural similarity between two images.

[0053] In the objective function, the values of the weight parameters satisfy: a + b = 1, a, b ∈ (0, 1). Among them, if the difference between the brightness mean and the saturation mean of all pixel points is greater than the set threshold, a > b, otherwise a ≤ b. Through research, it is found that the greater the difference between the brightness mean and the saturation mean of pixel points, the lower the corresponding visibility. By such a setting, when the visibility is extremely low, the weight of SSIM can be increased, and more attention is paid to the detail (completeness) of the structural information in the traffic sign image. When the visibility is relatively high, the weight of PSNR is increased, and more attention is paid to the similarity of the overall image, so that the overall similarity between the transformed image and the clear image is higher, thereby obtaining a better enhancement effect.

[0054] In this embodiment, both the clear image and the low-visibility image in the training samples are obtained by camera shooting.

[0055] As a preference, a genetic algorithm is used to optimize the weights and biases of the multi-layer perceptron model to obtain an optimized multi-layer perceptron model. The specific steps are as follows:

[0056] s1. Initialize the weights and biases of each layer of the multilayer perceptron model to obtain multiple arrays of weights and biases of each layer as the initial population. Each array corresponds to a multilayer perceptron model.

[0057] s2. Calculate the fitness of each individual in the current population;

[0058] f i =aSSIM i +bPSNR i ;

[0059] Among them, f i Represents the fitness of the i-th individual in the population, SSIM i Indicates the SSIM evaluation index value corresponding to the i-th individual in the population, PSNR i Represents the PSNR evaluation index value corresponding to the i-th individual in the population.

[0060] In this embodiment, each iteration process adopts a batch training method, that is, each time a training sample is used for training, and then all the output results of each batch are applied to image enhancement (transformation). After calculating the SSIM and PSNR of all the output results, the average SSIM of the batch is taken as the SSIM. i , take the average PSNR of the batch as PSNR i This setting can minimize errors in the training process.

[0061] s3. Select multiple individuals with large fitness in the population as parents, and use the parents to generate the next generation population through crossover and mutation.

[0062] s4. Repeat s2-s3 until the maximum number of iterations is reached; the array corresponding to the individual with the largest fitness is selected as the weight and bias of the multilayer perceptron model to obtain the optimized multilayer perceptron model.

[0063] 3. Use the optimized multi-layer perceptron model to output the optimal gamma factor of power-law transformation.

[0064] During vehicle driving, after obtaining a traffic sign image in real time through an on-board camera, the mean brightness and saturation values of all pixels in the traffic sign image are calculated and input into an optimized multi-layer perceptron model; the optimized multi-layer perceptron model outputs an optimal gamma factor for performing a power-law transformation on the traffic sign image.

[0065] Fourth, using the optimal gamma factor to perform power law transformation on the traffic sign image to obtain an enhanced traffic sign image.

[0066] 5. Perform target detection on the enhanced traffic sign image based on the YOLOv8 network model to identify the traffic signs in the image.

[0067] YOLOv8 is an advanced algorithm in the current YOLO series, adopting a number of new strategies. It uses the C2f module in the Backbone part, combines the "anchor-free + Decoupled-head" technology in the Head part, and the loss function combines the classification BCE Loss, regression CIOU Loss, and DFL Loss. In addition, the label assignment strategy changes from static matching to dynamic label matching. It uses three detectors of different sizes to select and detect the image content, and outputs prediction results of three sizes through reparameterization.

[0068] For the problems that the useful information in traffic sign images is relatively concentrated, there are many similar parts in some traffic sign images, which easily lead to misdetection, and the requirement for the recognition speed of traffic signs by autonomous vehicles is high. The present invention improves the YOLOv8 model and performs object detection based on the improved YOLOv8 model.

[0069] The method for improving the YOLOv8 model is as follows: Replace the 2nd to 4th C2f modules in the Backbone part with Ghost modules; in the Neck part, add CBAM modules after each C2f module and before each Upsample structure. The improved YOLOv8 model is as Figure 1 shown.

[0070] The YOLOv8 model with Ghost modules effectively reduces the overall parameters of the model and compresses the model volume. Ghost mainly consists of two steps: First, halve the feature map through conventional convolution operations, and then generate the other half of the feature map with more similar features through linear transformation to form a Ghost feature map. Finally, these two parts of the feature maps are concatenated and output. This method not only improves the convolution efficiency but also compresses the number of parameters and the amount of computation.

[0071] The CBAM module not only pays attention to channel information but also attaches importance to spatial information. First, it performs pooling, convolution, and normalization operations through the input channel module, then uses the Sigmoid activation function to output the redefined channel attention weights. Then, the feature map enters the spatial attention module through global average pooling (GAP) and global maximum pooling (GMP) operations, and is processed by a multi-layer perceptron to obtain spatial attention weights. Finally, the Sigmoid function is used to multiply these attention weights by the original channel weights to achieve weight recalibration, and the feature map is calibrated through the spatial attention module. Fusing the attention mechanism CBAM in the neck network and enhancing the attention to the target can solve the problems of missed detection and misdetection.

[0072] Experimental example

[0073] Ten different types of traffic sign images are used as test samples. Each traffic sign is photographed by an in-vehicle camera under low visibility conditions such as rain, snow, haze, etc. The number of each traffic sign image is 100, that is, there are a total of 1000 test samples.

[0074] The experiment is carried out in the Windows 11 system environment. The hardware device is an Intel i5-12400F processor, and the software environment is Python 3.8. The PyTorch 1.11.0 framework is used.

[0075] After the optimized multi-layer perceptron model outputs the gamma factor with a power-law change, this gamma factor is used to perform a power-law transformation on the traffic sign image. Then, the improved YOLOv8 model is used to perform object detection on the transformed traffic sign image, and the final detection results are output.

[0076] At the same time, a fixed gamma factor and the unimproved YOLOv8 model are used as the control group. The experimental results are shown in Table 1.

[0077] Recognition method Detection accuracy Fixed gamma factor + Improved YOLOv8 model 89.6% Optimized gamma factor + Unimproved YOLOv8 model 94.2% Optimized gamma factor + Improved YOLOv8 model 97.8%

[0078] The experimental results show that the accuracy of detecting traffic sign images using the method of the present invention (optimized gamma factor + improved YOLOv8 model) has been greatly improved compared to the detection method with a fixed gamma factor; at the same time, using the improved YOLOv8 model can further improve the accuracy of the detection results compared to the unimproved YOLOv8 model. By comparing the detection speeds of the optimized gamma factor + unimproved YOLOv8 model and the optimized gamma factor + improved YOLOv8 model, it is found that the average detection speed of the improved YOLOv8 model has increased by 5.8% compared to the unimproved YOLOv8 model. This fully shows that the present invention using the optimized gamma factor + improved YOLOv8 model can improve the accuracy of traffic sign recognition and enhance the recognition efficiency of traffic sign recognition.

[0079] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. A traffic sign recognition method for autonomous driving, characterized in that It includes the following steps: Step 1: Use the average brightness value and average saturation value of all pixel points in the traffic sign image as input parameters, and the gamma factor for power-law transformation of the traffic sign image as the output parameter to construct a multi-layer perceptron model; Step 2: Optimize the parameters of the multi-layer perceptron model with the maximum of the objective function as the optimization goal to obtain an optimized multi-layer perceptron model; Among them, the objective function is: F = aSSIM + bPSNR; In the formula, SSIM and PSNR are respectively the SSIM evaluation index value and PSNR evaluation index value of the traffic sign image after power-law transformation, and both a and b are weight coefficients; Step 3: During the vehicle driving process, after obtaining the traffic sign image in real time through the in-vehicle camera, calculate the average brightness value and average saturation value of all pixel points in the traffic sign image, and input them into the optimized multi-layer perceptron model; the optimized multi-layer perceptron model outputs the optimal gamma factor for power-law transformation of the traffic sign image; Step 4: Use the optimal gamma factor to perform power-law transformation on the traffic sign image to obtain an enhanced traffic sign image; Step 5: Based on the YOLOv8 network model, perform object detection on the enhanced traffic sign image to identify the traffic signs in the image.

2. The traffic sign recognition method for autonomous driving according to claim 1, wherein In Step 2, the genetic algorithm is used to optimize the weights and biases of the multi-layer perceptron model, including the following steps: Step 1: Initialize the weights and biases in each layer of the multi-layer perceptron model to obtain multiple arrays composed of the weights and biases of each layer as the initial population; Among them, each array corresponds to a multi-layer perceptron model; Step 2: Calculate the fitness of each individual in the current population; f i = aSSIM i + bPSNR i ; Among them, f i represents the fitness of the i-th individual in the population, and SSIM i represents the SSIM evaluation index value corresponding to the i-th individual in the population, and PSNR i represents the PSNR evaluation index value corresponding to the i-th individual in the population; Step 3: Select the individuals with large fitness in the population as the parents, and use the parents to generate the next generation population through crossover and mutation; Step 4: Loop through Steps 2 - 3 until the maximum number of iterations is reached; select the array corresponding to the individual with the largest fitness as the weights and biases of the multi-layer perceptron model to obtain an optimized multi-layer perceptron model.

3. The traffic sign recognition method for autonomous driving according to claim 2, characterized in that The hidden layer of the multi-layer perceptron model is 2 layers.

4. The traffic sign recognition method for autonomous driving according to claim 3, wherein The number of hidden layer nodes is determined according to the following formula: Among them, N hid is the number of hidden layer nodes; N in represents the number of input layer nodes; N out represents the number of output layer nodes; m represents any integer between 2 and 6, and [] represents rounding up.

5. The traffic sign recognition method for autonomous driving according to claim 3 or 4, characterized in that, In the objective function, the values of the weight parameters satisfy: a + b = 1, a, b ∈ (0, 1); Among them, if the difference between the average brightness value and average saturation value of all pixel points is greater than the set threshold, a > b, otherwise a ≤ b.

6. The traffic sign recognition method for autonomous driving according to claim 5, characterized in that In Step 5, it also includes improving the YOLOv8 model and performing object detection based on the improved YOLOv8 model; The method for improving the YOLOv8 model is: Replace the 2nd to 4th C2f modules in the Backbone part with Ghost modules; in the Neck part, add CBAM modules after each C2f module and before each Upsample structure.