Noise image processing method based on improved YOLOv8 model and application thereof

By improving the YOLOv8 model, adding SCDown, SEAM and SADetect modules, combining dark channel priors and ACE image enhancement technology, the problem of poor image quality in bad weather is solved, and the target detection accuracy and safety of driverless cars are improved.

CN120235779APending Publication Date: 2025-07-01WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411682754.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-05
Filing Date
2024-11-22
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing single-stage object detection algorithm has poor image quality after processing input images in harsh weather environments, resulting in driverless cars being unable to accurately identify vehicles or obstacles in front, reducing the safety and efficiency of driverless cars.

Method used

Build an improved YOLOv8 model, add SCDown module to improve image detail characteristics accuracy, SEAM attention mechanism module improves image recognition accuracy in occlusion, and introduces SADetect module to improve information capture focus, and combines dark channel prior algorithm and ACE image enhancement technology for denoising and enhancement processing.

Benefits of technology

In severe weather environments, the image quality and the accuracy of target detection are significantly improved, and the decision-making ability of driverless cars in complex environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235779A_ABST
    Figure CN120235779A_ABST
Patent Text Reader

Abstract

The invention provides a noise image processing method based on an improved YOLOv8 model, and relates to the technical field of image recognition. The method comprises the following steps: firstly, performing denoising enhancement processing on a noise image in a severe weather environment; an improved YOLOv8 model is constructed, and an SCDown module for improving the accuracy of captured image detail features, an SEAM attention mechanism module for improving the image recognition precision under the shielding condition and an SADeect module for improving information capture concentration are added on the basis of the YOLOv8 model; the improved YOLOv8 model is trained, and a trained improved YOLOv8 model is obtained; and inputting the image subjected to denoising enhancement processing into the trained improved YOLOv8 model to obtain a clear image. According to the invention, the quality of the image after the input image is processed under the environment noise is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and more particularly relates to a method for processing noise images based on an improved YOLOv8 model and its application. Background Art

[0002] With the continuous progress of technology, driverless technology has received extensive attention from the industrial and academic communities. Driverless cars perceive the surrounding environment through sensors and cameras. However, in harsh weather conditions, due to poor visibility, it is difficult for sensors to accurately identify the vehicles or obstacles ahead, resulting in the inability of driverless cars to make correct decisions and execute actions. Moreover, environmental factors such as reduced visibility lead to degraded and blurred images, object occlusion, and a low proportion of small target pixels, which also reduce the accuracy of real-time detection, weaken the overall performance of visual algorithms, and thus pose a great risk to vehicle driving.

[0003] Object detection is an important part of driverless technology. It ensures the safety and efficiency of autonomous driving by identifying and locating target objects in images or videos. Existing vehicle object detection algorithms are divided into traditional algorithms and deep learning algorithms. Traditional algorithms, due to relying on complex features designed manually, cannot fully and accurately represent the targets in images, resulting in low detection accuracy and poor real-time accuracy. With the rapid development of image processing and object detection technologies, deep learning algorithms have gradually matured. Common object detection algorithms can be divided into single-stage detection and two-stage detection according to the execution process of the algorithms. Two-stage detection methods such as R-CNN, Faster R-CNN, and DETR are based on the candidate box strategy, integrating feature extraction, region extraction, and classifiers in one network, which greatly improves the comprehensive performance, but the detection speed is slow. Another is the single-stage detection method, such as YOLO, SSD and other algorithms. The single-stage detection method is based on the regression method and directly predicts the regression bounding box and class probability from the output image using the convolutional neural network structure, which greatly improves the detection speed and can classify and detect objects in real time. However, existing single-stage object detection algorithms still face the problem of poor image quality after processing the input images under environmental noise. Summary of the Invention

[0004] In order to solve the problem of poor image quality after processing the input images under environmental noise by current single-stage object detection algorithms, the present invention proposes a method for processing noise images based on an improved YOLOv8 model to improve the image quality.

[0005] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0006] S1: Denoise and enhance the noise images in harsh weather conditions;

[0007] S2: Construct an improved YOLOv8 model by adding the SCDown module for improving the accuracy of capturing detailed features of images, the SEAM attention mechanism module for enhancing the image recognition accuracy in case of occlusion, and the SADetect module for improving the focus of information capture on the basis of the YOLOv8 model;

[0008] S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;

[0009] S4: Input the image after denoising and enhancement processing in step S1 into the trained improved YOLOv8 model to obtain a clear image.

[0010] Furthermore, the denoising and enhancement processing operation in S1 includes: using the dark channel prior algorithm to process the noisy image in a bad weather environment to obtain the image dark channel, and the expression of the image dark channel is:

[0011]

[0012] In the formula, x represents the pixel point of the noisy image in a bad weather environment, Ω(x) represents the local area centered on x, y represents any pixel in the local area, R, G, and B represent three color channels, and J c represents a certain color channel value among the R, G, and B channels of the noisy image in a bad weather environment, and J dark represents the dark primary color channel value of the noisy image in a bad weather environment;

[0013] Estimate the transmittance t(x, y) and the atmospheric light value A of the noisy image in a bad weather environment by the dark channel prior algorithm, select the pixel points of the top 0.1% brightest part in the image dark channel of the noisy image in a bad weather environment, and use the area corresponding to the pixel points in the noisy image in a bad weather environment as the target area, and define the maximum pixel value in the target area as the atmospheric light value A; introduce the environmental noise imaging model, and the expression is:

[0014] I(x,y)=J(x,y)t(x,y)+A(1 - t(x,y))

[0015] In the formula, (x, y) represents the image pixel coordinate value, I(x, y) represents the noisy image in a bad weather environment, J(x, y) represents the image after denoising, and J c represents a certain color channel value of the noisy image in a bad weather environment;

[0016] Perform two minimization operations on the environmental noise imaging model, and the expressions are:

[0017]

[0018] Wherein, t(x) represents the transmittance, A represents the atmospheric light value and is a constant;

[0019] Introduce a buffer factor ω to retain a small amount of noise, and let J dark tend to 0 to obtain the transmittance formula:

[0020]

[0021] Wherein, t0 represents the lower limit of the transmittance;

[0022] Substitute the transmittance and the atmospheric light value into the environmental noise imaging model to obtain the image denoising formula:

[0023]

[0024] Wherein, t0 represents the lower limit of the transmittance, A represents the maximum pixel value in the target area, t(x) represents the transmittance, I(x) represents the noisy image, and t(x) represents the transmittance formula.

[0025] Furthermore, the denoising and enhancement processing operation described in S1 further includes: processing the noisy image in a bad weather environment by using the ACE image enhancement technology;

[0026] The ACE image enhancement technology includes: region adaptive filtering for color and spatial domain adjustment and global reconstruction stretching for improving the brightness of the noisy image in a bad weather environment;

[0027] Apply region adaptive filtering for color and spatial domain adjustment to complete the color difference correction of the noisy image in a bad weather environment and obtain the spatially reconstructed image. The calculation expression of the region adaptive filtering is:

[0028]

[0029] Wherein, Subset represents the pixel set of the noisy image in a bad weather environment, Y c represents the spatially reconstructed image, j represents the point in the adjacent area of the pixel point x of the noisy image in a bad weather environment, I c (x)-I c (j) represents the gray difference between the two pixel points x and j, d(x,j) represents the Euclidean distance, and G represents the brightness representation function;

[0030] The expression of the brightness representation function is:

[0031] Wherein, α represents the adjustment parameter of the color saturation;

[0032] Perform single-channel hue reorganization stretching on the noisy images in harsh weather environments and extend it to the three channels of the RGB color space; finally, stretch and map the intermediate quantity obtained in the regional adaptive filtering to the interval [0, 255]. The global reorganization stretching expression is:

[0033]

[0034] In the formula, L(x) represents global reorganization stretching, [min(Y), max(Y)] represents the domain of L(x), and Y c (x) represents regional adaptive filtering.

[0035] Furthermore, the improved YOLOv8 model consists of three parts: the Backbone main network, the Neck neck network, and the Head network;

[0036] The Backbone main network includes: a first Conv module, a first SCDowm module, a first C2f module, a second SCDowm module, a second C2f module, a third SCDowm module, a third C2f module, a fourth SCDowm module, a fourth C2f module, and a first SPPF module connected in sequence;

[0037] The Neck neck network includes: a first Upsample module, a first Concat module, a fifth C2f module, a second Upsample module, a second Concat module, a sixth C2f module, a first SEAM attention mechanism module, a fifth SCDowm module, a third Concat module, a seventh C2f module, a second SEAM attention mechanism module, a sixth SCDowm module, a fourth Concat module, an eighth C2f module, and a third SEAM attention mechanism module connected in sequence; the output end of the second C2f module is connected to the second Concat module, the output end of the third C2f module is connected to the first Concat module, the output end of the first SPPF module is respectively connected to the first Upsample module and the fourth Concat module, the output end of the fifth C2f module is connected to the third Concat module, and the output end of the sixth C2f module is connected to the first SEAM attention mechanism module;

[0038] The Head network includes a first SADetect module, a second SADetect module, and a third SADetect module; the output end of the first SEAM attention mechanism module is connected to the first SADetect module, the output end of the second SEAM attention mechanism module is connected to the second SADetect module, and the output end of the third SEAM attention mechanism module is connected to the third SADetect module.

[0039] Further, the SCDown module includes: a first pointwise convolution module and a first depth convolution module connected in sequence.

[0040] Further, the SEAM attention mechanism module includes: a first CSMN module, a second CSMN module, a third CSMN module, a first average pooling module, a first fully connected network module, a second fully connected network module, and a first exponential normalization module; the first CSMN module, the second CSMN module, and the third CSMN module are in parallel, and the output ends of the first CSMN module, the second CSMN module, and the third CSMN module are concatenated and then connected to the first average pooling module, and the first average pooling module, the first fully connected network module, the second fully connected network module, and the first exponential normalization module are connected in sequence.

[0041] Further, any one of the first CSMN module, the second CSMN module, and the third CSMN module includes: a first depthwise separable convolution module, a first GELU module, a first batch normalization module, a second depth convolution module, a second GELU module, a second batch normalization module, a second depth convolution module, a third GELU module, and a third batch normalization module connected in sequence; the output end of the first batch normalization module and the output end of the second batch normalization module are concatenated and then connected to the second depth convolution module.

[0042] Further, the SADetect module is a Self-Attention self-attention mechanism module;

[0043] The Self-Attention self-attention mechanism module includes: a second Conv module, a first MHSA module, and a third Conv module connected in sequence, a fourth Conv module, a first BatchNorm module, and a first ReLU module connected in sequence; the output ends of the third Conv module and the first BatchNorm module are concatenated and then connected to the first ReLU module.

[0044] Further, during the training process of the improved YOLOv8 model, a linear interval mapping mechanism is introduced, and the loss function of the improved YOLOv8 model is constructed as the Focaler-IoU loss function.

[0045] During the construction of the Focaler-IoU loss function, IoU is the loss function based on the YOLOv8 model, and the IoU loss is reconstructed, and the expression is:

[0046] L Focaler-IoU = 1 - IoU focaler

[0047] Wherein, d represents the lower threshold value, and u represents the upper threshold value;

[0048] The expression of the Focaler-CIoU loss function is:

[0049] L Focaler-CIoU = L CIoU + IoU - IoU focaler

[0050] Wherein, IoU focaler represents the reconstructed IoU loss.

[0051] The present invention also provides an application of a noise image processing method based on an improved YOLOv8 model. The image processing method is applied to vehicle target detection in complex weather, and the detection target is a vehicle.

[0052] Compared with the prior art, the beneficial effects of this method are:

[0053] The present invention proposes a noise image processing method based on an improved YOLOv8 model and its application. First, denoising and enhancement processing is performed on the noise image in a harsh weather environment to correct the pixel values of the image. An improved YOLOv8 model is constructed. On the basis of the YOLOv8 model, an SCDonw module is added to enrich the downsampling information; an SEAM attention mechanism module is added to improve the image recognition accuracy in case of occlusion; an SADetect module is added to improve the ability to integrate local information while reducing the computational complexity. The present invention generally adopts an improved YOLOv8 model to improve the quality of the image after processing the input image under environmental noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 represents the flow chart of the noise image processing method based on the improved YOLOv8 model proposed in the embodiment of the present invention;

[0055] Figure 2 represents the structural diagram of the improved YOLOv8 model proposed in the embodiment of the present invention;

[0056] Figure 3 represents the structural schematic diagram of the SCDown module proposed in the embodiment of the present invention;

[0057] Figure 4 represents the structural schematic diagram of the SEAM attention mechanism module proposed in the embodiment of the present invention;

[0058] Figure 5 represents the schematic diagram of the CSMM module proposed in the embodiment of the present invention;

[0059] Figure 6A schematic diagram showing a Self-Attention mechanism module proposed in an embodiment of the present invention;

[0060] Figure 7 An example diagram showing a noise image in a severe weather environment used in an embodiment of the present invention;

[0061] Figure 8 It represents an image processed by the improved YOLOv8 model proposed in the embodiment of the present invention. DETAILED DESCRIPTION

[0062] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0063] In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the actual size;

[0064] It is understandable to those skilled in the art that descriptions of certain well-known contents in the drawings may be omitted.

[0065] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0066] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent;

[0067] Example 1

[0068] This embodiment proposes a noisy image processing method based on an improved YOLOv8 model. Figure 1 The flowchart of the method shown in the figure, the noise image processing method proposed in this embodiment generally includes the following steps:

[0069] S1: De-noise and enhance the noisy image in severe weather conditions;

[0070] S2: Build an improved YOLOv8 model. On the basis of the YOLOv8 model, add the SCDown module to improve the accuracy of capturing image detail features, the SEAM attention mechanism module to improve the image recognition accuracy under occlusion, and the SADetect module to improve the concentration of information capture;

[0071] S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;

[0072] S4: Input the image after denoising and enhancement processing in step S1 into the trained improved YOLOv8 model to obtain a clear image.

[0073] The composition of the improved YOLOv8 model is specifically described below. In this embodiment,Figure 2 The improved YOLOv8 model structure shown, the improved YOLOv8 model consists of three parts: the Backbone main network, the Neck neck network, and the Head network.

[0074] The Backbone main network includes: a first Conv module, a first SCDowm module, a first C2f module, a second SCDowm module, a second C2f module, a third SCDowm module, a third C2f module, a fourth SCDowm module, a fourth C2f module, and a first SPPF module connected in sequence;

[0075] The Neck neck network includes: a first Upsample module, a first Concat module, a fifth C2f module, a second Upsample module, a second Concat module, a sixth C2f module, a first SEAM attention mechanism module, a fifth SCDowm module, a third Concat module, a seventh C2f module, a second SEAM attention mechanism module, a sixth SCDowm module, a fourth Concat module, an eighth C2f module, and a third SEAM attention mechanism module connected in sequence; the output end of the second C2f module is connected to the second Concat module, the output end of the third C2f module is connected to the first Concat module, the output end of the first SPPF module is respectively connected to the first Upsample module and the fourth Concat module, the output end of the fifth C2f module is connected to the third Concat module, and the output end of the sixth C2f module is connected to the first SEAM attention mechanism module;

[0076] The Head network includes a first SADetect module, a second SADetect module, and a third SADetect module; the output end of the first SEAM attention mechanism module is connected to the first SADetect module, the output end of the second SEAM attention mechanism module is connected to the second SADetect module, and the output end of the third SEAM attention mechanism module is connected to the third SADetect module.

[0077] As Figure 3 shown, the SCDown module includes: a first pointwise convolution module and a first depth convolution module connected in sequence.

[0078] The image enters the Backbone network. First, it passes through the first Conv module, which consists of a Conv convolution, BN (batchnorm2d), and a silu activation function. The feature map processed by the first Conv module is input into the SCDown module and the C2f module. The SCDown module decouples the operations of spatial reduction and channel increase. First, it uses the first pointwise convolution module to increase the channel dimension from C to 2C, and then uses the first depthwise convolution module to perform spatial downsampling of the height and width (adjusted from H*W to H / 2*W / 2). The SCDown module reduces the computational cost from 9 / 2HWC^2 to 2HWC^2 + 9 / 2HWC, and the number of parameters from 18C^2 to 2C^2 + 18C, and maximally preserves information during the downsampling process, ensuring that the model can capture image detail features more accurately. The C2f module adopts an efficient aggregation network structure. Through multi-branch cross-layer connections, it enriches the gradient flow of the model, improves the utilization rate of parameters, and increases the network depth. After passing through four groups of SCDown modules and C2f modules, it enters the SPPF module to process feature maps of different pixel sizes to complete feature fusion.

[0079] After the image obtains deep feature information through the Backbone network, it enters the Neck network. The Neck network adopts the PAFPN structure, which consists of two parts: the Feature Pyramid Networks (FPN) and the Path Aggregation Network (PAN), and adds the SEAM attention mechanism module. As Figure 4 shown, the SEAM attention mechanism module includes: the first CSMN module, the second CSMN module, the third CSMN module, the first average pooling module, the first fully connected network module, the second fully connected network module, and the first exponential normalization module; the outputs of the first CSMN module, the second CSMN module, and the third CSMN module are concatenated and then connected to the first average pooling module. The first average pooling module, the first fully connected network module, the second fully connected network module, and the first exponential normalization module are connected in sequence. The output of the first exponential normalization module is multiplied by the outputs of the first CSMN module, the second CSMN module, and the third CSMN module. As Figure 5 shown, the CSMN module includes: the first depthwise separable convolution module, the first GELU module, the first batch normalization module, the second depthwise convolution module, the second GELU module, the second batch normalization module, the second depthwise convolution module, the third GELU module, and the third batch normalization module connected in sequence; the output of the first batch normalization module is concatenated with the output of the second batch normalization module and then connected to the second depthwise convolution module.

[0080] The input image is first input into three CSMM (Channel and Spatial Mixing Module) modules of different scales by the SEAM attention mechanism module. The CSMM module adopts the form of residual connection and the cascade of the first depthwise separable convolution module and the second depth convolution module, convolves each input channel separately, learns the important information of each channel and reduces the number of parameters, and then uses a 1*1 convolution to spatially mix the information of multiple channels and learn the relationship between different channels. After the outputs of the three CSMM modules pass through the first average pooling module, they are input into the first fully connected network module and the second fully connected network module to fuse the information of each channel, further strengthening the information connection between all channels. Finally, the attention information is output through the first exponential normalization module. The output attention information and the original input information are finally multiplied by a multiplier to perform detailed processing on the feature information in two dimensions of channel and space, enhancing the network's ability to handle object occlusion and target background interference problems while maintaining computational efficiency.

[0081] The detection results of the three SEAM attention mechanism modules in the Neck neck network are used as inputs to enter the Head network. In the Head network, a decoupled head structure is adopted to process the classification and regression task branches separately. The present invention adds a SADetect module on the basis of the decoupled head. The SADetect module is a Self-Attention self-attention mechanism module. As Figure 6 shown, the Self-Attention self-attention mechanism module includes: a second Conv module, a first MHSA module, and a third Conv module connected in sequence, a fourth Conv module, a first BatchNorm module, and a first ReLU module connected in sequence; the outputs of the third Conv module and the first BatchNorm module are spliced and then connected to the first ReLU module.

[0082] For the Self-Attention self-attention mechanism module, the dimension is first reduced by the 1*1 second Conv module to reduce the computational load of self-attention, and then it is input into the first MHSA module to obtain the information correlation between features. The number of channels is restored by the 1*1 third Conv module, and after batch normalization, the self-attention module is added to the initial input to form a residual connection, and the image features with self-attention are output. The second Conv module uses the residual connection idea. The image information is first downsampled and feature transformed by the 1*1 Conv module, and then upsampled to restore the spatial features and connected to the initial input, enabling the network to integrate local information under the self-attention mechanism and reduce the computational complexity.

[0083] During the training process of the improved YOLOv8 model, a linear interval mapping mechanism is introduced, and the loss function of the improved YOLOv8 model is constructed as the Focaler-IoU loss function.

[0084] In the process of constructing the Focaler-IoU loss function, IoU is the loss function based on the YOLOv8 model. The IoU loss is reconstructed, and the expression is:

[0085]

[0086] L Focaler-IoU = 1 - IoU focaler

[0087] In the formula, d represents the lower threshold, and u represents the upper threshold;

[0088] The expression of the Focaler-CIoU loss function is:

[0089] L Focaler-CIoU = L CIoU + IoU - IoU focaler

[0090] In the formula, IoU focaler represents the reconstructed IoU loss.

[0091] Example 2

[0092] This example describes the process of denoising and enhancing noisy images in a harsh weather environment.

[0093] The noisy images selected in this example come from the RTTS dataset. The RTTS dataset (RTTS Hazy Dataset) is a road traffic image dataset in a foggy environment. This dataset contains real-world foggy images with diverse scenes and lighting conditions, high image resolution, and can provide rich detailed information. The RTTS dataset contains 4,322 natural foggy images with 5 labeled object classes, namely people, bicycles, cars, buses, and motorcycles.

[0094] An example diagram of the noisy images selected in this example under a harsh weather environment is as Figure 7 shown.

[0095] First, use the dark channel prior algorithm to process the noisy images in a harsh weather environment to obtain the image dark channel. The expression of the image dark channel is:

[0096]

[0097] In the formula, x represents the pixel points of the noisy images in a harsh weather environment, Ω(x) represents the local area centered on x, y represents any pixel in the local area, R, G, B represent the three color channels, and J c represents a certain color channel value among the R, G, and B channels of the noisy images in a harsh weather environment, Jdark Represents the dark channel prior value of the noise image in a bad weather environment;

[0098] The transmission rate t(x, y) and the atmospheric light value A of the noise image in a bad weather environment are estimated by the dark channel prior algorithm. The pixel points of the top 0.1% brightest part in the image dark channel of the noise image in a bad weather environment are selected, and the corresponding area of the pixel points in the noise image in a bad weather environment is used as the target area. The maximum pixel value in the target area is defined as the atmospheric light value A; The environmental noise imaging model is introduced, and the expression is:

[0099] I(x,y) = J(x,y)t(x,y) + A(1 - t(x,y))

[0100] In the formula, (x,y) represents the image pixel coordinate value, I(x,y) represents the noise image in a bad weather environment, J(x,y) represents the denoised image, and J c Represents a certain color channel value of the noise image in a bad weather environment;

[0101] Two minimization operations are performed on the environmental noise imaging model, and the expression is:

[0102]

[0103] In the formula, t(x) represents the transmission rate, A represents the atmospheric light value and is a constant;

[0104] The buffer factor ω is introduced to retain a small amount of noise, and let J dark Tends to 0 to obtain the transmission rate formula:

[0105]

[0106] In the formula, t0 represents the lower limit of the transmission rate;

[0107] The transmission rate and the atmospheric light value are substituted into the environmental noise imaging model to obtain the image denoising formula:

[0108]

[0109] In the formula, t0 represents the lower limit of the transmission rate, A represents the maximum pixel value in the target area, t(x) represents the transmission rate, I(x) represents the noise image, and t(x) represents the transmission rate formula.

[0110] Next, to avoid problems such as halos and color casts in the high-brightness area of the denoising module, the ACE image enhancement technology is used to process the noise image. The regional adaptive filtering is used for color and spatial domain adjustment to complete the chromatic aberration correction of the image, and the spatial domain reconstructed image is obtained. The calculation expression of the regional adaptive filtering is:

[0111]

[0112] In the formula, Subset represents the pixel set of the noise image in the bad weather environment, Y c represents the airspace reconstructed image, j represents the point in the adjacent area of the pixel point x of the noise image in the bad weather environment, I c (x) - I c (j) represents the gray difference between the two pixel points x and j, d(x, j) represents the Euclidean distance, and G represents the brightness performance function;

[0113] The expression of the brightness performance function is:

[0114]

[0115] In the formula, α represents the adjustment parameter of the color saturation;

[0116] Perform single-channel hue reformation stretching on the noise image in the bad weather environment, and extend it to the three channels of the RGB color space; finally, stretch and map the intermediate quantity obtained in the regional adaptive filtering to the interval [0, 255], and the global reformation stretching expression is:

[0117]

[0118] In the formula, L(x) represents the global reformation stretching, [min(Y), max(Y)] represents the domain of L(x), and Y c (x) represents the regional adaptive filtering.

[0119] Example 3

[0120] In this example, the improved YOLOv8 model is trained using the training set to obtain the trained improved YOLOv8 model.

[0121] When training the DMC-YOLO model in this example, the basic platform and parameters are as follows:

[0122] Build the Pytorch framework, set the learning rate to 0.01, the weight decay coefficient to 0.0005, the image size to 640×640, the batch size to 8, and the training period to 300. Input the images in the dataset after denoising and enhancement processing into the improved YOLOv8 model for processing. As Figure 8 shown, the image processed by the improved YOLOv8 model. Use a rectangular box to mark the target vehicle in the image. Then, compare the results of the improved YOLOv8 model with the results of the YOLOv8 model. The results show that the results of the improved YOLOv8 model have less influence of water vapor on the image and improve the quality of the image.

[0123] Example 4

[0124] This embodiment provides an application of a noise image processing method based on an improved YOLOv8 model. The image processing method is applied to vehicle target detection in complex weather, and the detection target is a vehicle.

Claims

1. A noisy image processing method based on an improved YOLOv8 model, characterized in that: The following steps are involved: S1: De-noise and enhance the noisy image in severe weather conditions; S2: Build an improved YOLOv8 model. On the basis of the YOLOv8 model, add the SCDown module to improve the accuracy of capturing image detail features, the SEAM attention mechanism module to improve the image recognition accuracy under occlusion, and the SADetect module to improve the concentration of information capture; S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model; S4: Input the image after denoising and enhancement processing in step S1 into the trained improved YOLOv8 model to obtain a clear image.

2. The noisy image processing method based on the improved YOLOv8 model according to claim 1, characterized in that: The denoising and enhancement processing operation described in S1 includes: using a dark channel prior algorithm to process the noisy image in a bad weather environment to obtain an image dark channel, and the image dark channel expression is: In the formula, x represents the pixel of the noisy image under severe weather conditions, Ω(x) represents the local area centered on x, y represents any pixel in the local area, R, G, and B represent three color channels, and J c Represents the color channel value of one of the three channels R, G, and B of the noise image under severe weather conditions, J dark Represents the dark primary channel value of the noise image under severe weather conditions; The transmittance t(x, y) and the atmospheric light value A of the noise image under severe weather conditions are estimated by the dark channel prior algorithm. The top 0.1% brightest pixel points in the dark channel of the noise image under severe weather conditions are selected, and the area corresponding to the pixel point in the noise image under severe weather conditions is taken as the target area. The maximum pixel value in the target area is defined as the atmospheric light value A. The environmental noise imaging model is introduced, and the expression is: I(x,y)=J(x,y)t(x,y)+A(1-t(x,y)) In the formula, (x, y) represents the pixel coordinate value of the image, I(x, y) represents the noisy image under bad weather conditions, J(x, y) represents the denoised image, and J c Represents a color channel value of a noise image under severe weather conditions; The environmental noise imaging model is minimized twice, and the expression is: In the formula, t(x) represents the transmittance, A represents the atmospheric light value and is a constant; Introduce the buffer factor ω to retain a small amount of noise, and let J dark tends to 0, and the transmittance formula is obtained: Where t0 represents the lower limit of transmittance; Substitute the transmittance and atmospheric light values ​​into the environmental noise imaging model to obtain the image denoising formula: Where t0 represents the lower limit of transmittance, A represents the maximum pixel value in the target area, t(x) represents the transmittance, I(x) represents the noise image, and t(x) represents the transmittance formula.

3. The noisy image processing method based on the improved YOLOv8 model according to claim 1, characterized in that: The denoising and enhancement processing operation described in S1 also includes: using ACE image enhancement technology to process the noisy image in the severe weather environment; The ACE image enhancement technology includes: regional adaptive filtering for color and spatial domain adjustment and global reshaping and stretching for improving the brightness of noisy images in severe weather conditions; Regional adaptive filtering is used to adjust the color and spatial domain, complete the chromatic aberration correction of the noise image under severe weather conditions, and obtain the spatial domain reconstructed image. The calculation expression of regional adaptive filtering is: Where Subset represents the pixel set of the noisy image under severe weather conditions, Y c represents the spatial domain reconstructed image, j represents the point in the vicinity of the noise image pixel x under severe weather conditions, and I c (x)-I c (j) represents the grayscale difference between two pixels x and j, d(x,j) represents the Euclidean distance, and G represents the brightness performance function; The expression of the brightness performance function is: In the formula, α represents the adjustment parameter of color saturation; The single-channel tone reshaping and stretching of the noisy image under severe weather conditions is extended to the three channels of the RGB color space; finally, the intermediate amount stretching obtained in the regional adaptive filtering is mapped to the [0,255] interval, and the global reshaping and stretching expression is: Where L(x) represents the global reshape stretch, [min(Y), max(Y)] represents the domain of L(x), and Y c (x) represents region adaptive filtering.

4. The noisy image processing method based on the improved YOLOv8 model according to claim 1, characterized in that: The improved YOLOv8 model consists of three parts: Backbone network, Neck network and Head network; The Backbone network includes: a first Conv module, a first SCDowm module, a first C2f module, a second SCDowm module, a second C2f module, a third SCDowm module, a third C2f module, a fourth SCDowm module, a fourth C2f module and a first SPPF module connected in sequence; The Neck network includes: a first Upsample module, a first Concat module, a fifth C2f module, a second Upsample module, a second Concat module, a sixth C2f module, a first SEAM attention mechanism module, a fifth SCDowm module, a third Concat module, a seventh C2f module, a second SEAM attention mechanism module, a sixth SCDowm module, a fourth Concat module, an eighth C2f module and a third SEAM attention mechanism module connected in sequence; the output end of the second C2f module is connected to the second Concat module, the output end of the third C2f module is connected to the first Concat module, the output end of the first SPPF module is respectively connected to the first Upsample module and the fourth Concat module, the output end of the fifth C2f module is connected to the third Concat module, and the output end of the sixth C2f module is connected to the first SEAM attention mechanism module; The Head network includes a first SADetect module, a second SADetect module and a third SADetect module; the output end of the first SEAM attention mechanism module is connected to the first SADetect module, the output end of the second SEAM attention mechanism module is connected to the second SADetect module, and the output end of the third SEAM attention mechanism module is connected to the third SADetect module.

5. The noisy image processing method based on the improved YOLOv8 model according to claim 4, characterized in that: The SCDown module includes: a first point-by-point convolution module and a first depth convolution module connected in sequence.

6. A noisy image processing method based on an improved YOLOv8 model according to claim 4, characterized in that: The SEAM attention mechanism module includes: a first CSMN module, a second CSMN module, a third CSMN module, a first average pooling module, a first fully connected network module, a second fully connected network module and a first exponential normalization module; the first CSMN module, the second CSMN module and the third CSMN module are in parallel, the output ends of the first CSMN module, the second CSMN module and the third CSMN module are spliced ​​and connected to the first average pooling module, and the first average pooling module, the first fully connected network module, the second fully connected network module and the first exponential normalization module are connected in sequence.

7. The noisy image processing method based on the improved YOLOv8 model according to claim 6, characterized in that: Any one of the first CSMN module, the second CSMN module, and the third CSMN module includes: a first depthwise separable convolution module, a first GELU module, a first batch normalization module, a second depthwise convolution module, a second GELU module, a second batch normalization module, a second depthwise convolution module, a third GELU module, and a third batch normalization module connected in sequence; the output end of the first batch normalization module is spliced ​​with the output end of the second batch normalization module and then connected to the second depthwise convolution module.

8. The noisy image processing method based on the improved YOLOv8 model according to claim 4, characterized in that: The SADetect module is a Self-Attention mechanism module; The Self-Attention mechanism module includes: a second Conv module, a first MHSA module and a third Conv module connected in sequence, a fourth Conv module and a first BatchNorm module, and a first ReLU module connected in sequence; the output ends of the third Conv module and the first BatchNorm module are spliced ​​and connected to the first ReLU module.

9. The noisy image processing method based on the improved YOLOv8 model according to claim 1, characterized in that: In the training process of the improved YOLOv8 model, a linear interval mapping mechanism is introduced, and the loss function of the improved YOLOv8 model is constructed as the Focaler-IoU loss function. In the process of constructing the Focaler-IoU loss function, IoU is the loss function based on the YOLOv8 model. The reconstructed IoU loss is expressed as: L Focaler-IoU =1-IoU focaler In the formula, d represents the lower threshold, and u represents the upper threshold; The expression of Focaler-CIoU loss function is: L Focaler-CIoU =L CIoU +IoU-IoU focaler Where, IoU focaler represents the reconstruction IoU loss.

10. An application of the noisy image processing method based on the improved YOLOv8 model according to any one of claims 1 to 9, characterized in that: The image processing method is applied to vehicle target detection in bad weather, and the detection target is a vehicle.