High-robustness foggy day real-time target detection method and system based on defogging module

By designing the defog defog module and image sharpening processing, combined with the enhanced space pooling pyramid module, the target detection model is optimized, and the robustness and accuracy of target detection under foggy days is solved, and more efficient target detection in foggy days is achieved.

CN120339588APending Publication Date: 2025-07-18CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476289.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing object detection algorithms are difficult to effectively distinguish between target objects and image backgrounds under foggy conditions, and cannot effectively enhance semantic information hidden deep in the image, affecting the robustness and accuracy of the model.

Method used

A strong and robust real-time object detection method for fogging days is adopted based on the defogging module. By designing the defogging module and image sharpening processing, combined with the enhanced space pooling pyramid module, the object detection model is optimized, and feature extraction and edge information are enhanced.

Benefits of technology

It improves the accuracy and robustness of target detection in foggy conditions, solves the problem of haze blocking target objects and edges unclear, and improves the detection performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339588A_ABST
    Figure CN120339588A_ABST
Patent Text Reader

Abstract

The invention discloses a high-robustness foggy day real-time target detection method and system based on a defogging module, and relates to the field of image target detection. The method comprises the following steps: preprocessing data; a defogging module is designed to carry out defogging on a foggy image to obtain a defogged image, image sharpening is adopted to process the defogged image, the defogging module is composed of a feature mapper and a reflection emitter, jump connection is added between the mapper and the reflection emitter, and image sharpening is carried out by using a Laplace sharpening kernel to process the image; constructing a target detection model for performing target detection on the processed defogged image; model training is carried out, a loss function is designed for model optimization, the loss function comprises defogging loss of the defogging module and total loss of the target detection model, and joint optimization is carried out on the defogging module and the target detection model. According to the invention, through image defogging and sharpening processing, the problems that a target object is shielded by haze and the edge of the object is not clear under a foggy day condition are well solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image target detection, and particularly relates to a strong-robustness foggy-day real-time target detection method and system based on a dehazing module. Background Art

[0002] Target detection is an important branch in the field of computer science, which aims to identify and locate object instances of interest in images or videos. Under foggy-day conditions, there are problems such as low visibility of object instances, partial color distortion, and occlusion, which limit the performance of existing target detection algorithms. In real-world scenarios, foggy-day target detection has wide applications in fields such as autonomous driving, terrain exploration, industrial inspection, robot navigation, and video surveillance.

[0003] Currently, target detection under foggy-day conditions is mainly divided into four categories. The first category is the method based on enhancing the feature extraction ability of the model. By designing various network modules, the performance of the backbone or neck network of the model is improved. The second category is the method based on adaptive image enhancement. By designing a simple convolutional network module to predict image enhancement hyperparameters, and then training together with the target detection network. During the training process, the hyperparameters are continuously optimized. The third category is the method based on guiding the network with prior knowledge. By introducing various prior knowledge (such as the atmospheric scattering model, dark channel prior, etc.) into the detection model, the influence of haze on the model performance is eliminated, thereby improving the model performance. The fourth category is the method based on domain adaptation. By aligning features according to the feature distribution of the domain, the domain distribution difference between the training set and the test set can be reduced, thereby improving the detection accuracy of the model.

[0004] YOLOv10 is an advanced target detection algorithm, belonging to the YOLO (You Only Look Once) series. The YOLO series is famous for its efficient detection ability, and can quickly and accurately identify and locate targets in images in real-time scenarios. Based on inheriting the advantages of the previous generations of YOLO algorithms, YOLOv10 further optimizes the performance and efficiency, and has excellent performance. It adopts an end-to-end detection mechanism, directly predicting the position and category of the target from the image, without multi-scale detection or sliding windows. This makes the detection speed very fast and has strong real-time performance. In addition, it cancels the non-maximum suppression (NMS) method, further accelerating the inference speed of the model. However, YOLOv10 still has the following deficiencies in the foggy-day target detection task:

[0005] 1) Although YOLOv10 has improved the detection accuracy and inference speed of the model in ordinary scenarios based on YOLOv8. However, under foggy conditions, the quality of the image deteriorates, the edges of the targets are unclear, and it is difficult to extract features. Haze will obscure some features of the target object, and there will be certain distortions in the shape and color of the object, making the model unable to effectively distinguish the target object from the image background and reducing the accuracy of the model. These problems have not been effectively solved in YOLOv10.

[0006] 2) In the backbone network of YOLOv10, the SPPF module enhances the multi-scale representation ability of features through Spatial Pyramid Pooling (SPP). It aims to perform max pooling on the image features extracted by the backbone network using a 5×5 window multiple times to enhance the semantic information of the features. However, under foggy conditions, haze will hide the semantic information of the target object deep in the image, and the environment of foggy images is relatively complex, and the shapes of various objects are also very different. Simply using a window of the same size for pooling cannot effectively enhance the semantic information hidden deep in the image, seriously affecting the robustness of this method.

[0007] Therefore, the present invention proposes a strong-robustness foggy-day real-time target detection method and system based on a dehazing module. By adopting image dehazing and sharpening processing, it well solves problems such as haze obscuring the target object and unclear object edges under foggy conditions. Summary of the Invention

[0008] The purpose of the present invention is to provide a strong-robustness foggy-day real-time target detection method and system based on a dehazing module to solve problems in the existing target detection methods, such as the inability to effectively distinguish the target object from the image background and reduce the accuracy, and the inability to effectively enhance the semantic information hidden deep in the image and affect the robustness.

[0009] To achieve the above purpose, the technical solutions adopted by the present invention are as follows:

[0010] The present invention proposes a strong-robustness foggy-day real-time target detection method based on a dehazing module in the first aspect, including the following steps:

[0011] S1. Data preprocessing; collecting foggy images and clear images as image data, and performing image enhancement on the image data;

[0012] S2. Designing a dehazing module to dehaze the foggy image to obtain a dehazed image, and performing image sharpening processing on the dehazed image;

[0013] The defogging module consists of a mapper and a demapper. A skip connection is added between the mapper and the demapper. Both the mapper and the demapper include three two-dimensional convolution operations; image sharpening uses a Laplacian sharpening kernel to process the image;

[0014] S3. Construct a target detection model for performing target detection on the defogged image after processing; the target detection model includes a backbone network, a neck network, and a decoupled detection head;

[0015] S4. Train the model and design a loss function for model optimization; the loss function includes the defogging loss of the defogging module and the total loss of the target detection model, and jointly optimize the defogging module and the target detection model;

[0016] S5. Target detection; use the optimized target detection model based on the defogging module to perform target detection on the foggy image to be detected.

[0017] Preferably, the data preprocessing in S1 is specifically as follows: reset the collected image to a fixed size, and then increase the accuracy and robustness of the model through data augmentation operations such as equal-proportion scaling, central symmetry, horizontal flipping, random occlusion, and translation.

[0018] Preferably, the defogging module in S2 defogs the foggy image, specifically as follows:

[0019] The mapper uses three two-dimensional convolution operations to gradually expand the channel dimension of the image to 12 dimensions, obtaining F E1 , F E2 , F E3 , F E4 ; the demapper then uses three two-dimensional convolution operations to gradually compress the channel dimension of the image to 3 dimensions that conform to an RGB image, obtaining F D1 , F D2 , F D3 , F D4 ; during the decoding process, perform an addition operation on F Ei and the processed F Di-1 to obtain F Di , i = 2, 3, 4.

[0020] Furthermore, the two-dimensional convolution is a linear combination of 2DConv, BatchNorm normalization, and SiLu activation.

[0021] Preferably, the image sharpening process in S2 is specifically as follows: the Laplacian sharpening kernel and the sharpening formula are expressed as follows:

[0022]

[0023] F(x, ω) = (1 - ω)·I(x) + ω·Lap(I(x)) (ω = 0.5)

[0024] Among them, Kernel represents a 3×3 sharpening convolution kernel, F(x,ω) represents the sharpened image, I(x) represents the input image, Lap(I(x)) represents Laplacian filtering, and ω represents the sharpening weight parameter.

[0025] Preferably, the object detection model in S3 is constructed based on the YOLOv10 model and the backbone network is improved. A strengthened spatial pyramid pooling module is added to the backbone network to fuse the extracted semantic features through pooling operations of different depths and scales.

[0026] Furthermore, the processing of the strengthened spatial pyramid pooling module is as follows:

[0027] Let the input feature map be After extracting the image features through Conv, we get Then, use max pooling operations with a window size of 7×7 and multiple 5×5 to enhance the feature information to obtain M3, M4, M5. Then, use the Concat operation to splice the information of M2, M3, M5, M6 on the channel dimension dim = 1 to form Finally, fuse the semantic information of different-sized target objects through a Conv operation to obtain the final output The process is expressed as follows:

[0028] M2 = Conv(M1)

[0029] M3 = MaxPool5×5(M2)

[0030] M5 = MaxPool5×5(MaxPool7×7(M3))

[0031] M6 = MaxPool5×5(M5)

[0032] M7 = Convat(M2, M3, M5, M6)

[0033] M8 = Conv(M7)

[0034] Among them, Conv is a linear combination of 2DConv, BatchNorm normalization, and SiLu activation.

[0035] Preferably, the object detection in S3 is specifically as follows:

[0036] Feature extraction: Use the improved backbone network to extract semantic features of different scales from the processed dehazed image;

[0037] Feature fusion: The neck network is used to post-process and fuse feature maps at multiple scales to obtain semantic features at different scales;

[0038] Information post-processing decoding and prediction: Each semantic feature is processed to obtain two tensors storing classification information and localization information respectively.

[0039] Preferably, the loss function in S4 is specifically as follows:

[0040] The defogging module uses the MSE formula to calculate the distance between the defogged image and the clear image, and takes it as the defogging loss of the defogging module;

[0041] The defogging loss and the total loss of the object detection model are expressed as follows:

[0042]

[0043] Loss=L cls +L CIoU +L DFL +L Defog

[0044] where n is the total number of pixels, represents the pixel value of the clear image at position i, represents the pixel value of the defogged image at position i; L cls represents the class loss, L CIoU and L DFL represent the localization loss, L Defog represents the defogging loss, which is calculated by MSE.

[0045] In the second aspect, the present invention proposes a strongly robust foggy day real-time object detection system based on a defogging module, including:

[0046] A defogging module, composed of a mapper and a demapper, both the mapper and the demapper include three two-dimensional convolution operations for image defogging processing;

[0047] An image sharpening module, using a Laplacian sharpening kernel for image sharpening processing;

[0048] An object detection model, including a BackBone backbone network, a Neck neck network and a Decoupled Head decoupled detection head. The BackBone backbone network is added with a strengthened spatial pooling pyramid module for image object detection.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] (1) In the foggy day real-time target detection method of the present invention, a dehazing module FMDBlock is designed, which is an innovative design of the present invention. It consists of a feature mapper and a reflection mapper, each containing three two-dimensional convolution operations. By adding a skip connection between the mapper and the reflection mapper, the effective semantic information of the original image is retained. And after dehazing is completed, the dehazed image is compared with the clear image, and the dehazing loss is calculated using the MSE formula to optimize the module.

[0051] (2) The foggy day real-time target detection method in the present invention uses Laplacian image sharpening to enhance the edge information of objects, which can solve the problem of unclear edges of target objects in foggy day images. It uses the second derivative of the calculated image to highlight the high-frequency components in the image. The present invention uses a Laplacian sharpening kernel to process the image, which can effectively solve the problem that haze will obscure objects and make their shape edges blurred under foggy conditions.

[0052] (3) The foggy day real-time target detection method in the present invention designs an enhanced spatial pyramid pooling module ESPPF. The role of ESPPF is to fuse the semantic features extracted by the backbone through pooling operations of different depths and scales; it can improve the detection ability of the model for targets of different sizes and enhance robustness. In addition, ESPPF optimizes the computational efficiency compared with the traditional SPPF, and speeds up the inference speed of the model while maintaining the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the structural block diagram of the target detection system in the present invention;

[0054] Figure 2 is the structural block diagram of the dehazing module in the present invention;

[0055] Figure 3 is the schematic diagram of image sharpening processing in the present invention;

[0056] Figure 4 is the structural block diagram of the enhanced spatial pyramid pooling module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.

[0058] Embodiment 1:

[0059] A robust real-time target detection method for foggy weather based on a defogging module can be divided into four stages: image data preprocessing, feature extraction, feature pyramid fusion and information analysis. In the data preprocessing stage, the input images of different sizes are first reset to a fixed size, and then the accuracy and robustness of the model are increased by operations such as proportional scaling, central symmetry, horizontal flipping, random occlusion and translation. In addition, the present invention proposes image defogging and sharpening processing, which well solves the problems of haze blocking target objects and unclear object edges under foggy conditions. In the feature extraction stage, the model will expand the image features from the low-dimensional space to the high-dimensional space and obtain three feature maps of different sizes. In the feature fusion stage, the model uses the feature pyramid to fuse the three feature maps of different sizes through upsampling and downsampling methods and regenerate three feature maps of different sizes to adapt to target objects of different sizes. In the information analysis stage, the feature maps obtained in the feature pyramid will be analyzed, and each feature map will be analyzed to obtain two information tensors, including positioning information and classification information respectively. The present invention adopts a target detection system to perform the above process. The target detection system includes a defogging module, a backbone network, a neck network and a decoupling detection head. The overall network structure of the proposed model is as follows: Figure 1 shown.

[0060] The method comprises at least the following steps:

[0061] Step 1: Data preprocessing.

[0062] First of all, in reality, the images collected by different devices are different. In order to facilitate the model to process the images, the input images need to be set to the same size. After the image is input into the model, it is first resized to 640×640. In order to preserve the original information of the image as much as possible and prevent the image from being distorted, the long side of the image is generally set to 640, and then the short side is scaled proportionally, and the RGB value of the part without image information is set to 0 (blackened).

[0063] Then, the scaled image is enhanced using methods such as 4×4 mosaic center symmetry (randomly selecting a point in the image as the center and moving it to the center of the image, cropping the overflow to fill the gap), horizontal flipping, random occlusion, and translation. In the model inference phase, no other image enhancement processing is required except resizing the image to simulate the actual operation of the model. These image enhancement methods can effectively increase the difficulty of the model training phase, and use simpler images for prediction in the inference phase to effectively improve the accuracy and robustness of the model.

[0064] Step 2: Dehaze and sharpen the foggy image.

[0065] The present invention proposes an image dehazing module and uses image sharpening processing, which can well process image data under foggy conditions.

[0066] During the dehazing process, the present invention expands the semantic information of the image into a high-dimensional space, and this process is called mapping. Then, the semantic features in the high-dimensional space are mapped into a low-dimensional space, and this process is called inverse mapping. In these two processes, the present invention adds skip connections to preserve some features of the original image. The image dehazing module adopts a feature mapping dehazing block (FMDBlock), which is independently constructed by the present invention, and the module structure is as Figure 2 shown.

[0067] The dehazing module FMDBlock consists of a feature mapper and an inverse mapper, each of which contains three two-dimensional convolution operations. The two-dimensional convolution is a linear combination of 2DConv, BatchNorm normalization, and SiLu activation. To preserve the effective semantic information of the original image, a skip connection is added between the mapper and the inverse mapper. The mapper uses three convolution operations to expand the channel dimension of the image to 12 dimensions. The inverse mapper then uses three convolution operations to compress the channel dimension of the image to 3 dimensions that conform to RGB images. After dehazing is completed, the dehazed image and the clear image are compared, and the dehazing loss is calculated using the MSE formula.

[0068] Let the input feature be F E1 , during the mapping process, the FMDBlock module uses a normal convolution with a kernel size of 3×3 to gradually expand the channel dimension to 12 dimensions, obtaining F E2 , F E3 , F E4 . During the inverse mapping process, a similar method is used to project the channel dimension to 3 dimensions (RGB) again, obtaining F D1 , F D2 , F D3 , F D4 , During the decoding process, to preserve the effective semantic information of the original image, the present invention performs an addition operation on F E3 and the processed F D2 to obtain F D3 . The original foggy image is processed by FMDBlock to obtain the dehazed image F D4 .

[0069] The present invention uses the Mean Squared Error (MSE) formula to calculate the distance between the dehazed image and the clear image, and takes it as the dehazing loss of the FMDBlock. MSE is a commonly used error metric method, which has the advantages of being simple and easy to use, having good mathematical properties, and being sensitive to outliers. When the pixel values of the clear image and the dehazed image at the same position are very close, the value of MSE approaches 0, which is very suitable for calculating the loss of the dehazing module. Finally, the dehazing loss is added to the loss value of the detection model, and the detection model and the FMDBlock are jointly optimized.

[0070] During the process of image sharpening, the present invention selects to use a 3×3 Laplacian sharpening kernel to process the image, and the process is as Figure 3 shown. The Laplacian sharpening kernel is proposed based on the Laplacian operator, which aims to calculate the second derivative of the image to highlight the high-frequency components (such as edges and details) in the image. High-frequency components usually correspond to the rapidly changing regions in the image, while low-frequency components correspond to the smooth regions. By enhancing the high-frequency components, the Laplacian operator can improve the clarity and visual quality of the image. Under foggy conditions, the haze will obscure objects, making their shape edges blurred. Using the Laplacian sharpening kernel to process the image can effectively solve this problem. The Laplacian sharpening kernel and the sharpening formula are expressed as follows:

[0071]

[0072] F(x,ω)=(1-ω)·I(x)+ω·Lap(I(x)) (ω=0.5)

[0073] where Kernel represents the 3×3 sharpening convolution kernel, F(x,ω) represents the sharpened image, I(x) represents the input image, Lap(I(x)) represents the Laplacian filtering, and ω represents the sharpening weight parameter.

[0074] Step 3: Feature extraction.

[0075] The BackBone in the object detection model extracts image features. The backbone network is the core part of the detection model. Extracting image features means gradually expanding the low-dimensional image features into a high-dimensional space. In this process, the image size will gradually decrease, but the dimension will gradually increase. An excellent backbone network will pay more attention to the semantic features of the target object and ignore the semantic features of the image background during the process of expanding the image from low dimension to high dimension.

[0076] The preprocessed image is input into the backbone network, and the backbone network will expand the semantic information of the RGB three channels into a high-dimensional space. Then, feature maps of three scales are generated These feature maps come from different levels of the network structure, so they contain semantic information at different levels, facilitating the network to identify target objects of different sizes for subsequent object detection and classification tasks.

[0077] Specifically, after the preprocessed image is input into the backbone network, target foggy day features of different sizes are extracted through layers of feature extraction modules. Then it is input into the Enhanced Spatial Pyramid Pooling - Fast (ESPPF) module. This module is an improvement based on the SPPF module of YOLOv10. Its function is to fuse the semantic features extracted by the backbone through pooling operations of different depths and scales. This operation can improve the model's detection ability for target objects of different sizes and enhance robustness. In addition, ESPPF optimizes the computational efficiency compared to the traditional SPPF, accelerating the model's inference speed while maintaining the accuracy. Its structure is as Figure 4 shown. The Conv in the figure is a linear combination of traditional 2DConv, BatchNorm normalization, and SiLu activation.

[0078] Let the input feature map be After extracting the image features through Conv, we get Then, max - pooling operations with a window size of 7×7 and multiple 5×5 are used to enhance the feature information to obtain M3, M4, M5, During the continuous use of max - pooling operations, semantic information of different target object sizes can be extracted layer by layer in a concatenated manner. Then, the information of M2, M3, M5, M6 is concatenated together in the channel dimension (dim = 1) using the Concat operation to form Finally, the semantic information of target objects of different sizes is fused through a Conv operation to obtain the final output The process can be expressed as follows:

[0079] M2 = Conv(M1)

[0080] M3 = MaxPool5×5(M2)

[0081] M5 = MaxPool5×5(MaxPool7×7(M3))

[0082] M6 = MaxPool5×5(M5)

[0083] M7 = Convat(M2, M3, M5, M6)

[0084] M8 = Conv(M7)

[0085] Step 4: Feature fusion.

[0086] The Neck network in the target detection model is responsible for feature fusion. The Neck network is located between the backbone network and the detection head, playing a connecting role. The Neck network uses a feature pyramid to post-process and fuse the three feature maps (P3, P4, and P5) mentioned above to enhance the representation of features, so that the model can better capture information about target objects of different sizes. In addition, the Neck network is also responsible for further processing the semantic features of the image, including upsampling, downsampling, and dimensionality adjustment operations to meet the needs of the detection head.

[0087] Step 5: Information post-processing decoding and prediction.

[0088] The Decoupled Head in the target detection model performs information post-processing decoding and prediction. The detection head is the key part responsible for the final target positioning and classification. The semantic features of different sizes extracted from the neck network are input into the decoupled detection head. After processing, each semantic feature input into the detection head will obtain two tensors, which store classification information and positioning information respectively. After parsing these two tensors, the final detection result of the current image can be obtained.

[0089] The above is the entire process of the model inference stage.

[0090] Step 6: Loss calculation.

[0091] During the training phase of the model, after each epoch of training is completed, the network will calculate the various losses (dehazing, positioning, classification, and confidence) based on the two tensors obtained by the detection head, and then calculate the gradient values of each module of the model based on the total loss obtained, and then update the parameters of various parts of the network through back propagation to achieve the purpose of optimizing the model.

[0092] The present invention adds the defogging loss to the loss value of the target detection model, and optimizes the detection model and FMDBlock together. The total loss of the MSE formula and the model can be expressed as follows:

[0093]

[0094] Loss = L cls +L CIoU +L DFL +L Defog

[0095] Where n is the total number of pixels, represents the pixel value of the clear image in position i, Represents the pixel value of the dehazed image at position i. cls represents the category loss, L CIoU and L DFL represents the positioning loss, LDefog Denotes the defogging loss, calculated by MSE.

[0096] It is worth mentioning that although the epoch of the present invention is set to 500, the present invention adopts an early stopping mechanism, that is, if the best model has not been updated within 100 epochs during the model training phase, it means that the model has converged and the model training is stopped. This greatly saves the time cost of researchers compared with a fixed epoch.

[0097] Based on the above invention content, the following evaluation of technical effects is proposed:

[0098] Performance evaluation, the present invention conducts experiments and evaluations on the Foggy Cityscapes dense fog dataset. This dataset contains 5000 synthetic foggy images with 2048×1024 pixels, where the training set contains 2975 images, the validation set contains 500 images, and the test set contains 1525 images. This dataset includes eight categories: bicycle, bus, car, motorcycle, pedestrian, cyclist, truck, and train. In addition, this synthetic foggy dataset includes three categories: light fog, medium fog, and dense fog. In order to verify the effectiveness of the method of the present invention, the present invention selects the challenging dense fog dataset for experiments.

[0099] Currently, the YOLOv10 model has multiple versions with different parameter types. Since the present invention requires the model to have real-time performance, after weighing the accuracy and complexity of the model, the present invention selects the model with size m as the benchmark model of the present invention. The comparison test results with the benchmark model on the Foggy Cityscapes dense fog dataset are shown in Table 1:

[0100] Table 1 Comparison test results on the Foggy Cityscapes dense fog dataset

[0101]

[0102] As can be seen from the table, compared with the benchmark model, although the number of parameters and the amount of calculation of the present invention increase by 5.1% and 6.8% respectively, the average accuracy (P), the average recall rate (R), and the average accuracy (mAP50) at a threshold of 0.5 increase by 11.5, 0.6, and 3.4 percentage points respectively. That is, the present invention exchanges a slight increase in model complexity for a huge improvement in model performance.

[0103] As described above, it is only used to help understand the method of the present invention and its core concept. However, the protection scope of the present invention is not limited thereto. For those of ordinary skill in the art within the technical scope disclosed by the present invention, any equivalent replacement or change made according to the technical solution of the present invention and its inventive concept should be covered within the protection scope of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A real-time target detection method for foggy days based on a strong and robust dehazing module, characterized in that It includes the following steps: S1. Data preprocessing: Collect foggy images and clear images as image data, and perform image enhancement on the image data; S2. Design a defogging module to defog the foggy image to obtain a defogged image, and perform image sharpening on the defogged image; The defogging module consists of a mapper and a remapper. A skip connection is added between the mapper and the remapper. Both the mapper and the remapper include three two-dimensional convolution operations; Image sharpening uses a Laplacian sharpening kernel to process the image; S3. Build an object detection model for object detection on the processed defogged image; The object detection model includes a backbone network, a neck network, and a decoupled detection head; S4. Model training and design a loss function for model optimization; The loss function includes the defogging loss of the defogging module and the total loss of the object detection model, and jointly optimizes the defogging module and the object detection model; S5. Object detection: Use the optimized object detection model based on the defogging module to perform object detection on the foggy image to be detected.

2. The method according to claim 1, wherein In S2, the defogging module defogs the foggy image as follows: The mapper gradually expands the channel dimension of the image to 12 dimensions using three two-dimensional convolutional operations to obtain F E1 , F E2 , F E3 , F E4 ; The reflector then uses three two-dimensional convolution operations to gradually compress the channel dimension of the image to 3 dimensions, which conforms to the RGB image, to obtain F D1 , F D2 , F D3 , F D4 ; During the decoding process, add F Ei and the processed F Di-1 to obtain F Di , where i = 2, 3, 4.

3. The method according to claim 2, wherein The two-dimensional convolution is a linear combination of 2DConv, BatchNorm normalization, and SiLu activation.

4. The method according to claim 1, wherein In S2, the image sharpening process is as follows: The Laplacian sharpening kernel and the sharpening formula are expressed as follows: F(x,ω)=(1-ω)·I(x)+ω·Lap(I(x)) (ω=0.5) Where, Kernel represents a 3×3 sharpening convolution kernel, F(x,ω) represents the sharpened image, I(x) represents the input image, Lap(I(x)) represents the Laplacian filtering, and ω represents the sharpening weight parameter.

5. The method according to claim 1, wherein In S3, the object detection model is built based on the YOLOv10 model and the backbone network is improved. A strengthened spatial pooling pyramid module is added to the backbone network to fuse the extracted semantic features through pooling operations of different depths and scales.

6. The method according to claim 5, characterized in that, The processing of the strengthened spatial pooling pyramid module is as follows: Let the input feature map be After extracting image features through Conv, we get Then, we use max pooling operations with a window size of 7×7 and multiple 5×5 to enhance the feature information and obtain Next, we use the Concat operation to concatenate the information of M2, M3, M5, and M6 on the channel dimension dim = 1 to form Finally, we use a Conv operation to fuse the semantic information of target objects of different sizes to obtain the final output The process is as follows: M2=Conv(M1) M3=MaxPool5×5(M2) M5=MaxPool5×5(MaxPool7×7(M3)) M6=MaxPool5×5(M5) M7=Convat(M2,M3,M5,M6) M8=Conv(M7) Where, Conv is a linear combination of 2DConv, BatchNorm normalization, and SiLu activation.

7. The method according to claim 5, wherein In S3, the object detection is as follows: Feature extraction: Use the improved backbone network to extract semantic features of different scales from the processed defogged image; Feature Fusion: Use the neck network to perform post-processing and fusion on the feature maps of multiple scales to obtain semantic features of different scales; Information post-processing decoding and prediction: Each semantic feature is processed to obtain two tensors storing classification information and localization information respectively.

8. The method according to any one of claims 1-7, characterized in that, In S4, the loss function is as follows: The defogging module uses the MSE formula to calculate the distance between the defogged image and the clear image, and uses it as the defogging loss of the defogging module; The defogging loss and the total loss of the object detection model are expressed as follows: Loss=L cls +L CIoU +L DFL +L Defog where n is the total number of pixels, represents the pixel value of the clear image at position i, represents the pixel value of the defogged image at position i; L cls represents the class loss, L CIoU and L DFL represent the localization loss, L Defog represents the defogging loss, which is calculated by MSE.

9. A strongly robust foggy day real-time object detection system based on a defogging module applied to the method according to claim 8, characterized in that, It includes: The dehazing module, consisting of a mapper and a demapper, where both the mapper and the demapper include three two-dimensional convolution operations for image dehazing processing; The image sharpening module, using a Laplacian sharpening kernel for image sharpening processing; The object detection model, including a BackBone backbone network, a Neck neck network, and a Decoupled Head decoupled detection head. The BackBone backbone network incorporates a strengthened spatial pooling pyramid module for image object detection.