Deep learning-based small target detection method in haze scene

By designing the FASOD-Net network, using technical means such as multi-scale fusion, large convolution kernel and parallel enhanced attention module, the accuracy and robustness of small object detection in haze scenarios are solved, and the detection accuracy and robustness are significantly improved.

CN120014233AActive Publication Date: 2025-05-16GUANGZHOU ZHONGNENG DIGITAL INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510091973.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In haze scenarios, traditional target detection methods are difficult to effectively detect small targets, and existing deep learning models are difficult to distinguish targets from backgrounds when extracting features, resulting in mis-detection and missed detection problems.

Method used

A small object detection method in haze scenarios based on deep learning is designed, called FASOD-Net. The network includes a defog removal and shallow feature extraction module, a fine feature capture multi-channel convolution module and a multi-scale feature integration enhancement detection module. Through technical means such as multi-scale fusion, large convolution kernel and parallel enhanced attention module, it enhances the ability to extract and distinguish small-target features.

Benefits of technology

It significantly improves the accuracy and robustness of small target detection in haze scenarios, effectively reduces the interference of haze on small target characteristics, enhances the detection ability of small and slender targets, and significantly improves the detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014233A_ABST
    Figure CN120014233A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target detection, and particularly relates to a deep learning-based small target detection method in a haze scene, which comprises the following specific steps of: constructing a small target detection original data set in the haze scene, or using an existing disclosed small target detection data set in the haze scene; the method comprises the following steps: firstly, collecting small target image data in a haze environment, then carrying out frame selection on the position of a small target by using a labeling tool, and carrying out label labeling; according to the method, a small target detection original data set in a haze scene is preprocessed, a preprocessed small target data set in the haze scene is obtained, and the data set is divided, so that a deep learning-based small target detection model in the haze scene is trained subsequently. When a small target detection task in a haze scene is processed, detail features of the small target can be captured and recognized more accurately, and interference of haze on the features of the small target is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a small target detection method in haze scenes based on deep learning. Background Art

[0002] As a serious environmental pollution problem, haze usually occurs in cities or industrialized areas with poor air quality and low visibility. Its high pollution concentration and long-term impact not only endanger human health, but also bring great challenges to traffic safety, environmental monitoring and other fields. Especially in the fields of traffic monitoring, public safety and environmental management, the detection of small targets (such as pedestrians, vehicles, key equipment, etc.) in haze scenes has become the key to ensuring safety and improving management efficiency. However, due to the occlusion effect of haze, traditional target detection methods show obvious deficiencies in the detection of small targets in images under low visibility conditions. New technical means are urgently needed to improve the accuracy and robustness of small target detection in haze scenes.

[0003] In haze scenes, target detection faces multiple challenges: on the one hand, haze causes the contrast of target images to decrease, edges to blur, and details to be lost, making it difficult for traditional feature engineering-based methods (such as edge detection, color segmentation, and other technologies) to effectively extract target information; on the other hand, small targets in haze environments are more easily ignored by traditional detection methods due to their small size and high degree of integration with background information, resulting in a significant increase in missed detection rates. In addition, complex ambient lighting changes and optical scattering effects in haze scenes further increase the difficulty of target detection.

[0004] Traditional small target detection methods, such as image processing algorithms based on manually designed features, often require a lot of manual debugging, and it is difficult to achieve efficient and reliable detection results when facing complex haze scenes. At the same time, these methods usually cannot take into account the detection needs of multi-scale targets and lack adaptability to the morphological diversity and complex backgrounds of targets. In recent years, with the rapid development of deep learning technology, especially the widespread application of models such as convolutional neural networks (CNN) in the field of target detection, small target detection methods based on deep learning have shown significant performance advantages. Deep learning can better capture the texture information and spatial distribution characteristics of the target by automatically extracting multi-level features from the data, providing a new solution for small target detection in haze scenes.

[0005] Although deep learning technology has made significant progress in target detection, small target detection in haze scenes still faces many technical bottlenecks. For example, due to the degradation of image quality caused by haze, existing deep learning models may have difficulty distinguishing targets from backgrounds during feature extraction, resulting in false detection and missed detection. In addition, small targets often occupy a relatively low proportion of the image size, and existing methods often lose key information due to downsampling operations, affecting detection accuracy. Therefore, how to design a deep learning model that is more adaptable to low-quality images in haze scenes, especially to enhance its ability to extract small target features, is an important direction of current research.

[0006] Based on the above, the existing small target detection model in haze scenes based on deep learning has the following problems:

[0007] 1. The interference of haze scenes on images leads to decreased detection performance: Existing target detection methods perform poorly in haze scenes, mainly because haze causes reduced image contrast, blurred details and loss of target features. In this case, it is difficult for traditional models to extract reliable target features, especially for small targets, which are prone to missed detection or false detection.

[0008] 2. Small targets are intertwined with complex backgrounds, resulting in insufficient feature expression capabilities: Small targets are often intertwined with complex backgrounds, making it difficult for existing target detection models to distinguish between targets and backgrounds. Especially in complex scenes, the features of small targets are easily masked by background noise, resulting in high false detection and missed detection rates. Existing methods have weak expression capabilities for small target features and are unable to fully extract the details and boundary information of small targets. There is a lack of effective feature enhancement mechanisms, and target features are easily submerged by background information during feature extraction. The detection accuracy and reliability of small targets are insufficient, especially in the case of tiny targets or complex scenes.

[0009] 3. Insufficient integration of multi-scale features leads to limited model recognition ability: Small target detection requires the integration of multi-level and multi-scale feature information, but the existing models are insufficient in feature integration and find it difficult to capture the correlation between features at different levels, especially for the detection of tiny and slender targets.

[0010] 4. Traditional methods are not capable of processing multi-scale and multi-level features: Traditional target detection methods use fixed convolution structures or simple feature fusion techniques, which are difficult to adapt to the characteristic changes of small targets at different scales, resulting in unsatisfactory detection effects for multi-scale targets.

[0011] Therefore, a small target detection method in haze scenes based on deep learning is invented. Summary of the invention

[0012] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:

[0013] A method for detecting small targets in haze scenes based on deep learning includes the following specific steps:

[0014] S1: Construct an original dataset for small target detection in haze scenes, or use an existing publicly available dataset for small target detection in haze scenes; first collect image data of small targets in haze environments, then use annotation tools to select the positions of small targets and annotate them with labels;

[0015] S2: preprocessing the original data set of small target detection in haze scenes to obtain a preprocessed data set of small target detection in haze scenes, and dividing the data set to facilitate subsequent training of a small target detection model in haze scenes based on deep learning;

[0016] S3: Construct a small target detection network FASOD-NET in haze scenes;

[0017] S4: Based on the constructed FASOD-Net model of small target detection network in haze scenes, train and update the parameters of each layer; first initialize all neural network parameters and set model-related hyperparameters, including training rounds, batch size, weight decay coefficient, learning rate, and total number of iterations;

[0018] S5: After the small target detection model for haze scenes is trained, the trained small target detection model for haze scenes is used to detect small targets in haze scene images; finally, the small target detection results in the haze scene are output, including the small target bounding box coordinates, small target category labels and confidence scores.

[0019] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the specific steps of S2 are as follows:

[0020] S21: The haze scene images in the original dataset of small target detection in haze scenes are cropped to a uniform size. Since the original dataset is small, in order to ensure that the dataset has enough samples for training, verification and testing, and to enhance the robustness of the small target detection model in haze scenes, the sensitivity of the model to image quality is reduced;

[0021] S22: Then, data enhancement processing is performed to better train the deep learning model, and finally a data-enhanced small target data set in the haze scene is obtained, which is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0022] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the specific steps of S3 are as follows:

[0023] S31: Input the haze scene image into the defogging and shallow feature extraction module;

[0024] S32: input F4 into the multi-channel convolution module for fine feature capture;

[0025] S33: Input F2, F3, and F5 into the multi-scale feature integration enhancement detection module.

[0026] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the construction and execution process of the defogging and shallow feature extraction module in S31 is as follows:

[0027] S311: applying a CBS operation including a 3×3 convolution, normalization, and activation function to the input haze scene image to obtain a haze shallow feature map F1;

[0028] S312: Input F1 to the multi-scale fusion large convolution kernel enhancement module MSE-LKC;

[0029] S313: Input F2 into the parallel enhanced attention module PEA;

[0030] S314: Input F3 and F1 into Fusion for channel dimension fusion, and obtain the final output feature map F4 by fusing features at different levels.

[0031] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the operation and construction process of the MSE-LKC module in S312 is as follows:

[0032] S3121: Perform batch normalization on the input F1 to obtain a feature map F1_1, then use a convolution kernel of size 7×7 to perform a convolution operation on F1_1 to extract features to obtain a feature map F1_2, and then use a convolution kernel of size 5×5 to perform a convolution operation on F1_2 to obtain a feature map F1_3 of the first stage;

[0033] S3122: Input F1_3 to a deep dilated convolution module with a dilation rate of 7×7, a deep dilated convolution module with a dilation rate of 13×13, and a deep dilated convolution module with a dilation rate of 19×19 to extract multi-scale features, and generate three corresponding feature sub-graphs F1_4, F1_5, and F1_6 respectively; then, concatenate F1_4, F1_5, and F1_6 in the channel dimension through the C operation to obtain a fused feature graph F1_7;

[0034] S3123: Input F1_7 into a convolution kernel of size 1×1 for convolution operation to obtain feature map F1_8, process F1_8 through GELU activation function to obtain feature map F1_9, perform Conv1×1 convolution operation on F1_9 to obtain the final output feature map F1_10;

[0035] S3124: F1 and F1_10 are added through residual connection to form a feature map F2 with stronger representation ability.

[0036] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the construction and execution process of the PEA module in S313 is as follows:

[0037] S3131: perform batch normalization on F2 to standardize feature distribution, improve the stability and convergence speed of the model, and obtain the normalized feature map F2_1;

[0038] S3132: Input F2_1 into three attention modules respectively:

[0039] S31321: The pixel-level attention module processes the spatial dimension of F2_1, extracts local saliency information, and obtains the enhanced feature sub-graph F2_2;

[0040] S31322: The channel-level attention module operates on the channel dimension of F2_1 to capture global context information and obtain the weight-adjusted feature subgraph F2_3;

[0041] S31323: The simple pixel-level attention module uses a simplified mechanism to process F2_1, supplementing the feature extraction capabilities of other modules to obtain the feature sub-graph F2_4;

[0042] S3133: concatenate F2_2, F2_3 and F2_4 in the channel dimension to obtain a fused feature map F2_5;

[0043] S3134: Perform a Conv1×1 convolution operation on F2_5 to achieve feature dimensionality reduction, and perform a nonlinear transformation on the feature map through a GELU activation function to obtain a processed feature map F2_6; then perform a Conv1×1 convolution operation on F2_6 to further extract features and adapt to subsequent tasks to obtain a feature map F2_7;

[0044] S3135: Add F2 and F2_7 through residual connection to obtain the final output feature map F3;

[0045] Among them, the process of applying PEA to process F2 to obtain F3 is shown in the following formula:

[0046] F2_1=BatchNorm(F2)

[0047] F3=F2+Conv1_1(GELU(Conv1_1(C(PA(F2_1),CA(F2_1),SPA(F2_1)))))

[0048] Among them, BatchNorm represents batch normalization operation; Conv1_1 represents a convolution operation of size 1×1; GELU represents the application of activation function GELU; SPA represents a simple pixel-level attention module; CA represents a channel-level attention module; PA represents a pixel-level attention module; C represents a splicing operation on the channel dimension; + represents a residual connection operation.

[0049] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the construction and execution process of the pixel-level attention module in S31321 is as follows:

[0050] S313211: Perform a Conv1×1 convolution operation on F2_1 to adjust the number of channels of the input feature map; then, apply the activation function GELU to further process F2_1 to obtain the activated feature map F2_1_1;

[0051] S313212: Perform Conv1×1 convolution operation on F2_1_1 to obtain the intermediate feature map F2_1_2; then, apply the Sigmoid activation function to normalize F2_1_2 to obtain the pixel-level attention weight;

[0052] S313213: Perform element-by-element multiplication of F2_1 and pixel attention weights, weight the feature map, and obtain the final feature map F2_2; this module can enhance the expression of small target features while suppressing irrelevant background information, thus improving the model’s ability to detect small targets in haze scenes;

[0053] Among them, the process of applying the pixel-level attention module to process F2_1 and obtaining F2_2 is shown in the following formula:

[0054] F2_2=F2_1×S(Conv1_1(GELU(Conv1_1(F2_1))))

[0055] Among them, Conv1_1 represents a convolution operation of size 1×1; GELU represents the application of the activation function GELU; S represents the Sigmoid activation function; × represents an element-by-element multiplication operation;

[0056] The construction and execution process of the channel-level attention module in S31322 is as follows:

[0057] S313221: Perform global average pooling on F2_1 to generate global information in the channel dimension; then, perform Conv1×1 convolution operation on the pooling result to obtain feature map T1, and perform nonlinear processing on T1 through the activation function GELU to generate the intermediate expression feature map T2;

[0058] S313222: After performing a Conv1×1 convolution operation on T2, a channel attention weight is generated through a Sigmoid activation function; F2_1 and the channel attention weight are multiplied element by element to achieve weighted output on the channel, and an enhanced output feature map F2_3 is obtained;

[0059] Among them, the process of applying the channel-level attention module to process F2_1 to obtain F2_3 is shown in the following formula:

[0060] F2_3=F2_1×S(Conv1_1(GELU(Conv1_1(GAP(F2_1)))))

[0061] Among them, GAP represents the global average pooling operation; Conv1_1 represents the convolution operation of size 1×1; GELU represents the application of the activation function GELU; S represents the Sigmoid activation function; × represents the element-by-element multiplication operation;

[0062] The construction and execution process of the simple pixel-level attention module in S31323 is as follows:

[0063] S313231: Perform a Conv1×1 convolution operation on F2_1 to obtain a preliminary pixel feature map X1; then, X1 is convolved with another convolution kernel of size 3×3 to further extract local pixel correlation information to obtain a feature map X2;

[0064] S313232: Perform a Conv1×1 convolution operation directly on F2_1 through another branch to obtain feature map X3; then, add X2 and X3 and generate pixel attention weights through the Sigmoid activation function;

[0065] S313233: perform element-by-element multiplication of F2_1 and the pixel attention weight to obtain an enhanced output feature map F2_4;

[0066] Among them, the process of applying a simple pixel-level attention module to process F2_1 to obtain F2_4 is shown in the following formula:

[0067] F2_4=F2_1×s(Conv3_3(Conv1_1(F2_1))+Conv1_1(F2_1))

[0068] Among them, Conv1_1 represents a convolution operation of size 1×1; Conv3_3 represents a convolution operation of size 3×3; S represents the Sigmoid activation function; + represents the addition operation of the results of the two branches; × represents the element-by-element multiplication operation.

[0069] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the construction and execution process of the fine feature capture multi-channel convolution module in S32 is as follows:

[0070] S321: Perform a Conv1×1 convolution operation on F4 to adjust the number of channels and extract preliminary features to generate a feature map F4_1;

[0071] S322: feature splitting of F4_1 into three feature processing paths: in the first path, firstly, two Conv 3×3 convolution operations are performed to extract local deep features and generate feature map F4_2; in the second path, Conv3×3 convolution operations and Conv5×5 convolution operations are performed in sequence to further extract features of a larger receptive field and generate feature map F4_3; in the third path, Conv5×5 convolution operations are performed to further extract features of a larger receptive field and generate feature map F4_4;

[0072] S323: splice F4_2, F4_3 and F4_4 output from the three paths in the channel dimension to obtain a fused feature map F4_5; perform channel integration on F4_5 through a Conv1×1 convolution operation to generate a final small target deep feature map F4_6;

[0073] S324: Perform a Conv3×3 convolution operation on F4_6 to further integrate cross-path information and optimize local feature expression to generate feature map F5. Through the above steps, its module combines multi-scale convolution kernels and diversion structures to effectively extract deep semantic features of small targets and enhance the model's detection capability for small targets.

[0074] As a preferred solution of the method for detecting small targets in haze scenes based on deep learning described in the present invention, the construction and execution process of the multi-scale feature integration enhancement detection module in S33 is as follows:

[0075] S331: The input F2 is subjected to a deformable convolution operation to dynamically adapt to the morphological characteristics of small targets and complex background noise in the haze scene to generate an enhanced feature map F6; the input F3 is subjected to a deformable convolution operation to generate an enhanced feature map F7;

[0076] S332: concatenate F5 and F7 in the channel dimension to form a fused feature map F8, which enhances the multi-scale expression capability of the features and takes into account the unification of details and global features; then, perform a Conv 3×3 convolution operation on F8 to further integrate cross-path information and optimize the multi-level expression of features to generate an optimized feature map F9; then concatenate F9 and F6 in the channel dimension to form a fused feature map F10;

[0077] S333: Perform a Conv3×3 operation on F10 to further integrate cross-path information and optimize the multi-level expression of features to obtain a feature map F11; input F11 into the target detection head Head, use Head to detect F11, and output a tensor containing detection information, where each row corresponds to a detection result, and the detection result is a small target detection result in a haze scene, including a bounding box coordinate, a small target category label, and a confidence score.

[0078] Compared with existing technologies:

[0079] When processing small target detection tasks in haze scenes, the present invention can more accurately capture and identify the detailed features of small targets, effectively reduce the interference of haze on the features of small targets, and solve the problem of unclear expression of small target features caused by haze in the prior art; the present invention can more accurately distinguish small targets from backgrounds through multi-scale feature extraction and multi-level feature fusion, and solves the problem that small targets are easily confused with background information in complex haze scenes; in addition, the present invention enhances the detection capability of tiny and slender small targets, significantly reduces the problem of small target feature loss caused by excessive pooling and convolution operations in the prior art, and significantly improves the accuracy and robustness of small target detection in haze scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is a schematic diagram of the process of the present invention;

[0081] Figure 2 It is a schematic diagram of the main structure of FASOD-Net of the present invention;

[0082] Figure 3 It is a schematic diagram of the defogging and shallow feature extraction module of the present invention;

[0083] Figure 4 It is a schematic diagram of the MSE-LKC structure of the present invention;

[0084] Figure 5 It is a schematic diagram of the PEA structure of the present invention;

[0085] Figure 6 Schematic diagram of the structure of the pixel-level attention module of the present invention;

[0086] Figure 7 Schematic diagram of the channel-level attention module structure of the present invention;

[0087] Figure 8 This is a schematic diagram of the structure of a simple pixel-level attention module of the present invention;

[0088] Fig. 9 Schematic diagram of a multi-channel convolution module for capturing fine features of the present invention;

[0089] Fig.10 Schematic diagram of the multi-scale feature integration and enhanced detection module of the present invention. DETAILED DESCRIPTION

[0090] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0091] The present invention provides a small target detection method in a haze scene based on deep learning, which is used to detect and identify small targets in a haze environment. The present invention designs a small target detection network in a haze scene called FASOD-Net (Fog-Aware Small Object Detection Network). The network can effectively remove the interference of haze on the image, deeply analyze the foggy image, capture and extract the features of small targets, and finally generate a target detection result image. In the target detection result image, the detected small targets are marked with a bounding box, and the background area remains unchanged; please refer to Figure 1-Figure 10 ;

[0092] The specific steps are as follows:

[0093] S1: Construct an original dataset for small target detection in haze scenes, or use an existing publicly available dataset for small target detection in haze scenes; first collect image data of small targets in haze environments (e.g., by using drones or cameras to take field photos at different times and locations, or by using existing publicly available datasets), then use annotation tools to select the positions of small targets (i.e., determine the horizontal and vertical coordinates of the upper left and lower right corners of the frame), and annotate them with labels;

[0094] S1 includes but is not limited to the following embodiments: the richer the types of small targets contained in the data set, the more small targets the model can detect; based on the self-collected original data set of small target detection in haze scenes, a total of 2000 images, including 800 small vehicle targets, 600 small personnel targets, 350 small animal targets, and 250 small building detail targets; the labels here can be divided into four categories: vehicles, personnel, animals, and building details; use the LabelImg tool to annotate the data set and export the xml label file, which contains the label name and annotation box information of the target small target;

[0095] Among them, the LabelImg tool is an open source image annotation tool, which is mainly used for image dataset annotation for tasks such as target detection and image segmentation;

[0096] S2: preprocessing the original data set of small target detection in haze scenes to obtain a preprocessed data set of small target detection in haze scenes, and dividing the data set to facilitate subsequent training of a small target detection model in haze scenes based on deep learning;

[0097] The specific steps of S2 are as follows:

[0098] S21: The haze scene images in the original dataset of small target detection in haze scenes are cropped to a uniform size. Since the original dataset is small, in order to ensure that the dataset has enough samples for training, verification and testing, and to enhance the robustness of the small target detection model in haze scenes, the sensitivity of the model to image quality is reduced;

[0099] S22: Then, data enhancement processing is performed to better train the deep learning model, and finally a data set of small targets in the haze scene is obtained after data enhancement, and the data is divided into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0100] S2 includes but is not limited to the following embodiments: the original data set for small target detection in haze scenes has a total of 2000 images, and the images are first cropped to a uniform size of 512×512 pixels; the original data set is expanded by a data enhancement method, specifically including adding random pixels, Gaussian noise, salt and pepper noise, image contrast adjustment, image flipping, etc., which can be implemented in Python language; random pixels use the cv2.randu() function, Gaussian noise is achieved through the ImageFilter.GaussianBlur() method in the PIL library, salt and pepper noise can use the cv2.randSalt() function, image contrast adjustment can use the ImageEnhance.Contrast method, and flipping operation is achieved through the cv2.flip() function; the final data set is enhanced and expanded to 12,000 images, and divided according to a ratio of 8:1:1 to obtain 9,600 training sets, 1,200 validation sets, and 1,200 test sets;

[0101] S3: Construct a small target detection network FASOD-NET in haze scenes;

[0102] Input the haze scene images from the small target dataset under haze scene constructed in Step 2 into FASOD-Net;

[0103] The present invention designs a small target detection network FASOD-Net (Fog-Aware Small Object Detection Network) in haze scenes. The process of using FASOD-Net to detect small targets in haze scene images is divided into three steps: defogging and small target shallow feature extraction; fine feature capture; multi-level feature fusion and small target detection; the above three main steps correspond to the three modules of the FASOD-Net network model: defogging and shallow feature extraction module, fine feature capture multi-channel convolution module, and multi-scale feature integration enhanced detection module;

[0104] The specific steps of S3 are as follows:

[0105] S31: Input the haze scene image into the defogging and shallow feature extraction module. The defogging and shallow feature extraction module is as follows: Figure 3 As shown;

[0106] The construction and execution process of the dehazing and shallow feature extraction modules in S31 are as follows:

[0107] Input the haze scene image into the dehazing and shallow feature extraction module; perform CBS operation on the haze scene image to obtain the haze shallow feature map F1; then, input F1 into the MSE-LKC module to obtain the feature map F2; then, input F2 into the PEA module to obtain the feature map F3; finally, input F3 and F1 into Fusion to obtain the final output feature map F4;

[0108] S311: Apply a CBS operation including a 3×3 convolution, normalization, and activation function to the input haze scene image to obtain a haze shallow feature map F1; wherein the CBS operation is a basic convolution module, which extracts local features of the image through a convolution kernel of size 3×3, and then performs batch normalization (BatchNorm) to stabilize the feature distribution, and introduces nonlinearity through the activation function to enhance the feature expression ability of the model;

[0109] S312: Input F1 to the multi-scale fusion large convolution kernel enhancement module MSE-LKC. The structure of the MSE-LKC module is as follows: Figure 4 As shown;

[0110] The operation and construction process of the MSE-LKC module in S312 is:

[0111] S3121: Perform batch normalization on input F1 (i.e. Figure 4 The BatchNorm operation in the above example is used to obtain the feature map F1_1, and then a convolution operation with a convolution kernel of size 7×7 is performed on F1_1 to extract features (i.e. Figure 4 Conv7×7 in the image is used to obtain the feature map F1_2, and then a convolution kernel of size 5×5 is used to convolve F1_2 to obtain the feature map F1_3 of the first stage;

[0112] S3122: Input F1_3 to the depth expansion convolution module with expansion rate of 7×7 (i.e. Figure 4 DWDConv7 in ), a deep dilated convolution module with a dilation rate of 13×13 (i.e. Figure 4 DWDConv13 in ), a deep dilated convolution module with a dilation rate of 19×19 (i.e. Figure 4 DWDConv19 in the above code is used to extract multi-scale features and generate three corresponding feature sub-graphs F1_4, F1_5 and F1_6 respectively; then, F1_4, F1_5 and F1_6 are concatenated in the channel dimension through the C operation (i.e., the Concat operation) to obtain the fused feature graph F1_7;

[0113] S3123: Input F1_7 into a convolution kernel of size 1×1 for convolution operation (i.e., Conv1×1 operation in the figure) to obtain feature map F1_8, process F1_8 through GELU activation function to obtain feature map F1_9, and perform Conv1×1 convolution operation on F1_9 to obtain the final output feature map F1_10;

[0114] S3124: F1 and F1_10 are added through residual connection (i.e., the + operation in the figure) to form a feature map F2 with stronger representation ability;

[0115] Among them: BatchNorm is a commonly used deep learning technology. It stabilizes the activation distribution of the neural network by standardizing the features of each batch of data to make its mean 0 and variance 1. It dynamically adjusts the scale of the network input during training, alleviates the problem of gradient vanishing or exploding, accelerates model convergence, and improves the stability and generalization ability of training. In addition, BatchNorm also has a certain regularization effect, which helps to reduce dependence on other regularization methods.

[0116] The GELU (Gaussian Error Linear Unit) activation function is a smooth nonlinear activation function that combines the advantages of ReLU and Sigmoid. Its mathematical expression is: GELU(x) = x·Φ(x), where Φ(x) is the cumulative distribution function of the standard normal distribution. Compared with ReLU and Sigmoid, GELU smoothly adjusts the activation according to the size of the input value. Larger values ​​pass through and smaller values ​​are suppressed, providing a more natural gradient flow. It is widely used in deep learning models (such as Transformer) to improve training effects and performance.

[0117] S313: Input F2 into the parallel enhanced attention module PEA. The structure of the PEA module is as follows: Figure 5 As shown;

[0118] The construction and execution process of the PEA module in S313 is as follows:

[0119] S3131: Batch normalization of F2 (i.e. Figure 5 The Batch Norm operation in the model is used to standardize the feature distribution, improve the stability and convergence speed of the model, and obtain the normalized feature map F2_1;

[0120] S3132: Input F2_1 into three attention modules respectively:

[0121] S31321: The pixel-level attention module processes the spatial dimension of F2_1, extracts local saliency information, and obtains the enhanced feature sub-graph F2_2;

[0122] The construction and execution process of the pixel-level attention module in S31321 is as follows:

[0123] Input F2_1 into the pixel-level attention module (Pixel Attention, PA). The structure of the Pixel Attention module is as follows: Figure 6 As shown;

[0124] S313211: Perform a Conv1×1 convolution operation on F2_1 to adjust the number of channels of the input feature map; then, apply the activation function GELU to further process F2_1 to obtain the activated feature map F2_1_1;

[0125] S313212: Perform Conv1×1 convolution operation on F2_1_1 to obtain the intermediate feature map F2_1_2; then, apply the Sigmoid activation function to normalize F2_1_2 to obtain the pixel-level attention weight;

[0126] S313213: Perform element-by-element multiplication of F2_1 and pixel attention weights, weight the feature map, and obtain the final feature map F2_2; this module can enhance the expression of small target features while suppressing irrelevant background information, thus improving the model’s ability to detect small targets in haze scenes;

[0127] Among them, the process of applying the pixel-level attention module to process F2_1 and obtaining F2_2 is shown in the following formula:

[0128] F2_2=F2_1×S(Conv1_1(GELU(Conv1_1(F2_1))))

[0129] Among them, Conv1_1 represents a convolution operation of size 1×1; GELU represents the application of the activation function GELU; S represents the Sigmoid activation function; × represents an element-by-element multiplication operation;

[0130] S31322: The channel-level attention module (Channel Attention) operates on the channel dimension of F2_1 to capture global context information and obtain the weighted feature subgraph F2_3;

[0131] The construction and execution process of the channel-level attention module in S31322 is as follows:

[0132] Input F2_1 into the channel-level attention module (Channel Attention, CA). The structure of the ChannelAttention module is as follows Figure 7 As shown;

[0133] S313221: Perform global average pooling (i.e., GAP operation in the figure) on F2_1 to generate global information in the channel dimension; then, perform Conv1×1 convolution operation on the pooling result to obtain feature map T1, and perform nonlinear processing on T1 through activation function GELU to generate intermediate expression feature map T2;

[0134] S313222: After performing a Conv1×1 convolution operation on T2, a channel attention weight is generated through a Sigmoid activation function; F2_1 and the channel attention weight are multiplied element by element to achieve weighted output on the channel (i.e., the multiplication operation in the figure), and an enhanced output feature map F2_3 is obtained;

[0135] Among them, the process of applying the channel-level attention module to process F2_1 to obtain F2_3 is shown in the following formula:

[0136] F2_3=F2_1×s(Conv1_1(GELU(conv1_1(GAP(F2_1)))))

[0137] Among them, GAP represents the global average pooling operation; Conv1_1 represents the convolution operation of size 1×1; GELU represents the application of the activation function GELU; S represents the Sigmoid activation function; × represents the element-by-element multiplication operation;

[0138] S31323: Simple Pixel Attention uses a simplified mechanism to process F2_1, supplementing the feature extraction capabilities of other modules to obtain feature subgraph F2_4;

[0139] The construction and execution process of the simple pixel-level attention module in S31323 is as follows:

[0140] Input F2_1 into the Simple Pixel Attention (SPA) module. The structure of the SimplePixel Attention module is as follows: Figure 8 As shown;

[0141] S313231: Perform a Conv1×1 convolution operation on F2_1 to obtain a preliminary pixel feature map X1; then, X1 is convolved with another convolution kernel of size 3×3 (i.e., the Conv3×3 operation in the figure) to further extract local pixel correlation information and obtain a feature map X2;

[0142] S313232: Perform a Conv1×1 convolution operation directly on F2_1 through another branch to obtain feature map X3; then, add X2 and X3 and generate pixel attention weights through the Sigmoid activation function;

[0143] S313233: Perform element-wise multiplication of F2_1 and pixel attention weight (i.e. Figure 8 The multiplication operation in ( ) is performed to obtain the enhanced output feature map F2_4;

[0144] Among them, the process of applying a simple pixel-level attention module to process F2_1 to obtain F2_4 is shown in the following formula:

[0145] F2_4=F2_1×s(Conv3_3(Conv1_1(F2_1))+Conv1_1(F2_1))

[0146] Among them, Conv1_1 represents a convolution operation of size 1×1; Conv3_3 represents a convolution operation of size 3×3; S represents the Sigmoid activation function; + represents the addition operation of the results of the two branches; × represents the element-by-element multiplication operation;

[0147] S3133: Concatenate F2_2, F2_3 and F2_4 in the channel dimension (i.e. Figure 5 C operation in the figure), and obtain the fused feature map F2_5;

[0148] S3134: Perform a Conv1×1 convolution operation on F2_5 to achieve feature dimensionality reduction, and perform a nonlinear transformation on the feature map through a GELU activation function to obtain a processed feature map F2_6; then perform a Conv1×1 convolution operation on F2_6 to further extract features and adapt to subsequent tasks to obtain a feature map F2_7;

[0149] S3135: Through residual connection (i.e. Figure 5 The “+” operation in the figure adds F2 and F2_7 to obtain the final output feature map F3;

[0150] Among them, the process of applying PEA to process F2 to obtain F3 is shown in the following formula:

[0151] F2_1=BatchNorm(F2)

[0152] F3=F2+Conv1_1(GELU(Conv1_1(C(PA(F2_1),CA(F2_1),SPA(F2_1)))))

[0153] Among them, BatchNorm represents batch normalization operation; Conv1_1 represents a convolution operation of size 1×1; GELU represents the application of activation function GELU; SPA represents a simple pixel-level attention module; CA represents a channel-level attention module; PA represents a pixel-level attention module; C represents a concatenation operation on the channel dimension; + represents a residual connection operation;

[0154] S314: Input F3 and F1 into Fusion for channel dimension fusion, and obtain the final output feature map F4 by fusing features at different levels;

[0155] S32: Input F4 into the fine feature capture multi-channel convolution module. The structure of the fine feature capture multi-channel convolution module is as follows: Fig. 9 As shown;

[0156] The construction and execution process of the fine feature capture multi-channel convolution module in S32 is as follows:

[0157] S321: Perform a Conv1×1 convolution operation on F4 to adjust the number of channels and extract preliminary features to generate a feature map F4_1;

[0158] S322: feature splitting of F4_1 into three feature processing paths: in the first path, firstly, two Conv 3×3 convolution operations are performed to extract local deep features and generate feature map F4_2; in the second path, Conv3×3 convolution operations and Conv5×5 convolution operations are performed in sequence to further extract features of a larger receptive field and generate feature map F4_3; in the third path, Conv5×5 convolution operations are performed to further extract features of a larger receptive field and generate feature map F4_4;

[0159] S323: Concatenate F4_2, F4_3 and F4_4 output from the three paths in the channel dimension (i.e. Fig. 9 C operation in the figure), and obtain the fused feature map F4_5; perform channel integration on F4_5 through the Conv1×1 convolution operation to generate the final small target deep feature map F4_6;

[0160] S324: Perform a Conv3×3 convolution operation on F4_6 to further integrate cross-path information and optimize local feature expression to generate feature map F5. Through the above steps, the module combines multi-scale convolution kernels and shunt structures to effectively extract deep semantic features of small targets and enhance the model's detection capability for small targets.

[0161] S33: Input F2, F3, and F5 into a multi-scale feature integration enhancement detection module. The structure of the multi-scale feature integration enhancement detection module is as follows: Fig.10 As shown;

[0162] The construction and execution process of the multi-scale feature integration enhanced detection module in S33 is as follows:

[0163] S331: The input F2 is subjected to a deformable convolution operation to dynamically adapt to the morphological characteristics of small targets and complex background noise in the haze scene to generate an enhanced feature map F6; the input F3 is subjected to a deformable convolution operation to generate an enhanced feature map F7;

[0164] S332: concatenate F5 and F7 in the channel dimension to form a fused feature map F8, which enhances the multi-scale expression capability of the features and takes into account the unification of details and global features; then, perform a Conv 3×3 convolution operation on F8 to further integrate cross-path information and optimize the multi-level expression of features to generate an optimized feature map F9; then concatenate F9 and F6 in the channel dimension to form a fused feature map F10;

[0165] S333: Perform a Conv3×3 operation on F10 to further integrate cross-path information and optimize the multi-level expression of features to obtain a feature map F11; input F11 into the target detection head Head, use Head to detect F11, and output a tensor containing detection information, where each row corresponds to a detection result, and the detection result is a small target detection result in the haze scene, including a bounding box coordinate, a small target category label, and a confidence score;

[0166] The head is responsible for directly detecting the category, position (including the coordinates of the bounding box) and the confidence of the existence of the object from the feature map. The output of the detection head is a three-dimensional tensor, in which each row contains the category probability, the coordinates of the bounding box and the confidence of the existence of the object for each detection box at each spatial position (relative to the feature map).

[0167] Bounding box prediction: Each cell is responsible for predicting the coordinates and confidence of several bounding boxes; the coordinates of the bounding box include: the center point (x, y) (relative to the position of the feature map cell), width and height (w, h) (scaled relative to the width and height of the entire image);

[0168] Category and confidence prediction: In addition to coordinate predictions, each bounding box also has a confidence prediction, which is used to indicate the probability of the existence of an object of a specific category within the box;

[0169] S3 includes but is not limited to the following embodiments: inputting a haze scene image under a haze scene into a dehazing and shallow feature extraction module; the size of the haze scene image is 512×512 pixels and 3 channels; using a convolution kernel of size 3×3 to perform a convolution operation on the haze image, and then performing batch normalization and activation function processing (i.e., the CBS operation in the figure) to obtain a preliminary haze feature map F1, the size of F1 is 256×256 pixels and 32 channels; inputting F1 into a multi-scale fusion large convolution kernel enhancement module MSE-LKC to obtain a generated feature map F2, the size of which is 256×256 pixels and 32 channels; inputting F2 into a parallel enhanced attention module PEA to obtain a feature map F3, the size of which is 256×256 pixels and 32 channels; then inputting F1 and F3 into a feature fusion module Fusion, and integrating information across paths to obtain an enhanced feature map F4, the size of which is 256×256 pixels and 64 channels;

[0170] In the MSE-LKC module, the input feature map F1 is batch normalized (BatchNorm), and then convolved with a 7×7 convolution kernel (Conv7×7) to extract preliminary features; then convolved with a 5×5 convolution kernel (Conv5) to generate the first-stage feature map F1_3, which is 256×256 pixels and 32 channels; F1_3 is input into the deep dilated convolution module with a dilation rate of 7×7 (i.e. Figure 4 DWDConv7 in ), a deep dilated convolution module with a dilation rate of 13×13 (i.e. Figure 4 DWDConv13 in ), a deep dilated convolution module with a dilation rate of 19×19 (i.e. Figure 4 DWDConv19 in the extract multi-scale features to generate three corresponding feature sub-maps F1_4, F1_5 and F1_6, the size of these sub-maps are 256×256 pixels and 32 channels; use the splicing operation (C) to splice these sub-maps in the channel dimension to obtain the fused feature map F1_7, whose size is 256×256 pixels and 96 channels; the fused feature map F1_7 is input into a convolution kernel of size 1×1 for convolution operation, and then processed by GELU activation function, and then convolved again by a 1×1 convolution kernel to generate an output feature map F1_10, whose size is 256×256 pixels and 32 channels; at the same time, through the residual connection (ResidualConnection), the initial input feature map F1 is added to F1_10 to obtain the feature map F2, whose size is 256×256 pixels and 32 channels;

[0171] In the PEA module, firstly, batch normalization operation (BatchNorm) is performed on F2 to generate the normalized feature map F2_1; the feature map F2_1 is input into the pixel-level attention module for processing, and a convolution kernel of size 1×1 is used to convolve F2_1 (Conv1×1), the number of channels is adjusted, and the intermediate feature map is generated. After that, the intermediate feature map is nonlinearly transformed through the GELU activation function to generate the feature map F2_1_1; a 1×1 convolution kernel is used to convolve F2_1_1 again to generate the intermediate feature map F2_1_2, and it is normalized through the Sigmoid activation function to generate the pixel-level attention weight; the input feature map F2_ 1 and the pixel attention weight are multiplied element by element (×) to generate the enhanced feature sub-graph F2_2; the feature graph F2_1 is input into the channel-level attention module for processing, firstly, F2_1 is subjected to global average pooling (GAP) to generate global information in the channel dimension, and then a convolution kernel of size 1×1 is used to convolve the pooling result to obtain the feature graph T1, and the intermediate feature graph T2 is generated by the GELU activation function; the feature graph T2 is processed with the intermediate features by the 1×1 convolution kernel, and the channel attention weight is generated by the Sigmoid activation function; the channel attention weight is multiplied element by element with F2_1 in the channel dimension (×) to generate the feature sub-graph F2_3 after weight adjustment; The feature map F2_1 is input into the simple pixel attention module for processing. A convolution kernel of size 1×1 is used to convolve F2_1 to generate a preliminary pixel feature map X1. A convolution kernel of size 3×3 is used to further extract local pixel correlations from the feature map X1 to generate a local feature map X2. Through another branch, a convolution kernel of 1×1 is used to directly convolve F2_1 to obtain a feature map X3. After adding the results of feature maps X2 and X3, a simple pixel attention weight is generated through the Sigmoid activation function. The simple pixel attention weight is multiplied element-by-element by F2_1 (×) to generate a feature sub-map F2_4. F2_2, F2_3, F2_4 is concatenated (Concat operation) in the channel dimension to generate a fused feature map F2_5, which is 256×256 pixels and 96 channels in size; F2_5 is subjected to a 1×1 convolution operation (Conv1×1), and after the dimension is reduced to 32 channels, it is subjected to nonlinear processing through the GELU activation function to generate a feature map F2_6; F2_6 is convolved again through a 1×1 convolution kernel to generate a feature map F2_7; the initial input feature map F2 is added to F2_7 through a residual connection (ResidualConnection) to generate the final output feature map F3, which is 256×256 pixels and 32 channels in size;

[0172] In the fine feature capture multi-channel convolution module, the input feature map F4 is first convolved through a convolution kernel (Conv1×1) of size 1×1 to generate a feature map F4_1 of size 256×256 pixels and 32 channels; then, F4_1 is feature split (Split) into three feature processing paths: the first path: F4_1 is convolved through two convolution kernels (Conv3×3) of size 3×3 in sequence to generate a feature map F4_2 of size 256×256 pixels and 32 channels; the second path: F4_1 is convolved through a convolution kernel of size 3×3 and a convolution kernel of size 5×5 (Conv5×5) in sequence to generate a feature map F4_3 of size 256×256 pixels and 32 channels; the third path: Path: Convolve F4_1 directly through a 5×5 convolution kernel (Conv5×5) to generate a feature map F4_4 with a size of 256×256 pixels and 32 channels; concatenate the output feature maps F4_2, F4_3 and F4_4 of the three paths in the channel dimension (C) to obtain a fused feature map F4_5 with a size of 256×256 pixels and 160 channels; then, convolve F4_5 through a 1×1 convolution kernel (Conv1×1) for channel integration to generate a feature map F4_6 with a size of 256×256 pixels and 64 channels; finally, input F4_6 into a 3×3 convolution kernel (Conv3×3) to generate the final output feature map F5 with a size of 256×256 pixels and 64 channels;

[0173] In the multi-scale feature integration enhanced detection module, a 3×3 deformable convolution (DConv) operation is applied to the feature map F2 to obtain an enhanced feature map F6, which has a size of 256×256 pixels and 32 channels. A 3×3 deformable convolution (DConv) operation is applied to the feature map F3 to obtain an enhanced feature map F7, which has a size of 256×256 pixels and 32 channels. The feature maps F5 and F7 are spliced ​​in the channel dimension to form a fused feature map F8, which has a size of 256×256 pixels and 96 channels. A 3×3 convolution kernel is applied to F8 to obtain an enhanced feature map F7. Convolution operation is performed to obtain the optimized feature map F9, which has a size of 256×256 pixels and 64 channels. F9 and F6 are concatenated in the channel dimension to form a fused feature map F10, which has a size of 256×256 pixels and 96 channels. A convolution operation is performed on F10 using a convolution kernel of size 3×3 to obtain the final optimized feature map F11, which has a size of 256×256 pixels and 64 channels. F11 is input into the target detection head (Head), which outputs a tensor containing detection information, where each row corresponds to a detection result, including bounding box coordinates, small target category labels, and confidence scores.

[0174] S4: Based on the constructed FASOD-Net model of small target detection network in haze scenes, train and update the parameters of each layer; first initialize all neural network parameters and set model-related hyperparameters, including training rounds, batch size, weight decay coefficient, learning rate, and total number of iterations;

[0175] S4 includes but is not limited to the following embodiments: training the small target detection network FASOD-Net in the haze scene, the whole training process is set to 100 epochs, the weight decay coefficient is set to 0.01, the learning rate is set to 0.00001, and the batch size is set to 8; the number of iterations of the first 50 epochs is 300, the number of iterations of the next 50 epochs is 400, and the total number of iterations is 35000 times; in each round of training, Mosaic data enhancement is turned on, and the brightness and contrast of the image are randomly cropped, rotated, scaled, and adjusted to enhance the generalization ability of the model; the resolution of the input image is uniformly adjusted to a format of 512×512 to meet the network input requirements; after initializing the parameters, the training set and verification set data of the small target in the haze scene processed by Step 2 are divided into multiple batches, and each batch of data is input into FASOD-Net for training, the training loss value (loss) of the batch is calculated, and the parameters of the model are updated according to the loss value; after completing the training of all batches of data in the entire training set, the verification stage is entered;

[0176] In the verification phase, the verification set data is input into FASOD-Net in batches, and the verification loss value (batch_loss) of each batch is calculated; FASOD-Net dynamically adjusts the learning rate and other optimization parameters according to the loss of each training and the batch_loss value of verification to improve the training efficiency and model performance; the entire training process will continue for multiple rounds of iterations until the verification loss value (batch_loss) tends to converge, indicating that the model has good detection capabilities and has completed training; finally, the trained FASOD-Net can achieve high-precision small target detection in complex haze scenes;

[0177] S5: After the small target detection model for haze scenes is trained, the trained small target detection model for haze scenes is used to detect small targets in haze scene images; finally, the small target detection results in the haze scene are output, including the small target bounding box coordinates, small target category labels and confidence scores.

[0178] Based on the above, the present invention designs a deep learning network FASOD-Net (Fog-Aware Small Object Detection Network) for small target detection in haze scenes;

[0179] (1) In the dehazing and shallow feature extraction module of FASOD-Net, a multi-scale fusion large convolution kernel enhancement module MSE-LKC and a parallel enhanced attention module PEA are designed; the MSE-LKC module extracts multi-scale shallow features through multi-scale hole convolution kernel combined with batch normalization and activation function; the PEA module combines pixel-level attention and channel-level attention mechanisms with simple pixel-level attention to enhance the shallow feature expression ability from the spatial and channel dimensions, effectively remove haze interference and improve the shallow feature capture ability of small targets;

[0180] (2) In the fine feature capture multi-channel convolution module of FASOD-Net, a fine feature capture method combining multi-channel feature diversion and multi-scale convolution kernel is designed; this module enhances the local feature expression ability of small targets through multi-path feature extraction, while optimizing the global semantic expression of the target, effectively improving the detection accuracy of small targets in complex backgrounds;

[0181] (3) In the multi-scale feature integration enhanced detection module of FASOD-Net, a feature integration method combining deformable convolution and multi-path feature fusion is designed; this module optimizes the expression of target boundaries and texture features by effectively integrating features of different scales and levels, significantly improving the accuracy and robustness of small target detection in complex haze scenes.

[0182] Although the present invention has been described above with reference to the embodiments, various modifications may be made thereto and parts thereof may be replaced by equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A small target detection method in haze scene based on deep learning, characterized in that: The specific steps are as follows: S1: Construct an original dataset for small target detection in haze scenes, or use an existing publicly available dataset for small target detection in haze scenes; first collect image data of small targets in haze environments, then use annotation tools to select the positions of small targets and annotate them with labels; S2: preprocessing the original data set of small target detection in haze scenes to obtain a preprocessed data set of small target detection in haze scenes, and dividing the data set to facilitate subsequent training of a small target detection model in haze scenes based on deep learning; S3: Construct a small target detection network FASOD-NET in haze scenes; S4: Based on the constructed FASOD-Net model of small target detection network in haze scenes, train and update the parameters of each layer; first initialize all neural network parameters and set model-related hyperparameters, including training rounds, batch size, weight decay coefficient, learning rate, and total number of iterations; S5: After the small target detection model for haze scenes is trained, the trained small target detection model for haze scenes is used to detect small targets in haze scene images; finally, the small target detection results in the haze scene are output, including the small target bounding box coordinates, small target category labels and confidence scores.

2. According to the method for detecting small targets in haze scenes based on deep learning in claim 1, it is characterized in that: The specific steps of S2 are as follows: S21: The haze scene images in the original dataset of small target detection in haze scenes are cropped to a uniform size. Since the original dataset is small, in order to ensure that the dataset has enough samples for training, verification and testing, and to enhance the robustness of the small target detection model in haze scenes, the sensitivity of the model to image quality is reduced; S22: Then, data enhancement processing is performed to better train the deep learning model, and finally a data-enhanced small target data set in the haze scene is obtained, which is divided into a training set, a validation set, and a test set in a ratio of 8:1:

1.

3. According to the method for detecting small targets in haze scenes based on deep learning in claim 1, it is characterized in that: The specific steps of S3 are as follows: S31: Input the haze scene image into the defogging and shallow feature extraction module; S32: input F4 into the multi-channel convolution module for fine feature capture; S33: Input F2, F3, and F5 into the multi-scale feature integration enhancement detection module.

4. The method for detecting small targets in haze scenes based on deep learning according to claim 3 is characterized in that: The construction and execution process of the defogging and shallow feature extraction module in S31 is as follows: S311: applying a CBS operation including a 3×3 convolution, normalization, and activation function to the input haze scene image to obtain a haze shallow feature map F1; S312: Input F1 to the multi-scale fusion large convolution kernel enhancement module MSE-LKC; S313: Input F2 into the parallel enhanced attention module PEA; S314: Input F3 and F1 into Fusion for channel dimension fusion, and obtain the final output feature map F4 by fusing features at different levels.

5. The method for detecting small targets in haze scenes based on deep learning according to claim 4, characterized in that: The operation and construction process of the MSE-LKC module in S312 is as follows: S3121: Perform batch normalization on the input F1 to obtain a feature map F1_1, then use a convolution kernel of size 7×7 to perform a convolution operation on F1_1 to extract features to obtain a feature map F1_2, and then use a convolution kernel of size 5×5 to perform a convolution operation on F1_2 to obtain a feature map F1_3 of the first stage; S3122: Input F1_3 to a deep dilated convolution module with a dilation rate of 7×7, a deep dilated convolution module with a dilation rate of 13×13, and a deep dilated convolution module with a dilation rate of 19×19 to extract multi-scale features, and generate three corresponding feature sub-graphs F1_4, F1_5, and F1_6 respectively; then, concatenate F1_4, F1_5, and F1_6 in the channel dimension through the C operation to obtain a fused feature graph F1_7; S3123: Input F1_7 into a convolution kernel of size 1×1 for convolution operation to obtain feature map F1_8, process F1_8 through GELU activation function to obtain feature map F1_9, perform Conv1×1 convolution operation on F1_9 to obtain the final output feature map F1_10; S3124: F1 and F1_10 are added through residual connection to form a feature map F2 with stronger representation ability.

6. The method for detecting small targets in haze scenes based on deep learning according to claim 4, characterized in that: The construction and execution process of the PEA module in S313 is as follows: S3131: perform batch normalization on F2 to standardize feature distribution, improve the stability and convergence speed of the model, and obtain the normalized feature map F2_1; S3132: Input F2_1 into three attention modules respectively: S31321: The pixel-level attention module processes the spatial dimension of F2_1, extracts local saliency information, and obtains the enhanced feature sub-graph F2_2; S31322: The channel-level attention module operates on the channel dimension of F2_1 to capture global context information and obtain the weight-adjusted feature subgraph F2_3; S31323: The simple pixel-level attention module uses a simplified mechanism to process F2_1, supplementing the feature extraction capabilities of other modules to obtain the feature sub-graph F2_4; S3133: concatenate F2_2, F2_3 and F2_4 in the channel dimension to obtain a fused feature map F2_5; S3134: Perform a Conv1×1 convolution operation on F2_5 to achieve feature dimensionality reduction, and perform a nonlinear transformation on the feature map through a GELU activation function to obtain a processed feature map F2_6; then perform a Conv1×1 convolution operation on F2_6 to further extract features and adapt to subsequent tasks to obtain a feature map F2_7; S3135: Add F2 and F2_7 through residual connection to obtain the final output feature map F3; Among them, the process of applying PEA to process F2 to obtain F3 is shown in the following formula: F2_1=BatchNorm(F2) F3=F2+Conv1_1(GELU(Conv1_1(C(PA(F2_1),CA(F2_1),SPA(F2_1)))))) Among them, BatchNorm represents batch normalization operation; Conv1_1 represents a convolution operation of size 1×1; GELU represents the application of activation function GELU; SPA represents a simple pixel-level attention module; CA represents a channel-level attention module; PA represents a pixel-level attention module; C represents a splicing operation on the channel dimension; + represents a residual connection operation.

7. The method for detecting small targets in haze scenes based on deep learning according to claim 6, characterized in that: The construction and execution process of the pixel-level attention module in S31321 is as follows: S313211: Perform a Conv1×1 convolution operation on F2_1 to adjust the number of channels of the input feature map; then, apply the activation function GELU to further process F2_1 to obtain the activated feature map F2_1_1; S313212: Perform Conv1×1 convolution operation on F2_1_1 to obtain the intermediate feature map F2_1_2; then, apply the Sigmoid activation function to normalize F2_1_2 to obtain the pixel-level attention weight; S313213: perform element-by-element multiplication of F2_1 and pixel attention weight, weight the feature map, and obtain the final feature map F2_2; Its module can enhance the expression of small target features while suppressing irrelevant background information, thus improving the model's ability to detect small targets in haze scenes. Among them, the process of applying the pixel-level attention module to process F2_1 and obtaining F2_2 is shown in the following formula: F2_2=F2_1×S(Conv1_1(GELU(Conv1_1(F2_1)))) Among them, Conv1_1 represents a convolution operation of size 1×1; GELU represents the application of the activation function GELU; S represents the Sigmoid activation function; × represents an element-by-element multiplication operation; The construction and execution process of the channel-level attention module in S31322 is as follows: S313221: Perform global average pooling on F2_1 to generate global information in the channel dimension; then, perform Conv1×1 convolution operation on the pooling result to obtain feature map T1, and perform nonlinear processing on T1 through the activation function GELU to generate the intermediate expression feature map T2; S313222: After performing a Conv1×1 convolution operation on T2, a channel attention weight is generated through a Sigmoid activation function; F2_1 and the channel attention weight are multiplied element by element to achieve weighted output on the channel, and an enhanced output feature map F2_3 is obtained; Among them, the process of applying the channel-level attention module to process F2_1 to obtain F2_3 is shown in the following formula: F2_3=F2_1×S(Conv1_1(GELU(Conv1_1(GAP(F2_1))))) Among them, GAP represents the global average pooling operation; Conv1_1 represents the convolution operation of size 1×1; GELU represents the application of the activation function GELU; S represents the Sigmoid activation function; × represents the element-by-element multiplication operation; The construction and execution process of the simple pixel-level attention module in S31323 is as follows: S313231: Perform a Conv1×1 convolution operation on F2_1 to obtain a preliminary pixel feature map X1; then, X1 is convolved with another convolution kernel of size 3×3 to further extract local pixel correlation information to obtain a feature map X2; S313232: Perform a Conv1×1 convolution operation directly on F2_1 through another branch to obtain feature map X3; then, add X2 and X3 and generate pixel attention weights through the Sigmoid activation function; S313233: perform element-by-element multiplication of F2_1 and the pixel attention weight to obtain an enhanced output feature map F2_4; Among them, the process of applying a simple pixel-level attention module to process F2_1 to obtain F2_4 is shown in the following formula: F2_4=F2_1×S(Conv3_3(Conv1_1(F2_1))+Conv1_1(F2_1)) Among them, Conv1_1 represents a convolution operation of size 1×1; Conv3_3 represents a convolution operation of size 3×3; S represents the Sigmoid activation function; + represents the addition operation of the results of the two branches; × represents the element-by-element multiplication operation.

8. The method for detecting small targets in haze scenes based on deep learning according to claim 3, characterized in that: The construction and execution process of the fine feature capture multi-channel convolution module in S32 is as follows: S321: Perform a Conv1×1 convolution operation on F4 to adjust the number of channels and extract preliminary features to generate a feature map F4_1; S322: feature splitting of F4_1 into three feature processing paths: in the first path, firstly, two Conv 3×3 convolution operations are performed to extract local deep features and generate feature map F4_2; in the second path, Conv3×3 convolution operations and Conv5×5 convolution operations are performed in sequence to further extract features of a larger receptive field and generate feature map F4_3; in the third path, Conv5×5 convolution operations are performed to further extract features of a larger receptive field and generate feature map F4_4; S323: splice F4_2, F4_3 and F4_4 output from the three paths in the channel dimension to obtain a fused feature map F4_5; perform channel integration on F4_5 through a Conv1×1 convolution operation to generate a final small target deep feature map F4_6; S324: Perform a Conv3×3 convolution operation on F4_6 to further integrate cross-path information and optimize local feature expression to generate feature map F5. Through the above steps, its module combines multi-scale convolution kernels and diversion structures to effectively extract deep semantic features of small targets and enhance the model's detection capability for small targets.

9. The method for detecting small targets in haze scenes based on deep learning according to claim 3, characterized in that: The construction and execution process of the multi-scale feature integration enhanced detection module in S33 is as follows: S331: The input F2 is subjected to a deformable convolution operation to dynamically adapt to the morphological characteristics of small targets and complex background noise in the haze scene to generate an enhanced feature map F6; the input F3 is subjected to a deformable convolution operation to generate an enhanced feature map F7; S332: F5 and F7 are concatenated in the channel dimension to form a fused feature map F8, which enhances the multi-scale expression capability of features and takes into account the unification of details and global features; Subsequently, a Conv 3×3 convolution operation is performed on F8 to further integrate cross-path information and optimize the multi-level expression of features to generate an optimized feature map F9; F9 and F6 are then concatenated in the channel dimension to form a fused feature map F10; S333: Perform a Conv3×3 operation on F10 to further integrate cross-path information and optimize the multi-level expression of features to obtain a feature map F11; input F11 into the target detection head Head, use Head to detect F11, and output a tensor containing detection information, where each row corresponds to a detection result, and the detection result is a small target detection result in a haze scene, including a bounding box coordinate, a small target category label, and a confidence score.

Citation Information

Patent Citations

  • Image defogging method based on adaptive feature fusion

    CN114627002A

  • Hyperspectral image haze removal method and system

    CN117670718A

  • Haze scene driving vision enhancement and target detection method

    CN117974497A

  • Convolutional neural network defogging method based on prior information assistance

    CN118710547A

  • Wheat scab spore target detection method in shielding scene

    CN118823773A

Cited By

  • Power demand prediction method based on deep learning

    CN120764767A

  • Power grid operation scene road foreign matter intrusion detection method based on deep learning

    CN121305467A