Image detection method and device in foggy day scene
By using a multi-path defogging model in foggy scenes, the accuracy of object detection in foggy scenes is solved, and the visibility and recognition of the target object are significantly improved.
Patent Information
- Application Number
- CN202510204639.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The prior art is difficult to achieve accurate target detection in foggy scenes, mainly due to the scattering and absorption of light by fog, which leads to the decrease in image clarity, contrast and color fidelity, which brings great difficulties to target detection.
The multi-path defogging model is adopted to extract the common features/different features of different categories of objects to be detected under different scales through the first propagation path, the contour features and posture features of the objects to be detected through the second propagation path, and the relative position features and spatial position features of the objects to be detected through the third propagation path, and these features are fused to generate global fusion features to remove the target image to be detected.
Through the application of the multi-path defogging model, the visibility and recognition of the target object in foggy scenes are significantly improved, the clarity and detail restoration capabilities of the image are improved, and the model's adaptability to complex foggy scenes is enhanced.
Smart Images

Figure CN120125949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image defogging processing, and particularly to an image detection method and device in a foggy scene. Background Art
[0002] With the wide application of computer vision technology in various fields, the requirements for the accuracy and robustness of object detection are increasing day by day. However, the performance of existing object detection methods is difficult to meet the requirements when facing different scenes, especially foggy scenes. This not only limits the application of object detection technology in bad weather, but also affects the development of related fields such as intelligent transportation and security monitoring.
[0003] In a foggy scene, due to the scattering and absorption of light by fog, the clarity, contrast, and color fidelity of the image are significantly reduced. This makes the contour of the target object blurred and the detailed information lost, bringing great difficulties to object detection. Traditional object detection algorithms are usually trained based on clear images and have poor detection effects in foggy scenes. Although some studies have tried to solve the domain difference problem, in a foggy scene, how to effectively narrow the huge domain gap between normal weather and foggy scenes and achieve accurate object detection is still a key problem to be solved urgently.
[0004] Therefore, there is an urgent need for an image detection method and device in a foggy scene. Summary of the Invention
[0005] This application provides an image detection method and device in a foggy scene, which solves the problem of how to effectively narrow the huge domain gap between normal weather and foggy scenes and achieve accurate object detection in a foggy scene.
[0006] In the first aspect of this application, an image detection method in a foggy scene is provided. The method includes: obtaining a target image to be detected, where the target image to be detected includes at least one object to be detected; inputting the target image to be detected into a multi-path defogging model, where the multi-path defogging model includes a first propagation path, a second propagation path, and a third propagation path; obtaining a first feature in the target image to be detected through the first propagation path, where the first feature is the common / different features of objects to be detected of different categories at different scales; obtaining a second feature in the target image to be detected through the second propagation path, where the second feature includes the contour features and pose features of objects to be detected of different categories; obtaining a third feature in the target image to be detected through the third propagation path, where the third feature includes the relative position features and spatial position features of objects to be detected of different categories; performing feature fusion on the first feature, the second feature, and the third feature, and outputting the global fusion feature corresponding to the object to be detected through the multi-path defogging model; and outputting the defogged target image to be detected through the multi-path defogging model according to the global fusion feature.
[0007] Optionally, the multi-path dehazing model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, and a first residual downsampling layer. The first feature in the target image to be detected is obtained through the first propagation path, specifically including: inputting the target image to be detected into the first feature extraction layer, and outputting a first intermediate feature image through the first feature extraction layer; inputting the target image to be detected into the second feature extraction layer, and outputting a second intermediate feature image through the second feature extraction layer; inputting the first intermediate feature image into the first residual downsampling layer, and outputting a third intermediate feature image through the first residual downsampling layer; inputting the second intermediate feature image and the third intermediate feature image into the third feature extraction layer simultaneously, and outputting the first feature in the target image to be detected through the third feature extraction layer.
[0008] Optionally, the multi-path dehazing model further includes a fourth feature extraction layer and a second residual downsampling layer. The second feature in the target image to be detected is obtained through the second propagation path, specifically including: inputting the second intermediate feature image into the second residual downsampling layer, and outputting a fourth intermediate feature image through the second residual downsampling layer; inputting the first feature and the fourth intermediate feature image into the fourth feature extraction layer simultaneously, and outputting the second feature through the fourth feature extraction layer.
[0009] Optionally, the multi-path dehazing model further includes a fifth feature extraction layer. The third feature in the target image to be detected is obtained through the third propagation path, specifically including: connecting the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer in series according to a preset order; sequentially inputting the target image to be detected through the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer according to a preset order to obtain the third feature in the target image to be detected.
[0010] Optionally, the multi-path dehazing model further includes a batch normalization layer. The method further includes: performing normalization processing on each feature dimension in each intermediate feature image through the batch normalization layer. The expression of the normalization processing is as follows:
[0011]
[0012] where I conv (h, w, c 1 ) represents the intermediate feature image obtained through a preset two-dimensional convolution, that is, the normalized input image. h is the height index in the intermediate feature image, w is the width index in the intermediate feature image, and c 1 is the channel index of the intermediate feature image. K h is the height of the preset two-dimensional convolution, and K w is the width of the preset two-dimensional convolution. I in(h+i, w+j, k) is the pixel value corresponding to the coordinate (h+i, w+j) and the channel k in the intermediate feature image, where i is the local spatial abscissa of the convolution kernel in the preset two-dimensional convolution, and j is the local spatial ordinate of the convolution kernel in the preset two-dimensional convolution. represents the basic element feature map after normalization processing. is the mean value of the image data corresponding to the intermediate feature image in the channel c 1 dimension. is the variance of the image data corresponding to the intermediate feature image in the channel c 1 dimension, and ∈ is a protection parameter used to avoid a denominator of 0. I bn is the intermediate feature image after normalization processing, that is, the normalized output image. is the scaling factor of the image data corresponding to the intermediate feature image in the channel c 1 dimension. The image data corresponding to the intermediate feature image in the channel c 1 dimension offset.
[0013] Optionally, input the first intermediate feature image into the first residual downsampling layer and output the third intermediate feature image through the first residual downsampling layer, which specifically includes: performing a downsampling operation on the first intermediate feature image through a preset convolutional layer, and the downsampling operation is used to reduce the first intermediate feature image to an image of a preset ratio; performing normalization processing on the first intermediate feature image after the downsampling operation through a batch normalization layer; passing through the first residual branch of the first residual downsampling layer and adjusting the channels of the first intermediate feature map after normalization processing based on the channel screening mechanism of information entropy; passing through the second residual branch of the first residual downsampling layer and directly adjusting the channels of the first intermediate feature map based on the channel screening mechanism of information entropy, so that the number of channels in the first intermediate feature map is the same as that in the first intermediate feature map after normalization processing; fusing the first intermediate feature maps with the channels adjusted by the first residual branch and the second residual branch of the first residual downsampling layer respectively, and outputting the third intermediate feature image after feature fusion.
[0014] Optionally, input the second intermediate feature image into the second residual downsampling layer, and output the fourth intermediate feature image through the second residual downsampling layer. Specifically, it includes: performing a pooling operation on the second intermediate feature image through a preset average pooling layer, where the pooling operation is used to reduce the resolution of the second intermediate feature image; performing a normalization process on the second intermediate feature image after the pooling operation through a batch normalization layer; adjusting the channels of the second intermediate feature map after the normalization process through the first residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy; directly adjusting the channels of the second intermediate feature map through the second residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy, so that the number of channels in the second intermediate feature map is the same as that in the second intermediate feature map after the normalization process; fusing the second intermediate feature maps with the channels adjusted by the first residual branch of the second residual downsampling layer and the second residual branch of the second residual downsampling layer respectively, and outputting the fourth intermediate feature image after the feature fusion.
[0015] Optionally, based on the channel screening mechanism of information entropy, perform channel adjustment on the intermediate feature map. Specifically, it includes: obtaining the information entropy value output by the target channel in the intermediate feature map, where the target channel is any one channel in the intermediate feature map; determining whether the information entropy value is less than the preset information entropy value; if the information entropy value is less than or equal to the preset information entropy value, then delete the target channel for channel adjustment.
[0016] In the second aspect of the present application, an image detection device in a foggy scenario is provided. The device includes an acquisition module and a processing module, where,
[0017] The acquisition module is used to acquire the target image to be detected, and the target image to be detected includes at least one object to be detected.
[0018] The cross-scale extraction network module is used to input the target image to be detected into the multi-path dehazing model. The multi-path dehazing model includes a first propagation path, a second propagation path, and a third propagation path; obtain the first feature in the target image to be detected through the first propagation path, where the first feature is the common / different features of different categories of objects to be detected at different scales; obtain the second feature in the target image to be detected through the second propagation path, where the second feature includes the contour features and pose features of different categories of objects to be detected; obtain the third feature in the target image to be detected through the third propagation path, where the third feature includes the relative position features and spatial position features of different categories of objects to be detected.
[0019] The processing module is used to fuse the first feature, the second feature, and the third feature, and output the global fusion feature corresponding to the object to be detected through the multi-path dehazing model; according to the global fusion feature, output the dehazed target image to be detected through the multi-path dehazing model.
[0020] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of the above.
[0021] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to perform the method as described in any one of the above.
[0022] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0023] 1. Obtain the target image to be detected, input the target image to be detected into the multi-path dehazing model, so as to obtain the first feature in the target image to be detected through the first propagation path in the multi-path dehazing model; obtain the second feature in the target image to be detected through the second propagation path in the multi-path dehazing model; obtain the third feature in the target image to be detected through the third propagation path in the multi-path dehazing model; perform feature fusion on the first feature, the second feature, and the third feature, and output the global fusion feature corresponding to the object to be detected through the multi-path dehazing model; according to the global fusion feature, output the dehazed target image to be detected through the multi-path dehazing model, so that through multiple branch operation structures, for the low-contrast, low-illuminance, and low-colorfulness environment in foggy days, it is possible to fully extract the local detail features, global semantic features, and edge contour features in the target image to be detected, realize the efficient fusion of various feature information, thereby improving the adaptability of the model to complex foggy scenes, enhancing the clarity and detail restoration ability of the target image to be detected, and significantly improving the visibility and recognition of the target object.
[0024] 2. By setting two different downsampling layers with residual structures, namely the first residual downsampling layer and the second residual downsampling layer, and effectively extracting and fusing the deep semantic features of the input image through the first residual downsampling layer, further extracting fine-grained feature information through the second residual downsampling layer, and maintaining the structural consistency between features at the same time, so as to realize the efficient fusion of multi-scale features, enhance the detail capture ability and semantic understanding ability of the model for target objects in complex foggy scenes, and finally improve the accuracy and effect of the image dehazing process.
[0025] 3. By obtaining the information entropy value of the output of the target channel in the intermediate feature map, it is determined whether the information entropy value is less than the preset information entropy value. When the information entropy value is less than or equal to the preset information entropy value, the target channel is deleted for channel adjustment, so that the multi-path dehazing model can automatically remove redundant or invalid feature channels, reduce the computational amount, and further improve the computational efficiency and inference speed of the multi-path dehazing model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic flowchart of an image detection method in a foggy scene provided by an embodiment of the present application;
[0027] Figure 2 is a schematic flowchart of the processing of a first residual downsampling layer provided by an embodiment of the present application;
[0028] Figure 3 is a schematic flowchart of the processing of a second residual downsampling layer provided by an embodiment of the present application;
[0029] Figure 4 is an output schematic diagram of a multi-path dehazing model provided by an embodiment of the present application;
[0030] Figure 5 is a module schematic diagram of an image detection device in a foggy scene provided by an embodiment of the present application;
[0031] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0032] Description of the reference numerals: 51, acquisition module; 52, cross-scale extraction network module; 53, processing module; 601, processor; 602, communication bus; 603, user interface; 604, network interface; 605, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0034] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "the", "above", "the above-mentioned", "this" and "this one" are also intended to include the plural forms, unless there is a clear indication to the contrary in the context. It should also be understood that the term " / and / " used in the present application refers to and includes any or all possible combinations of one or more of the listed items.
[0035] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more than two.
[0036] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0037] Please refer to Figure 1 , which shows a schematic flowchart of an image detection method in a foggy scenario provided by an embodiment of the present application. The flowchart mainly includes the following steps: S101 to S107.
[0038] Step S101, obtain a target image to be detected, where the target image to be detected includes at least one object to be detected.
[0039] Specifically, obtain a target image to be detected, where the target image to be detected is an image taken in a specific foggy environment, and the specific foggy environment is a foggy environment with low contrast, low illuminance, and low chromaticity. The target image to be detected includes at least one object to be detected, and the object to be detected includes but is not limited to vehicles, pedestrians, road signs, buildings, or other objects with specific target features.
[0040] Step S102, input the target image to be detected into a multi-path dehazing model, where the multi-path dehazing model includes a first propagation path, a second propagation path, and a third propagation path.
[0041] Specifically, the target image to be detected is input into the multi-path dehazing model. The multi-path dehazing model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, a fourth feature extraction layer, a fifth feature extraction layer, a first residual downsampling layer, and a second residual downsampling layer. Through the combination of the above different feature extraction layers and residual downsampling layers, three different propagation paths are formed, namely the first propagation path, the second propagation path, and the third propagation path. The target image to be detected passes through the above three different propagation paths respectively to perform different feature extractions. The structure of each feature extraction layer is similar. When the target image to be detected is input into the feature extraction layer, it first undergoes a convolution layer operation. The convolution layer is a 3*3 convolution layer, where the convolution kernel size of the convolution layer is 1 and the stride is 1. After passing through this convolution layer, the height and width of the image are reduced by half, reducing the amount of data for subsequent processing and preliminarily extracting the local features of the image. Through the Batch Normalization layer, each feature dimension in each batch of data is normalized to make the data distribution more stable, accelerate network convergence. Then, the sine linear unit is used as the activation function. Assuming the introduction of non-linear factors, the expression ability of the model is increased, the training efficiency is improved, and the generalization ability of the model is enhanced. Through the Batch Normalization layer, each feature dimension in each intermediate feature image is normalized. The expression of the normalization process is as follows:
[0042]
[0043] where I conv (h, w, c 1 ) represents the intermediate feature image obtained through the preset two-dimensional convolution, that is, the normalized input image. h is the height index in the intermediate feature image, w is the width index in the intermediate feature image, c 1 is the channel index of the intermediate feature image, and h ∈ {0, 1, …, H in - 1}, w ∈ {0, 1, …, W in - 1}, c 1 ∈ {0, 1, …, C out1 - 1}, H in is the height of the intermediate feature image, W in is the width of the intermediate feature image, C out1 is the number of channels in the intermediate feature image, K h is the height of the preset two-dimensional convolution, K w is the width of the preset two-dimensional convolution, I in (h + i, w + j, k) is the pixel value corresponding to the coordinates (h + i, w + j) and channel k in the intermediate feature image. i is the local spatial abscissa of the convolution kernel in the preset two-dimensional convolution, and j is the local spatial ordinate of the convolution kernel in the preset two-dimensional convolution. represents the basic element feature map after normalization. The image data corresponding to the intermediate feature image is in channel c 1 The mean value in dimension, The image data corresponding to the intermediate feature image is in channel c 1 The variance in the dimension, ∈ is a protection parameter used to avoid the denominator being 0, I BN is the intermediate feature image after normalization, that is, the normalized output image. The image data corresponding to the intermediate feature image is in channel c 1 The scaling factor in dimension, The image data corresponding to the intermediate feature image is in channel c 1 The offset in the dimension. For using the sinusoidal linear unit as the activation function, it can be expressed by the following formula:
[0044]
[0045] Among them, I act (h,w,c 1 ) is the intermediate feature image after the activation function, and σ(x) is the positive linear unit SiLU formula.
[0046] Step S103, obtaining a first feature in the target image to be detected through the first propagation path, where the first feature is a common feature / difference feature of different categories of objects to be detected at different scales.
[0047] Specifically, the first feature in the target image to be detected is obtained through the first propagation path. The first feature is the common feature / difference feature of different categories of objects to be detected at different scales, that is, the features from different scales and different levels can be fused in the first propagation path. This fusion method can dig out the potential connections between different features and can better identify the common and difference features of objects of different categories at different scales, thereby improving the accuracy of classification. At the same time, the serial connection method in the third propagation path also ensures the continuity and progressiveness of feature extraction, so that the network can gradually transition from shallow features to deep and more abstract features, and optimize the overall feature extraction process. Through its internal convolutional layer, batch normalization, activation function, and residual branch operation, the feature representation is continuously optimized so that it can adapt to the feature extraction requirements of different image contents.
[0048] In a possible implementation, step S103 further includes: inputting the target image to be detected into the first feature extraction layer, and outputting a first intermediate feature image through the first feature extraction layer; inputting the target image to be detected into the second feature extraction layer, and outputting a second intermediate feature image through the second feature extraction layer; inputting the first intermediate feature image into the first residual downsampling layer, and outputting a third intermediate feature image through the first residual downsampling layer; inputting the second intermediate feature image and the third intermediate feature image into the third feature extraction layer simultaneously, and outputting the first feature in the target image to be detected through the third feature extraction layer.
[0049] Specifically, input the target image to be detected into the first feature extraction layer, extract and output a first intermediate feature image; simultaneously input the target image to be detected into the second feature extraction layer, extract and output a second intermediate feature image. Subsequently, input the first intermediate feature image into the first residual downsampling layer, and further extract features and output a third intermediate feature image through a multi-layer convolution and residual connection mechanism. Finally, input the second intermediate feature image and the third intermediate feature image into the third feature extraction layer simultaneously, integrate the feature information extracted from multiple paths, and output the first feature in the target image to be detected, laying a foundation for subsequent target analysis and detection.
[0050] Step S104, obtaining a second feature in the target image to be detected through a second propagation path, where the second feature includes contour features and pose features of objects to be detected of different categories.
[0051] Specifically, obtaining a second feature in the target image to be detected through a second propagation path, where the second feature includes contour features and pose features of objects to be detected of different categories, that is, deeper feature abstraction and fusion can be performed in the second propagation path, so as to extract the category features of the objects, such as distinguishing the contour features of an object cat and the pose features of the object.
[0052] In a possible implementation, step S104 further includes: inputting the second intermediate feature image into the second residual downsampling layer, and outputting a fourth intermediate feature image through the second residual downsampling layer; inputting the first feature and the fourth intermediate feature image into the fourth feature extraction layer simultaneously, and outputting the second feature through the fourth feature extraction layer.
[0053] Specifically, input the second intermediate feature image obtained through the second feature extraction layer into the second residual downsampling layer, use this layer to further reduce the image resolution and fuse and extract features to obtain a fourth intermediate feature image; subsequently, input the first feature obtained through the third feature extraction layer and the fourth intermediate feature image into the fourth feature extraction layer together, and perform deep fusion and feature enhancement of multi-scale features on the two through this layer, and finally output the second feature for further target detection or image analysis operations.
[0054] Step S105: Obtain the third feature in the target image to be detected through the third propagation path. The third feature includes the relative position feature and the spatial position feature of different categories of objects to be detected.
[0055] Specifically, obtaining the third feature in the target image to be detected through the third propagation path means that in the third propagation path, by further mining the features of the target image to be detected, at this time, the spatial relationship and semantic association features between objects are focused on, so as to obtain the relative position feature and the spatial position feature between different objects to be detected.
[0056] In a possible implementation manner, step S105 further includes: connecting the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer in series according to a preset order; sequentially passing the target image to be detected through the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer to obtain the third feature in the target image to be detected.
[0057] Specifically, connect the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer in series in a preset order to form a complete feature extraction network; then, use the target image to be detected as the input and sequentially pass through each feature extraction layer in the network to layer by layer extract and fuse multi-scale and multi-level image features, and finally output the third feature in the target image to be detected in the fifth feature extraction layer, providing a high-quality feature expression for subsequent target detection or classification tasks.
[0058] In a possible implementation manner, step S105 further includes: performing a downsampling operation on the first intermediate feature image through a preset convolutional layer. The sampling operation is used to reduce the first intermediate feature image to an image with a preset ratio; performing a normalization process on the first intermediate feature image after the downsampling operation through a batch normalization layer; passing through the first residual branch of the first residual downsampling layer and performing a channel adjustment on the first intermediate feature map after the normalization process based on the channel screening mechanism of information entropy; passing through the second residual branch of the first residual downsampling layer and directly performing a channel adjustment on the first intermediate feature map based on the channel screening mechanism of information entropy, so that the number of channels in the first intermediate feature map is the same as that in the first intermediate feature map after the normalization process; performing feature fusion on the first intermediate feature maps obtained by respectively performing channel adjustments on the first residual branch and the second residual branch of the first residual downsampling layer, and outputting the third intermediate feature image after feature fusion.
[0059] Specifically, in this application, two different downsampling layers with residual structures are designed to effectively retain the contour and local detail information of foggy images. Please refer to Figure 2 , which shows a schematic diagram of the processing flow of a first residual downsampling layer provided in an embodiment of this application. The processed first intermediate feature image input is downsampled through a preset convolutional layer. The downsampling operation is used to reduce the first intermediate feature image to an image with a preset ratio. The preset convolutional layer is a convolutional layer with a size of 3×3 and a stride of 2. The preset ratio means that both the height and width of the output image are reduced to half of the original first intermediate feature image. This step can reduce the amount of data and expand the receptive field, extracting more abstract and global features. Then, a batch normalization layer is introduced after this convolutional layer, and the first intermediate feature image after the downsampling operation is normalized to stabilize the data distribution and accelerate the network convergence. Then, LeakyReLU is used as the activation function, and its parameter is set to 0.2 to introduce non-linear factors to enhance the expression ability of the multi-path dehazing model. Subsequently, the first intermediate feature image after the above processing is subjected to a residual branch operation and passes through the first residual branch of the first residual downsampling layer and the second residual branch of the first residual downsampling layer respectively. In the first residual branch, the first intermediate feature image first passes through a Conv1×1 convolutional layer to adjust the channel dimension, halving the number of channels, and then passes through a convolutional layer with a size of 3×3 and a stride of 1 for feature extraction and fusion. Finally, it passes through a Conv1×1 convolutional layer to restore the number of channels to the number of channels of the original input. In the second residual branch, the original first intermediate feature image is directly passed through a 1×1 convolutional layer to adjust the channel dimension to match the number of channels after the first branch is processed. Finally, the results of the two branches are fused through the Element-wiseAddition operation to obtain the output image of the first residual downsampling layer, that is, the third intermediate feature image. The expression for downsampling through the first residual downsampling layer is:
[0060]
[0061] where Y1(h,w,c) is the third intermediate feature image, Conv 3×3,s=2 represents downsampling the first intermediate feature image in a way of 3*3 size and stride 2, represents a 1*1 convolutional operation that adjusts the number of channels of the first intermediate feature image from C to C / 2, represents a 1*1 matching dimension that adjusts the number of channels of the first intermediate feature image from C / 2 to C, where
[0062] In a possible implementation, step S105 further includes: performing a pooling operation on the second intermediate feature image through a preset average pooling layer, where the pooling operation is used to reduce the resolution of the second intermediate feature image; performing a normalization process on the second intermediate feature image after the pooling operation through a batch normalization layer; performing a channel adjustment on the second intermediate feature map after the normalization process through the first residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy; directly performing a channel adjustment on the second intermediate feature map through the second residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy, so that the number of channels in the second intermediate feature map is the same as that in the second intermediate feature map after the normalization process; fusing the second intermediate feature maps obtained by respectively performing channel adjustments on the first residual branch of the second residual downsampling layer and the second residual branch of the second residual downsampling layer, and outputting the fourth intermediate feature image after the feature fusion.
[0063] Specifically, in the present application, by designing two different downsampling layers with a residual structure, the contour and local detail information of the foggy image are effectively retained. The second residual downsampling layer has some differences in structure from the first residual downsampling layer. Please refer to Figure 3 , which shows a schematic diagram of the processing flow of a second residual downsampling layer provided in an embodiment of the present application. The second intermediate feature map first enters an average pooling layer with a size of 2×2 and a stride of 2, and the image resolution is quickly reduced through the pooling operation and preliminary feature fusion is performed. Then, a batch normalization layer is also introduced to stabilize the data distribution. The activation function uses ReLU to enhance the non-linear expression ability of the model. In terms of the residual branch operation, the first residual branch first uses a Conv1×1 convolutional layer to expand the number of channels in the second intermediate feature image to twice the original, and then passes through a convolutional layer with a size of 3×3 and a stride of 1 for the extraction and transformation of complex features. Finally, a Conv1×1 convolutional layer is used to restore the number of channels to the original input channel number. The second residual branch first performs a 1×1 convolutional operation on the original second intermediate feature image to adjust the channel dimension to be the same as the final channel number of the first branch. Finally, the results of the two branches are fused through a Concat splicing operation to form the output of the second residual downsampling layer, that is, the fourth intermediate feature image. The expression for downsampling through the second residual downsampling layer is:
[0064]
[0065] where Y2(h, w, c) is the fourth intermediate feature image, and Conv 3×3,s=1 represents downsampling the second intermediate feature image in a manner of 3*3 size and stride 1, A 1×1 convolutional operation that adjusts the number of channels of the second intermediate feature image from C to 2C. It represents a convolution with unchanged number of channels. BN represents the batch normalization layer, AvgPool represents the average pooling operation, and ⊕ represents feature fusion.
[0066] In the processing of the above first residual downsampling layer and second residual downsampling layer: in the first branch operation of the residual, the first intermediate feature image is flattened through a two-dimensional convolution operation of Conv1×1, and then enters n bottleneck layers. The n bottleneck layers are connected in series, which can effectively reduce the number of channels, thereby reducing the complexity of subsequent calculations.
[0067]
[0068] Among them, is the output result after the first intermediate feature image after flattening passes through n bottleneck layers. mid1 is the first intermediate feature image after flattening. O n () means that within each bottleneck layer, except for the first Conv1×1 convolution, the operations are repeated n times on the first intermediate feature image after flattening. In the second branch operation of the residual, the output of the second branch is directly connected to the output of the first branch through a two-dimensional convolution flattening, and they are merged through a Concat splicing fusion operation, and finally a feature map is output through a single Conv1×1. This feature map is the output of the feature extraction layer.
[0069]
[0070] Among them, I out (h, w, c final ) is the final output feature map of the feature extraction layer. W final (i, j, k, c final ) are the weight parameters of this convolution. C out2 is the number of channels in the feature map of the double-branch aggregation. c final is the number of channels in the final output feature map. I concat (h + i, w + j, k) is the pixel value corresponding to the coordinates (h + i, w + j) and channel k in the feature map of the double-branch aggregation.
[0071] In a possible implementation, step S105 further includes: obtaining the information entropy value of the output of the target channel in the intermediate feature map, where the target channel is any channel in the intermediate feature map; determining whether the information entropy value is less than a preset information entropy value; if the information entropy value is less than or equal to the preset information entropy value, then delete the target channel for channel adjustment.
[0072] Specifically, a channel screening mechanism based on information entropy is introduced. In each convolutional layer, the information entropy value of the output feature map of each channel is calculated, and the channels are sorted according to the information entropy size. If the information entropy value is less than or equal to the preset information entropy value, the channels with smaller information entropy are deleted. If the information entropy value is greater than the preset information entropy value, the channels with larger information entropy are retained. Let the number of channels in the original convolutional layer be n, and the number of channels retained after screening be n'. Then the width of the new convolutional layer becomes n' / n times the original. At the same time, in order to make up for the possible information loss caused by channel deletion, a cross-layer feature fusion module is introduced between adjacent convolutional layers. This module fuses the feature information of the deleted channels in the previous layer into the retained channels in the next layer in a weighted summation manner. Let the feature map of the deleted channels in the previous layer be Fd, and the feature map of the retained channels in the next layer be Fr. The fused feature map F = αFr + βFd, where α and β are learnable weight parameters. In this way, while reducing the model width, important feature information is retained as much as possible, improving the lightweight effect and detection accuracy of the model. The learnable weight parameters α and β are updated by the gradient descent algorithm, where τ is the learning rate and Loss is the loss function:
[0073]
[0074] Please refer to Figure 4 , which shows the output schematic diagram of a multi-path dehazing model provided in the embodiments of the present application, Figure 4 and shows the output processes of the first, second, and third propagation paths respectively.
[0075] Step S106: Perform feature fusion on the first feature, the second feature, and the third feature, and output the global fusion feature corresponding to the object to be detected through the multi-path dehazing model.
[0076] Specifically, first, alignment operations are performed on the first feature, the second feature, and the third feature respectively to ensure their consistency in the number of channels and spatial resolution. The alignment operations may include convolution, upsampling, or downsampling, etc. Then, the aligned features are fused according to preset rules, such as using fusion strategies such as feature weighted summation, concatenation, or attention mechanism, so as to fully integrate the feature information from different paths. After fusion, the features not only retain the independent characteristics of different paths but also contain global context information. Finally, through the output layer of the multi-path dehazing model, the fused features are further processed into the global fusion feature corresponding to the object to be detected to support the final object detection or recognition task. This process improves the accuracy and robustness of detecting and recognizing targets in complex scenes.
[0077] Step S107: According to the global fusion feature, output the dehazed image of the object to be detected through the multi-path dehazing model.
[0078] Specifically, using the multi-path dehazing model, the dehazed target image to be detected is generated based on the globally fused features that have been fused, so as to enhance the image clarity and provide a higher-quality image input for the target detection task.
[0079] By adopting the above method, this application obtains the target image to be detected, inputs the target image to be detected into the multi-path dehazing model, and obtains the first feature in the target image to be detected through the first propagation path in the multi-path dehazing model; obtains the second feature in the target image to be detected through the second propagation path in the multi-path dehazing model; obtains the third feature in the target image to be detected through the third propagation path in the multi-path dehazing model; performs feature fusion on the first feature, the second feature, and the third feature, and outputs the globally fused feature corresponding to the object to be detected through the multi-path dehazing model; according to the globally fused feature, outputs the dehazed target image to be detected through the multi-path dehazing model, so that through multiple branch operation structures, for the low-contrast, low-illuminance, and low-colorfulness environment in foggy days, it can fully extract the local detail features, global semantic features, and edge contour features in the target image to be detected, realize the efficient fusion of various feature information, thereby improving the adaptability of the model to complex foggy scenes, enhancing the clarity and detail restoration ability of the target image to be detected, and significantly improving the visibility and recognition of the target object.
[0080] Please refer to Figure 5 , which shows a module schematic diagram of an image detection device in a foggy scene provided by an embodiment of this application. The device includes an acquisition module 51, a cross-scale extraction network module 52, and a processing module 53, where
[0081] The acquisition module 51 is used to acquire the target image to be detected, and the target image to be detected includes at least one object to be detected.
[0082] The cross-scale extraction network module 52 is used to input the target image to be detected into the multi-path dehazing model. The multi-path dehazing model includes a first propagation path, a second propagation path, and a third propagation path; obtains the first feature in the target image to be detected through the first propagation path, and the first feature is the common / different features of different categories of objects to be detected at different scales; obtains the second feature in the target image to be detected through the second propagation path, and the second feature includes the contour features and pose features of different categories of objects to be detected; obtains the third feature in the target image to be detected through the third propagation path, and the third feature includes the relative position features and spatial position features of different categories of objects to be detected.
[0083] A processing module 53 is configured to perform feature fusion on the first feature, the second feature, and the third feature, and output a global fusion feature corresponding to the object to be detected through a multi-path dehazing model; according to the global fusion feature, output a dehazed image of the object to be detected through the multi-path dehazing model.
[0084] In a possible implementation manner, the multi-path dehazing model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, and a first residual downsampling layer. The cross-scale extraction network module 22 is configured to obtain the first feature in the image of the object to be detected through a first propagation path, specifically including: inputting the image of the object to be detected into the first feature extraction layer, and outputting a first intermediate feature image through the first feature extraction layer; inputting the image of the object to be detected into the second feature extraction layer, and outputting a second intermediate feature image through the second feature extraction layer; inputting the first intermediate feature image into the first residual downsampling layer, and outputting a third intermediate feature image through the first residual downsampling layer; inputting the second intermediate feature image and the third intermediate feature image into the third feature extraction layer at the same time, and outputting the first feature in the image of the object to be detected through the third feature extraction layer.
[0085] In a possible implementation manner, the multi-path dehazing model further includes a fourth feature extraction layer and a second residual downsampling layer. The cross-scale extraction network module 52 is configured to obtain the second feature in the image of the object to be detected through a second propagation path, specifically including: inputting the second intermediate feature image into the second residual downsampling layer, and outputting a fourth intermediate feature image through the second residual downsampling layer; inputting the first feature and the fourth intermediate feature image into the fourth feature extraction layer at the same time, and outputting the second feature through the fourth feature extraction layer.
[0086] In a possible implementation manner, the multi-path dehazing model further includes a fifth feature extraction layer. The cross-scale extraction network module 52 is configured to obtain the third feature in the image of the object to be detected through a third propagation path, specifically including: connecting the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer in series in a preset order; sequentially inputting the image of the object to be detected through the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer in a preset order to obtain the third feature in the image of the object to be detected.
[0087] In a possible implementation manner, the multi-path dehazing model further includes a batch normalization layer. The cross-scale extraction network module 52 is configured to perform normalization processing on each feature dimension in each intermediate feature image through the batch normalization layer. The expression of the normalization processing is as follows:
[0088]
[0089] where Iconv (h, w, c 1 ) represents the intermediate feature image obtained through a preset two-dimensional convolution, i.e., the normalized input image. Here, h is the height index in the intermediate feature image, w is the width index in the intermediate feature image, and c 1 is the channel index of the intermediate feature image, and K h is the height of the preset two-dimensional convolution, and K w is the width of the preset two-dimensional convolution, and I in (h + i, w + j, k) is the pixel value corresponding to the coordinates (h + i, w + j) and channel k in the intermediate feature image. Here, i is the local spatial abscissa of the convolution kernel in the preset two-dimensional convolution, and j is the local spatial ordinate of the convolution kernel in the preset two-dimensional convolution. represents the basic element feature map after normalization processing. is the mean value of the image data corresponding to the intermediate feature image in the channel c 1 dimension. is the variance of the image data corresponding to the intermediate feature image in the channel c 1 dimension. ∈ is a protection parameter used to avoid a zero denominator, and I bn is the intermediate feature image after normalization processing, i.e., the normalized output image. is the scaling factor of the image data corresponding to the intermediate feature image in the channel c 1 dimension. The offset of the image data corresponding to the intermediate feature image in the channel c 1 dimension.
[0090] In a possible implementation manner, the cross-scale extraction network module 52 is used to input the first intermediate feature image into the first residual downsampling layer and output the third intermediate feature image through the first residual downsampling layer. Specifically, it includes: performing a downsampling operation on the first intermediate feature image through a preset convolutional layer, and the sampling operation is used to reduce the first intermediate feature image to an image with a preset ratio; performing normalization processing on the first intermediate feature image after the downsampling operation through a batch normalization layer; adjusting the channels of the first intermediate feature map after normalization processing through the first residual branch of the first residual downsampling layer and based on the channel screening mechanism of information entropy; directly adjusting the channels of the first intermediate feature map through the second residual branch of the first residual downsampling layer and based on the channel screening mechanism of information entropy, so that the number of channels in the first intermediate feature map is the same as that in the first intermediate feature map after normalization processing; fusing the first intermediate feature maps with their channels adjusted by the first residual branch of the first residual downsampling layer and the second residual branch of the first residual downsampling layer respectively, and outputting the third intermediate feature image after feature fusion.
[0091] In a possible implementation, the cross-scale extraction network module 52 is used to input the second intermediate feature image into the second residual downsampling layer and output the fourth intermediate feature image through the second residual downsampling layer, which specifically includes: performing a pooling operation on the second intermediate feature image through a preset average pooling layer, where the pooling operation is used to reduce the resolution of the second intermediate feature image; performing a normalization process on the second intermediate feature image after the pooling operation through a batch normalization layer; adjusting the channels of the second intermediate feature map after normalization through the first residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy; directly adjusting the channels of the second intermediate feature map through the second residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy, so that the number of channels in the second intermediate feature map is the same as that in the second intermediate feature map after normalization; fusing the second intermediate feature maps with the channels adjusted by the first residual branch and the second residual branch of the second residual downsampling layer respectively, and outputting the fourth intermediate feature image after feature fusion.
[0092] In a possible implementation, the cross-scale extraction network module 52 is used to adjust the channels of the intermediate feature map based on the channel screening mechanism of information entropy, which specifically includes: obtaining the information entropy value output by the target channel in the intermediate feature map, where the target channel is any one channel in the intermediate feature map; determining whether the information entropy value is less than a preset information entropy value; if the information entropy value is less than or equal to the preset information entropy value, deleting the target channel for channel adjustment.
[0093] It should be noted that: when the device provided in the above embodiment realizes its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.
[0094] This application also provides an electronic device. Refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided in an embodiment of this application. The electronic device may include: at least one processor 601, at least one communication bus 602, a user interface 603, at least one network interface 604, and a memory 605.
[0095] Among them, the communication bus 602 is used to realize the connection and communication between these components.
[0096] Among them, the user interface 603 may include a display screen and a camera. Optionally, the user interface 603 may further include a standard wired interface and a wireless interface.
[0097] Among them, the network interface 604 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0098] Among them, the processor 601 may include one or more processing cores. The processor 601 connects various parts within the entire server through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 605, and by calling data stored in the memory 605, it executes various functions of the server and processes data. Optionally, the processor 601 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 601 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 601 and may be implemented separately by a single chip.
[0099] Among them, the memory 605 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 605 includes a non-transitory computer-readable storage medium. The memory 605 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 605 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments. Optionally, the memory 605 may further be at least one storage device located far from the aforementioned processor 601. Refer to Figure 6In the memory 605, which is a computer storage medium, an operating system, a network communication module, a user interface module, and an image detection application program in a foggy day scenario may be included.
[0100] In Figure 6 In the electronic device shown, the user interface 603 is mainly used to provide an input interface for the user to obtain the data input by the user; and the processor 601 can be used to call the image detection application program stored in the memory 605. When executed by one or more processors 601, the electronic device is caused to execute one or more of the methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0101] This application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, the electronic device is caused to execute one or more of the methods as described in the above embodiments.
[0102] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0103] In several implementation manners provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical or other form.
[0104] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0105] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, can exist physically alone for each unit, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0106] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0107] The above are only exemplary embodiments disclosed in the present application and cannot be used to limit the scope of the present application. That is, any equivalent changes and modifications made in accordance with the teachings disclosed in the present application still fall within the scope covered by the present application. After considering the specification and the disclosure of the practical truth, those skilled in the art will easily think of other implementation schemes of the present application.
[0108] The present application aims to cover any variations, uses, or adaptive changes of the present application, and these variations, uses, or adaptive changes follow the general principles disclosed in the present application and include common general knowledge or conventional technical means in the technical field not recorded in the present application.
Claims
1. An image detection method in a foggy scene, characterized in that: The method comprises: Acquire a target image to be detected, wherein the target image to be detected includes at least one object to be detected; Inputting the target image to be detected into a multi-path defogging model, wherein the multi-path defogging model includes a first propagation path, a second propagation path, and a third propagation path; Acquire a first feature in the target image to be detected through the first propagation path, where the first feature is a common feature / difference feature of the objects to be detected of different categories at different scales; Acquire a second feature in the target image to be detected through the second propagation path, wherein the second feature includes contour features and posture features of different categories of the target image to be detected; Acquire a third feature in the target image to be detected through the third propagation path, wherein the third feature includes relative position features and spatial position features of different categories of the objects to be detected; Performing feature fusion on the first feature, the second feature and the third feature, and outputting a global fusion feature corresponding to the object to be detected through the multipath defogging model; According to the global fusion feature, the defogging image of the target to be detected is outputted through the multi-path defogging model.
2. The method according to claim 1, characterized in that The multipath defogging model includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer and a first residual downsampling layer, and the first feature in the target image to be detected is obtained through the first propagation path, specifically including: Inputting the target image to be detected into the first feature extraction layer, and outputting a first intermediate feature image through the first feature extraction layer; Inputting the target image to be detected into the second feature extraction layer, and outputting a second intermediate feature image through the second feature extraction layer; Inputting the first intermediate feature image into the first residual downsampling layer, and outputting a third intermediate feature image through the first residual downsampling layer; The second intermediate feature image and the third intermediate feature image are simultaneously input into the third feature extraction layer, and the first feature in the target image to be detected is output through the third feature extraction layer.
3. The method according to claim 2, characterized in that The multipath defogging model further includes a fourth feature extraction layer and a second residual downsampling layer, and the obtaining of the second feature in the target image to be detected through the second propagation path specifically includes: Inputting the second intermediate feature image into the second residual downsampling layer, and outputting a fourth intermediate feature image through the second residual downsampling layer; The first feature and the fourth intermediate feature image are simultaneously input into the fourth feature extraction layer, and the second feature is output through the fourth feature extraction layer.
4. The method according to claim 3, characterized in that The multipath defogging model further includes a fifth feature extraction layer, wherein obtaining the third feature in the target image to be detected through the third propagation path specifically includes: Connecting the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer, and the fifth feature extraction layer in series according to a preset order; According to the preset order, the target image to be detected is sequentially passed through the first feature extraction layer, the second feature extraction layer, the third feature extraction layer, the fourth feature extraction layer and the fifth feature extraction layer to obtain the third feature in the target image to be detected.
5. The method according to claim 2, characterized in that: The multipath defogging model further includes a batch normalization layer, and the method further includes: Through the batch normalization layer, each feature dimension in each intermediate feature image is normalized, and the expression of the normalization is as follows: Among them, I conv (h, w, c1) represents the intermediate feature image obtained by the preset two-dimensional convolution, that is, the normalized input image, h is the height index in the intermediate feature image, w is the width index in the intermediate feature image, c1 is the channel index of the intermediate feature image, K h is the height of the preset two-dimensional convolution, K w The preset two-dimensional convolution width, I in (h+i,w+j,k) is the pixel value corresponding to the coordinate (h+i,w+j) and channel k in the intermediate feature image, i is the local spatial abscissa of the convolution kernel in the preset two-dimensional convolution, j is the local spatial ordinate of the convolution kernel in the preset two-dimensional convolution, represents the basic element feature map after the normalization process, is the mean value of the image data corresponding to the intermediate feature image in the channel c1 dimension, is the variance of the image data corresponding to the intermediate feature image in the channel c1 dimension, ∈ is a protection parameter used to avoid the denominator being 0, I bn is the intermediate feature image after the normalization process, that is, the normalized output image, is the scaling factor of the image data corresponding to the intermediate feature image in the channel c1 dimension, The offset of the image data corresponding to the intermediate feature image in the channel c1 dimension.
6. The method according to claim 5, characterized in that The step of inputting the first intermediate feature image into the first residual downsampling layer, and outputting the third intermediate feature image through the first residual downsampling layer specifically includes: Performing a downsampling operation on the first intermediate feature image through a preset convolutional layer, wherein the downsampling operation is used to reduce the first intermediate feature image to an image of a preset scale; Normalizing the first intermediate feature image after the downsampling operation by the batch normalization layer; Performing channel adjustment on the first intermediate feature map after normalization through the first residual branch of the first residual downsampling layer and based on a channel screening mechanism of information entropy; Directly performing channel adjustment on the first intermediate feature map through the second residual branch of the first residual downsampling layer and based on the channel screening mechanism of information entropy, so that the number of channels in the first intermediate feature map is consistent with the number of channels in the first intermediate feature map after normalization; The first residual branch of the first residual downsampling layer and the second residual branch of the first residual downsampling layer are respectively subjected to the first intermediate feature map after the channel adjustment for feature fusion, and the third intermediate feature image after feature fusion is output.
7. The method according to claim 6, characterized in that Inputting the second intermediate feature image into the second residual downsampling layer, and outputting the fourth intermediate feature image through the second residual downsampling layer specifically includes: Performing a pooling operation on the second intermediate feature image through a preset average pooling layer, wherein the pooling operation is used to reduce a resolution of the second intermediate feature image; Normalizing the second intermediate feature image after the pooling operation by the batch normalization layer; Performing the channel adjustment on the second intermediate feature map after the normalization processing through the first residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy; Directly performing the channel adjustment on the second intermediate feature map through the second residual branch of the second residual downsampling layer and based on the channel screening mechanism of information entropy, so that the number of channels in the second intermediate feature map is consistent with the number of channels in the second intermediate feature map after normalization; The first residual branch of the second residual downsampling layer and the second residual branch of the second residual downsampling layer are respectively subjected to the second intermediate feature map after the channel adjustment for feature fusion, and the fourth intermediate feature image after feature fusion is output.
8. The method according to claim 7, characterized in that Based on the channel screening mechanism of information entropy, the intermediate feature map performs the channel adjustment, which specifically includes: Obtaining an information entropy value output by a target channel in the intermediate feature map, where the target channel is any channel in the intermediate feature map; Determining whether the information entropy value is less than a preset information entropy value; If the information entropy value is less than or equal to the preset information entropy value, the target channel is deleted to perform the channel adjustment.
9. An image detection device in a foggy scene, characterized in that: The device includes an acquisition module, a cross-scale extraction network module and a processing module, wherein: The acquisition module is used to acquire a target image to be detected, wherein the target image to be detected includes at least one object to be detected; The cross-scale extraction network module is used to input the target image to be detected into a multi-path defogging model, and the multi-path defogging model includes a first propagation path, a second propagation path, and a third propagation path; obtain a first feature in the target image to be detected through the first propagation path, and the first feature is a common feature / difference feature of the objects to be detected of different categories at different scales; obtain a second feature in the target image to be detected through the second propagation path, and the second feature includes contour features and posture features of the objects to be detected of different categories; obtain a third feature in the target image to be detected through the third propagation path, and the third feature includes relative position features and spatial position features of the objects to be detected of different categories; The processing module is used to perform feature fusion on the first feature, the second feature and the third feature, and output the global fusion feature corresponding to the object to be detected through the multi-path defogging model; based on the global fusion feature, the defogging image of the target to be detected is output through the multi-path defogging model.
10. An electronic device, characterized in that: It includes a processor, a communication bus, a user interface, a network interface and a memory, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Haze scene target detection method based on similarity fusion attention mechanism
CN118397251A
Non-uniform real foggy day image defogging method, device, equipment, medium and product
CN118864312A
Low-illumination image defogging method and device
CN119151809A
Multi-scale fusion defogging method based on stacked hourglass network
US20240062347A1