Image defogging method, device and equipment based on cross fusion
Through the cross-fusion image defogging method, multi-level convolution operation and feature fusion are used to solve the limitations of relying on prior knowledge in the existing technology, efficient and accurate image defogging processing is achieved, and the original image details are retained and the scope of application is expanded.
Patent Information
- Application Number
- CN202510236742.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-25
AI Technical Summary
Existing image defogging methods rely on prior knowledge, resulting in limitations in processing various blurred images, making it difficult to recover clear images efficiently and accurately.
Using the image defogging method based on cross-fusion, multi-level feature images are obtained by performing N first convolution operations on the initial image, N-1 second convolution operations are performed in combination with the first feature fusion image and the initial image, and finally the initial image and the target feature image are added to realize the defogging treatment of the lightweight network.
It reduces the calculation and storage overhead of the model, improves the efficiency and accuracy of defog treatment, while maximizing the original image details, and enhances the wide range of applicability of the model.
Smart Images

Figure CN120374446A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing, and particularly relates to an image dehazing method, device, and equipment based on cross-fusion. Background Art
[0002] Image dehazing has a wide range of application fields, especially in the field of autonomous driving vehicles. Image dehazing is an important issue in the field of computer vision, and its goal is to recover clear and real images from haze-affected images. Traditional dehazing methods usually rely on physical models or statistical assumptions, but these methods often depend on prior knowledge and have certain limitations when dealing with various blurred images. Summary of the Invention
[0003] This application provides an image dehazing method, device, and equipment based on cross-fusion, which is used to reduce the computational and storage overhead of the model, improve the efficiency and accuracy of dehazing processing, while maximizing the retention of the details of the original image and enhancing the generality of the model's applicability.
[0004] This application provides an image dehazing method based on cross-fusion, including: Obtain an initial image; Perform N first convolution operations on the initial image to obtain N initial feature images, where the Kth initial feature image among the N initial feature images indicates that the initial image has continuously undergone K first convolution operations; Obtain a first feature fusion image according to the N initial feature images; Perform N - 1 second convolution operations on the first feature fusion image and the initial image to obtain a target feature image; Add the initial image and the target feature image to obtain a target image.
[0005] According to the cross - fusion - based image defogging method provided by the present application, the step of performing N - 1 second convolution operations on the first feature fusion image and the initial image to obtain a target feature image includes: splicing the first feature fusion image, the Nth initial feature image, the (N - 1)th initial feature image, and the initial image to obtain a second feature fusion image; performing the first first - convolution operation on the second feature fusion image to obtain a first reference feature image; performing N - 2 second convolution operations according to the first feature fusion image, the N initial feature images, the first reference feature image, and the initial image to obtain the target feature image, including: Step 1, splicing the (L - 1)th reference feature image, the first feature fusion image, the (N - L)th initial feature image, and the initial image to obtain a third feature fusion image, where L is an integer greater than or equal to 2 and less than or equal to N - 1; Step 2, performing the Lth first - convolution operation on the third fusion image to obtain the Lth reference feature image; increasing the value of L in sequence and repeating Step 1 - Step 2 until the (N - 1)th reference feature image is obtained; determining the target feature image according to the (N - 1)th reference feature image. According to the cross - fusion - based image defogging method provided by the present application, the step of obtaining a target feature image after performing M first - convolution operations on the third feature fusion image includes: obtaining the reference feature image output after the Mth first - convolution operation; splicing the reference feature image output after the Mth first - convolution operation, the first feature fusion image, and the initial image to obtain the third feature fusion image; performing a third convolution operation on the third feature fusion image to obtain the target feature image.
[0006] According to the cross - fusion - based image defogging method provided by the present application, the step of determining the target feature image according to the (N - 1)th reference feature image includes: splicing the (N - 1)th reference feature image, the first feature fusion image, and the initial image to obtain a fourth feature fusion image; performing a third convolution operation on the fourth feature fusion image to obtain the target feature image.
[0007] According to the cross - fusion - based image defogging method provided by the present application, the third convolution operation includes: calculating the input data through a first convolution kernel.
[0008] According to the cross - fusion - based image de - fogging method provided by the present application, before performing N first convolution operations on the initial image, the method further includes: calculating the initial image data by using a first convolution kernel on the initial image; performing normalization processing on the initial image data to obtain reference image data; processing the reference image data by using a preset activation function to obtain target image data; the performing N first convolution operations on the initial image includes: performing N first convolution operations according to the target image data.
[0009] According to the cross - fusion - based image de - fogging method provided by the present application, the first convolution operation includes: processing input data by using a second convolution kernel to obtain first output data; performing normalization processing on the first output data to obtain second output data; processing the second output data by using a preset activation function to obtain the output data of the current first convolution operation, and the input data of the next first convolution operation includes the output data of the current first convolution operation.
[0010] The present application also provides a cross - fusion - based image de - fogging device, including: A first acquisition unit, configured to acquire an initial image; A first convolution unit, configured to perform N first convolution operations on the initial image to obtain N initial feature images, and K of the N initial feature images indicate that the initial image has continuously undergone K first convolution operations; A second acquisition unit, configured to acquire a first feature fusion image according to the N initial feature images; A second convolution unit, configured to perform N - 1 second convolution operations on the first feature fusion image and the initial image to obtain a target feature image; A third acquisition unit, configured to add the initial image and the target feature image to obtain a target image.
[0011] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the cross - fusion - based image de - fogging method as described above is implemented.
[0012] The present application also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the cross - fusion - based image de - fogging method as described above is implemented.
[0013] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the cross - fusion - based image de - fogging method as described above is implemented.
[0014] The image dehazing method, device and equipment based on cross-fusion provided by this application obtain N initial feature images by performing N first convolution operations on the initial image, then obtain the first feature fusion image based on the N initial feature images, and then perform N-1 second convolution operations on the first initial feature fusion image and the initial image to obtain the target feature image. Finally, the initial image and the target feature image are added to obtain the target image. In this way, this solution adopts a lightweight network to quickly dehaze images based on relatively small parameters. It can not only reduce the computational and storage overhead of the model, improve the efficiency and accuracy of the dehazing process, but also maximize the retention of the details of the original image and enhance the generality of the model's applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic flowchart of an image dehazing method based on cross-fusion provided by this application.
[0017] Figure 2 It is a network structure diagram provided by this application.
[0018] Figure 3 It is one of the test result comparison diagrams provided by this application.
[0019] Figure 4 It is the second of the test result comparison diagrams provided by this application.
[0020] Figure 5 It is a schematic structural diagram of an image dehazing device based on cross-fusion provided by this application.
[0021] Figure 6 It is a schematic structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0023] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0024] Reference to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0025] Currently, when performing haze removal processing on an image, it is necessary to rely on prior knowledge, which makes it limited when processing various blurred images.
[0026] In view of the above problems, the embodiments of this application provide an image haze removal method, device, and equipment based on cross-fusion. The embodiments of this application will be introduced in detail below with reference to the drawings.
[0027] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an image haze removal method based on cross-fusion provided by this application. The image haze removal method based on cross-fusion includes the following steps.
[0028] S101, obtain an initial image.
[0029] Wherein, the initial image is the image that needs to be subjected to haze removal processing. For example, the initial image is an image containing haze. The initial image usually has three RGB channels.
[0030] S102, perform N first convolution operations on the initial image to obtain N initial feature images. The Kth initial feature image among the N initial feature images indicates that the initial image has continuously undergone K times of the first convolution operation.
[0031] Among them, N is a positive integer. During the N - th first convolution operation, the convolution kernel corresponding to each convolution operation is the same, and the input data for the next convolution operation is the output data of the previous convolution operation. That is, the input data for the first convolution operation is the initial image. After one first convolution operation, the first initial feature corresponding to the first first convolution operation is output. Taking this first initial feature as the input data, the second first convolution operation is performed to obtain the second initial feature image corresponding to the second first convolution operation, and so on, until the N - th initial feature image is obtained. Therefore, after N first convolution operations, N initial feature images can be obtained. The first initial feature image indicates that the initial image has undergone one first convolution operation, the second initial feature image indicates that the initial image has continuously undergone two first convolution operations, and the K - th initial feature image indicates that the initial image has continuously undergone K first convolution operations.
[0032] S103. Obtain a first feature fusion image according to the N initial feature images.
[0033] Among them, the number of first convolution operations corresponding to each initial feature image in the N initial feature images is different. Therefore, each initial feature image corresponds to the feature information of the initial image at a certain level. Obtaining the first feature fusion image based on the N initial feature images can enable the cross - fusion of multi - level feature information, enhancing the network's perception ability of details and global information. Then, de - fogging processing is performed based on the first feature fusion image, enabling the network to comprehensively utilize information at different levels, which can help extract and restore the image details of the initial image and improve the accuracy and robustness of the de - fogging effect.
[0034] S104. Perform N - 1 second convolution operations according to the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image.
[0035] Among them, the second convolution operation can be regarded as a process of reconstructing based on the feature image. This second convolution operation is similar to the above - mentioned first convolution operation, that is, the input data for the current second convolution operation includes the output data of the previous second convolution operation. However, the difference is that in the second convolution operation, the input data for the current convolution operation also includes the first feature fusion image, the initial feature image, and the initial image.
[0036] S105. Add the initial image and the target feature image to obtain a target image.
[0037] Among them, the target feature image, which is the feature image restored by the network after the above steps, can be added to the original image based on the residual connection to obtain the target image. The target image is the image output after defogging processing. This can make the output image better retain the structural and detailed information in the original image.
[0038] It can be seen that in this embodiment, by performing N first convolution operations on the initial image, N initial feature images are obtained, and then a first feature fusion image is obtained based on the N initial feature images. Then, N - 1 second convolution operations are performed according to the first initial feature fusion image and the initial image to obtain the target feature image. Finally, the initial image and the target feature image are added to obtain the target image. In this way, this solution uses a lightweight network to quickly perform defogging processing on the image based on fewer parameters. It can not only reduce the computational and storage overhead of the model, improve the efficiency and accuracy of defogging processing, but also maximize the retention of the details of the original image and enhance the generality of the model's applicability.
[0039] In a possible embodiment, the performing N - 1 second convolution operations according to the first feature fusion image, the N initial feature images, and the initial image to obtain the target feature image includes: splicing the first feature fusion image, the Nth initial feature image, the (N - 1)th initial feature image, and the initial image to obtain a second feature fusion image; performing the first first convolution operation on the second feature fusion image to obtain a first reference feature image; according to the first feature fusion image, the N initial feature images, the first reference feature image, and the initial image, performing N - 2 second convolution operations to obtain the target feature image, including: Step 1, splicing the (L - 1)th reference feature image, the first feature fusion image, the (N - L)th initial feature image, and the initial image to obtain a third feature fusion image, where L is an integer greater than or equal to 2 and less than or equal to N - 1; Step 2, performing the Lth first convolution operation on the third fusion image to obtain the Lth reference feature image; sequentially increasing the value of L, repeating Step 1 - Step 2 until the (N - 1)th reference feature image is obtained; determining the target feature image according to the (N - 1)th reference feature image.
[0040] Among them, please refer to Figure 2 , Figure 2 which is the network structure diagram provided by this application. Figure 2 The network structure on the left is a schematic diagram of performing multiple first convolution operations. Figure 2 The network structure on the right is a schematic diagram of performing multiple second convolution operations. When performing the first second convolution operation, the input data is the first feature fusion image (i.e., Figure 2in ①), the Nth initial feature image (i.e., Figure 2 in ③), the (N - 1)th initial feature image (i.e., Figure 2 in ④) and the initial image (i.e., Figure 2 in ②).
[0041] In a specific implementation, the second convolution operation includes two main steps. The first step is a splicing operation, and the second step is to perform the first convolution operation. That is, when performing the first second convolution operation, first splice the first feature fusion image, the Nth initial feature image, the (N - 1)th initial feature image and the initial image to obtain a second feature fusion image. Then, based on the second feature fusion image, perform the first convolution operation to obtain the first reference feature image (i.e., Figure 2 in ⑤).
[0042] In a possible embodiment, the first convolution operation includes: processing the input data through a second convolution kernel to obtain first output data; performing normalization processing on the first output data to obtain second output data; processing the second output data through a preset activation function to obtain the output data of the current first convolution operation, and the input data of the next first convolution operation includes the output data of the current first convolution operation.
[0043] Among them, the second convolution kernel can be a 3×3 convolution kernel. The first convolution operation includes a three-layer structure. The first layer is a convolution layer, that is, calculating the input data, such as the above-mentioned initial image or the second feature fusion image, through a 3×3 convolution kernel to obtain first output data. Then, the first output data is processed through a normalization layer to obtain second output data. In a specific implementation, the normalization method of this normalization layer can be instance normalization (InstanceNorm, IN). In particular, group normalization and other methods can also be used for normalization processing. After the normalization processing, it is then calculated through an activation function to obtain the output data after the current first convolution operation. In a specific implementation, this activation function can be a ReLU activation function.
[0044] It can be seen that in this embodiment, after the input data passes through convolution, normalization, and an activation function, one first convolution operation is completed, and different-level feature information in the image can be extracted, and image texture, edges, and semantic information can be captured, providing rich feature representations for subsequent feature communication. And through the normalization operation, the variation range of feature values can also be reduced, making the training more stable.
[0045] After obtaining the first reference feature image, the second second convolution operation will be performed again. At this time, first the first reference feature image (i.e., Figure 2 in ⑤), the (N - 2)th initial feature image (i.e., Figure 2in ⑥), the first feature fusion image (i.e., Figure 2 in ⑧), and the initial image (i.e., Figure 2 in ⑦) are stitched together to obtain the third feature fusion image. Then, a first convolution operation is performed based on the second reference feature image to obtain the second reference feature image. Then, the third convolution operation, the fourth convolution operation, and the steps of the (N - 1)-th convolution operation are the same as above. It should be noted that Figure 2 ① in Figure 2 indicates the same image content as Figure 2 ② in Figure 2 and ⑦ in
[0046] It can be seen that in this embodiment, when performing image reconstruction based on the fusion feature image, the convolution operation after stitching based on the original image, the feature image extracted by the network, and the fusion image after multi-level feature fusion can improve the image quality and detail retention ability, and improve the defogging effect.
[0047] In a possible embodiment, determining the target feature image according to the (N - 1)-th reference feature image includes: stitching the (N - 1)-th reference feature image, the first feature fusion image, and the initial image to obtain a fourth feature fusion image; performing a third convolution operation on the fourth feature fusion image to obtain the target feature image.
[0048] Wherein, after (N - 1) second convolution operations, a third convolution operation can be performed separately, and then the target feature image is output.
[0049] In a possible embodiment, the third convolution operation includes: calculating the input data through a first convolution kernel.
[0050] Wherein, the first convolution kernel can be a 1×1 convolution kernel. Outputting the fourth feature fusion image after calculation through a 1×1 convolution kernel can adjust the output number of channels and introduce a non-linear transformation, so that the network can learn and represent more complex patterns.
[0051] It can be seen that in this embodiment, after (N - 1) second convolution operations, a third convolution operation is performed based on the low-level feature fusion image to change the number of channels and introduce a non-linear transformation, which can improve the performance of the network without increasing a large amount of computational complexity.
[0052] In a possible embodiment, obtaining the first feature fusion image according to the N initial feature images includes: stitching the N initial feature images to obtain a stitched image; performing a third convolution operation on the stitched image to obtain the first feature fusion image.
[0053] Among them, when obtaining the first feature fusion image, the initial feature images at each different level can be spliced, and then the spliced image is calculated through a 1×1 convolution kernel and then output.
[0054] It can be seen that in this embodiment, after adjusting the number of channels after feature fusion and then outputting, it is convenient to splice the first feature fusion image, the initial image, the initial feature image, etc. during subsequent reconstruction.
[0055] In a possible embodiment, before performing the N first convolution operations on the initial image, the method further includes: calculating the initial image through a first convolution kernel to obtain initial image data; performing normalization processing on the initial image data to obtain reference image data; processing the reference image data through a preset activation function to obtain target image data; the performing the N first convolution operations on the initial image includes: performing the N first convolution operations according to the target image data.
[0056] Among them, before the network extracts the features of the initial image, the initial image can also be first passed through a 1×1 convolution kernel. In this way, the number of channels is adjusted and a non-linear transformation is introduced. Then, after instance normalization operation and ReLU activation function, the reference image data is output. Finally, the reference image data is subjected to multiple first convolution operations to extract the features of the initial image.
[0057] It can be seen that in this embodiment, before feature extraction, feature transformation is first performed at the channel level, which can improve the performance of the network without increasing a large amount of computational complexity.
[0058] Next, the present application will be introduced in detail with reference to Figure 2 the network structure shown below.
[0059] First, the "foggy image" is used as the initial image and input into the network (the network of this solution is named the CrossFsionNet network, and all the networks mentioned in this article refer to this network). First, it passes through the feature extraction layer. That is, the "foggy image" is first passed through a 1×1 convolution kernel, and then through instance normalization and ReLU activation function to obtain target image data. Then, the target image data is used as the input and passes through 4 first convolution operations to obtain 4 initial feature images.
[0060] Then it enters the feature fusion and reconstruction layer. That is, first, the 4 initial feature images are concatenated to obtain a concatenated image. Then, after calculating the concatenated image through a 1×1 convolutional kernel, the first feature fusion image is output. Then, multiple second convolution operations are performed, that is, the fourth initial feature image, the third initial feature image, the first feature fusion image, and the initial image are concatenated to obtain the second feature fusion image. Then, the second feature fusion image is passed through a 3×3 convolutional kernel, instance normalization operation, and ReLU activation function to obtain the first reference feature image. Then, the first reference feature image, the third initial feature image, the first feature fusion image, and the initial image are concatenated to obtain the third feature fusion image. Then, the third feature fusion image is passed through a 3×3 convolutional kernel, instance normalization operation, and ReLU activation function to obtain the second reference feature image. In this way, the third reference feature image is obtained.
[0061] After obtaining the third reference feature image, the first feature fusion image, the initial image, and the third reference feature image are concatenated to obtain the fourth feature fusion image. Then, after calculating the fourth feature fusion image through a 1×1 convolutional kernel, the "dehazed image" is obtained and output.
[0062] In specific implementation, the Mean Squared Error (MSE) loss function can be selected to measure the difference between the dehazed image generated by the network and the clear image. This loss function measures the similarity between them by calculating the square of the difference between each pixel of the two and taking the average. In short, the MSE loss function provides a quantitative way to evaluate the dehazing effect of the network, which is used to measure the pixel-level difference between the generated image and the clear image. A lower MSE value indicates that the difference between the dehazed image generated by the model and the clear image is smaller, that is, the dehazing ability of the model is better. Using the MSE loss function can help us optimize the model to make it generate a dehazing result closer to the clear image.
[0063] Based on the above steps, the network parameter table of the CrossFsionNet network provided by this solution is shown in Table 1 below.
[0064] Table 1 To further illustrate the effectiveness of this solution, the haze removal effect of the CrossFusionNet network provided by this solution can be tested based on a public dataset. 500 images from the Synthetic Object Tracking Scene (SOTS) subset of the publicly available Remote Sensing Image Dehazing Dataset (RESIDE) are used for the experiment, and the average values of the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) metrics for all the photos are calculated. The same dataset is used for training and experimental testing with the Dark Channel Prior (DCP) network, the DehazeNet, and the Atmospheric Object Detection Network (AOD-Net).
[0065] Figure 3 For the experimental results of the outdoor synthetic dataset SOTS, Figure 3 In column (a) of the figure are the hazy images, Figure 3 In column (b) of the figure are the images processed by DCP, Figure 3 In column (c) of the figure are the images processed by DehazeNet, Figure 3 In column (d) of the figure are the images processed by AOD-Net, Figure 3 In column (e) of the figure are the images processed by the CrossFusionNet network of this solution.
[0066] For Figure 3 Some algorithms are selected for comparison of the images in. It can be seen that after haze removal by the DCP algorithm, due to the Dark Channel Prior theory not fitting well with the sky region, the color distortion of the sky area in the haze-removed image is relatively serious, and there is a phenomenon that the overall color is on the darker side. There is an incomplete haze removal situation with DehazeNet. The AOD-Net haze removal algorithm also has the phenomenon of incomplete haze removal, such as Figure 3 in the third and fourth rows of. Such as Figure 3The experimental results of the CrossFusionNet network in this paper are initially similar to those of the AOD-Net model. However, upon closer inspection, the dehazing effect of CrossFusionNet in this solution is slightly brighter in color and light, and has a higher color and light restoration degree compared to the dehazing effect of AOD-Net. Moreover, AOD-Net performs parameter estimation based on the atmospheric scattering model, but there are issues where the atmospheric scattering model does not match in certain scenarios. In contrast, the image dehazing method based on cross-fusion in this solution does not rely on the atmospheric scattering model or any prior knowledge, so it does not have the above problems, fits a wider range of scenarios, and achieves good dehazing results for foggy images in most scenarios.
[0067] Please refer to Table 2. Table 2 shows the quantitative comparison of the image dehazing method based on cross-fusion in this solution with other algorithms on the synthetic fog image test set. As can be seen from Table 2, on the SOTS outdoor dataset, the PSNR and SSIM indicators of the CrossFusionNet model in this solution are the highest. According to Figure 3 It can also be observed that the dehazing effect of CrossFusionNet is slightly brighter in light and slightly more vivid in color compared to the dehazing effects of other algorithms, and has a higher reduction degree compared to the ground truth.
[0068] Table 2 Next, the dehazing effect of this solution will be evaluated based on real datasets.
[0069] Please refer to Figure 4 , 200 images are selected for testing in the Realistic Time-of-Flight Synthetic (RTTS) dataset of the Realistic Single Image Dehazing (RESIDE) real single-image dehazing dataset. By comparing with other algorithms, the effectiveness of the network in this solution can be verified. Figure 4 The pictures in column (a) are foggy images, Figure 4 The pictures in column (b) are the pictures after DCP processing, Figure 4 The pictures in column (c) are the pictures after DehazeNet processing, Figure 4 The pictures in column (d) are the pictures after AOD-Net processing, Figure 4 The pictures in column (e) are the pictures after being processed by the CrossFusionNet network in this solution.
[0070] Figure 4For the experimental results of the real dataset, it can be seen that DCP is severely distorted in the sky area, whether it is a blue sky or a cloudy day. Moreover, the color and light are significantly overly uneven, with obvious discontinuities, not conforming to the natural scene. The clarity is also low, the details of the picture are relatively rough, and the quality of the picture is poor. When the light is dim, the defogging light of DehazeNet becomes darker, losing information in the dark areas, and the defogging effect for distant and thick fog is not good. The defogging effect of AOD-Net is better, but there is also the problem of incomplete defogging, such as Figure 4 The first and second lines. And comparing with the experimental results of the CrossFusionNet network, visually, the clarity of CrossFusionNet is higher, such as Figure 4 The second, third, and fourth lines. Although there is also the problem of incomplete defogging in the distance, the defogging effect is better, and the color is slightly brighter, and the light brightness is better. The overall picture is smoother both in color and light, and the visual effect is clearer.
[0071] Next, the parameters and running time of CrossFusionNet are analyzed.
[0072] The comparison of the model parameters and running time of this solution and other defogging algorithms can be obtained through Table 3.
[0073] Table 3 It can be seen that the CrossFusionNet network corresponding to the image defogging method of this solution is a lightweight model. Compared with the traditional defogging algorithm DCP and the classic lightweight defogging algorithms DehazeNet and AOD-Net, it has the fewest number of parameters, has the smallest inference time when verified and tested on the SOTS dataset, and at the same time achieves the highest PSNR value in the image defogging task. The CrossFusionNet network corresponding to this solution has a lower complexity and can represent effective feature information with fewer parameters. And it has the smallest inference time, with high efficiency in practical applications.
[0074] The above experimental results show that the image defogging method based on cross-fusion provided by this solution has achieved excellent performance in terms of the number of parameters, inference time, and PSNR and SSIM metrics. And the network architecture has a lower complexity and high efficiency, and can provide excellent image defogging quality.
[0075] Next, a description is given of an image defogging device based on cross-fusion provided by this application. The image defogging device based on cross-fusion described below corresponds to and refers to the image defogging method based on cross-fusion described above.
[0076] Please refer to Figure 5 ,Figure 5 It is a schematic structural diagram of an image dehazing device based on cross - fusion provided by this application. The image dehazing device 500 based on cross - fusion includes: a first acquisition unit 501, configured to acquire an initial image; a first convolution unit 502, configured to perform N first convolution operations on the basis of the initial image to obtain N initial feature images, where K of the N initial feature images indicate that the initial image has continuously undergone K first convolution operations; a second acquisition unit 503, configured to acquire a first feature fusion image according to the N initial feature images; a second convolution unit 504, configured to perform N - 1 second convolution operations on the basis of the first feature fusion image and the initial image to obtain a target feature image; a third acquisition unit 505, configured to add the initial image and the target feature image to obtain a target image.
[0077] In a possible embodiment, in terms of performing N - 1 second convolution operations on the basis of the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image, the second convolution operation 504 specifically is: splicing the first feature fusion image, the Nth initial feature image, the (N - 1)th initial feature image, and the initial image to obtain a second feature fusion image; performing the first first convolution operation on the second feature fusion image to obtain a first reference feature image; performing N - 2 second convolution operations on the basis of the first feature fusion image, the N initial feature images, the first reference feature image, and the initial image to obtain the target feature image, including: Step 1, splicing the (L - 1)th reference feature image, the first feature fusion image, the (N - L)th initial feature image, and the initial image to obtain a third feature fusion image, where L is an integer greater than or equal to 2 and less than or equal to N - 1; Step 2, performing the Lth first convolution operation on the third fusion image to obtain the Lth reference feature image; sequentially increasing the value of L, repeating Step 1 - Step 2 until the (N - 1)th reference feature image is obtained; determining the target feature image according to the (N - 1)th reference feature image.
[0078] In a possible embodiment, in terms of determining the target feature image according to the (N - 1)th reference feature image, the convolution unit 504 specifically is: splicing the (N - 1)th reference feature image, the first feature fusion image, and the initial image to obtain a fourth feature fusion image; performing a third convolution operation on the fourth feature fusion image to obtain the target feature image.
[0079] In a possible embodiment, in terms of obtaining the first feature fusion image according to the N initial feature images, the second obtaining unit 503 is specifically configured to: splice the N initial feature images to obtain a spliced image; perform a third convolution operation on the spliced image to obtain the first feature fusion image.
[0080] In a possible embodiment, the third convolution operation includes: calculating the input data through a first convolution kernel.
[0081] In a possible embodiment, before performing the N first convolution operations according to the initial image, the first convolution unit 502 is further configured to: calculate the initial image through a first convolution kernel to obtain initial image data; perform normalization processing on the initial image data to obtain reference image data; process the reference image data through a preset activation function to obtain target image data; in terms of performing the N first convolution operations according to the initial image, the first convolution unit 502 is specifically configured to: perform the N first convolution operations according to the target image data.
[0082] In a possible embodiment, the first convolution operation includes: processing the input data through a second convolution kernel to obtain first output data; performing normalization processing on the first output data to obtain second output data; processing the second output data through a preset activation function to obtain the output data of the current first convolution operation, and the output data of the current first convolution operation is included in the input data of the next first convolution operation.
[0083] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the electronic device provided by this application. As Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete their mutual communication through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute an image dehazing method based on cross-fusion. The method includes: obtaining an initial image; performing N first convolution operations on the basis of the initial image to obtain N initial feature images, where the K-th initial feature image among the N initial feature images indicates that the initial image has continuously undergone K times of the first convolution operation; obtaining a first feature fusion image according to the N initial feature images; performing N - 1 second convolution operations according to the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image; adding the initial image and the target feature image to obtain a target image.
[0084] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0085] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the image dehazing method based on cross-fusion provided by the above-mentioned various methods. The method includes: obtaining an initial image; performing N first convolution operations on the basis of the initial image to obtain N initial feature images, where the K-th initial feature image among the N initial feature images indicates that the initial image has continuously undergone K times of the first convolution operation; obtaining a first feature fusion image according to the N initial feature images; performing N - 1 second convolution operations according to the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image; adding the initial image and the target feature image to obtain a target image.
[0086] In another aspect, the present application also provides a computer program product, including a computer program which, when executed by a processor, implements any of the above cross-fusion based image dehazing methods. The method includes: obtaining an initial image; performing N first convolution operations on the initial image to obtain N initial feature images, where the K-th initial feature image among the N initial feature images indicates that the initial image has continuously undergone K times of the first convolution operation; obtaining a first feature fusion image based on the N initial feature images; performing N-1 second convolution operations on the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image; and adding the initial image and the target feature image to obtain a target image.
[0087] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image dehazing method based on cross - fusion, characterized in that, Including: Obtain an initial image; Perform N first convolution operations on the initial image to obtain N initial feature images, where the K-th initial feature image among the N initial feature images indicates that the initial image has continuously undergone K first convolution operations; Obtain a first feature fusion image according to the N initial feature images; Perform N - 1 second convolution operations according to the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image; Add the initial image and the target feature image to obtain a target image.
2. The method according to claim 1, wherein The step of performing N - 1 second convolution operations according to the first feature fusion image, the N initial feature images, and the initial image to obtain a target feature image includes: Stitch the first feature fusion image, the N-th initial feature image, the (N - 1)-th initial feature image, and the initial image to obtain a second feature fusion image; Perform the first first convolution operation on the second feature fusion image to obtain a first reference feature image; According to the first feature fusion image, the N initial feature images, the first reference feature image, and the initial image, perform N - 2 second convolution operations to obtain the target feature image, including: Step 1, stitch the (L - 1)-th reference feature image, the first feature fusion image, the (N - L)-th initial feature image, and the initial image to obtain a third feature fusion image, where L is an integer greater than or equal to 2 and less than or equal to N - 1; Step 2, perform the L-th first convolution operation on the third fusion image to obtain the L-th reference feature image; Increase the value of L in sequence, and repeat Step 1 - Step 2 until the (N - 1)-th reference feature image is obtained; Determine the target feature image according to the (N - 1)-th reference feature image.
3. The method according to claim 2, wherein The step of determining the target feature image according to the (N - 1)-th reference feature image includes: Stitch the (N - 1)-th reference feature image, the first feature fusion image, and the initial image to obtain a fourth feature fusion image; Perform a third convolution operation on the fourth feature fusion image to obtain the target feature image.
4. The method according to claim 1, characterized in that, The step of obtaining a first feature fusion image according to the N initial feature images includes: Stitch the N initial feature images to obtain a stitched image; Perform a third convolution operation on the stitched image to obtain a first feature fusion image.
5. The method according to claim 3 or 4, characterized in that, The third convolution operation includes: calculating the input data through a first convolution kernel.
6. The method according to claim 1, wherein Before performing N first convolution operations on the initial image, the method further includes: Calculate the initial image through a first convolution kernel to obtain initial image data; Perform normalization processing on the initial image data to obtain reference image data; Process the reference image data through a preset activation function to obtain target image data; The step of performing N first convolution operations on the initial image includes: Perform N first convolution operations according to the target image data.
7. The method according to any one of claims 1-4, characterized in that, The first convolution operation includes: Process the input data through a second convolutional kernel to obtain first output data; Perform normalization processing on the first output data to obtain second output data; Process the second output data through a preset activation function to obtain the output data of the current first convolutional operation, and the output data of the current first convolutional operation is included in the input data of the next first convolutional operation.
8. An image dehazing device based on cross-fusion, characterized in that, Includes: A first acquisition unit for acquiring an initial image; A first convolutional unit for performing N first convolutional operations on the initial image to obtain N initial feature images, and K of the N initial feature images indicate that the initial image has continuously performed K first convolutional operations; A second acquisition unit for acquiring a first feature fusion image according to the N initial feature images; A second convolutional unit for performing N - 1 second convolutional operations on the first feature fusion image and the initial image to obtain a target feature image; A third acquisition unit for adding the initial image and the target feature image to obtain a target image.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the cross - fusion - based image de - fogging method according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the cross - fusion - based image de - fogging method according to any one of claims 1 to 7 are implemented.