Image restoration method and device, medium and equipment

By embedding multi-scale dynamic frequency refinement modules and gated residual blocks in neural networks, the problems of discomfortability and information loss in image recovery are solved, and high-quality image reconstruction is achieved, especially effective recovery of complex degraded images.

CN120339129APending Publication Date: 2025-07-18LISHUI UNIV

Patent Information

Application Number
CN202510418978.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems with unsuitability and information loss in image recovery. Traditional methods lead to excessive smooth reconstruction results, and the self-attention network computing complexity is high and it is difficult to use in high-resolution images.

Method used

A neural network including encoder and decoder is built, and a multi-scale dynamic frequency refinement module and a gated residual block are embedded. The low-frequency and high-frequency characteristics are separated by a multi-scale dynamic frequency refinement module. The gated module controls information flow, enhances non-redundant characteristics, and reduces excessive smoothing.

Benefits of technology

The quality of image recovery is improved, especially for images with spatial transformation degraded cores, and a higher quality reconstruction effect is achieved while reducing the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339129A_ABST
    Figure CN120339129A_ABST
Patent Text Reader

Abstract

The invention discloses an image restoration method and device, a medium and equipment, and relates to the technical field of image processing. In order to solve the problem that useful information is lost in the transmission process of a deep network, an improved gating residual module is provided. Wherein the residual sub-module is embedded with a multi-scale frequency refining module which is provided for solving the problem that a fuzzy reconstructed image is caused by the fact that a traditional convolutional network suppresses high-frequency detail information, and the multi-scale frequency refining module dynamically obtains a low-pass filter from original input features through average pooling and feature folding operation; and low-frequency and high-frequency feature maps are separated through convolution operation, and spliced features rich in high-frequency information are obtained after weighting. According to the module, input features are down-sampled to three different scales through global average pooling operation, after the different scales are subjected to the same frequency refining operation, the features of each scale are restored into the input scale through a bilinear interpolation method, and after elements are added one by one, output features are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image restoration method, apparatus, medium and device. Background Art

[0002] Due to the important role of image restoration in fields such as medical imaging, satellite remote sensing, and unmanned driving, it has received extensive attention from the academic and industrial communities. However, high-quality reconstruction of low-quality images results in multiple solutions, and this ill-posedness makes the solution of the image restoration task extremely challenging. Traditional methods usually use manually crafted priors to constrain the solution space. Existing methods are mainly based on convolutional networks, but their local characteristics and static weights will produce overly smooth results when dealing with degraded images. And for images with a spatially-varying degradation kernel, it is difficult to obtain good reconstruction results by only increasing the width and depth of the convolutional network, because a single convolutional structure will cause information loss in deep networks. Although the self-attention network can model global features and can better learn features, due to its computational complexity growing quadratically with the image size, it is difficult to be used for high-resolution images. Summary of the Invention

[0003] Based on this, in order to solve the technical problems in the prior art, the present invention provides an image restoration method, apparatus, medium and device.

[0004] The present invention provides an image restoration method, including:

[0005] Construct a neural network including an encoder and a decoder, where the encoder includes a convolutional layer and a gated residual block connected in sequence; the gated residual block includes an improved residual module and a gating module; the improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module, and the multi-scale dynamic frequency refinement module includes several parallel frequency branches and a fusion layer connected to the output ends of the several parallel frequency branches at the same time;

[0006] Collect low-quality - high-quality image pairs to construct a dataset, and use the dataset to train the constructed neural network to obtain a restoration model for restoring low-quality images to high-quality images;

[0007] Input the low-quality image to be restored into the restoration model, extract features through the convolutional layers in the encoder to obtain the initial feature map; perform feature extraction of different frequencies on the initial feature map through several parallel frequency branches in the multi-scale dynamic frequency refinement module, and fuse the outputs of each frequency branch through the fusion layer in the multi-scale dynamic frequency refinement module to obtain the multi-frequency feature map; suppress redundant features and enhance non-redundant features in the multi-frequency feature map through the gating module to obtain the encoded feature map; decode the encoded feature map through the decoder to obtain the high-quality image corresponding to the low-quality image to be restored.

[0008] Further, the dynamic frequency refinement layer includes an input layer, a feature extraction branch, a dynamic convolutional layer, and a separation layer connected in parallel to the output end of the input layer, where the feature extraction branch includes a global average pooling layer, a 1×1 convolutional layer, a batch normalization layer, a Softmax activation function layer, and a folding layer connected in sequence;

[0009] The dynamic frequency refinement layer performs feature extraction of different frequencies on the initial feature map, specifically including:

[0010] Compress the feature dimension of the input to the dynamic frequency refinement layer by using the global average pooling layer; expand the dimension of the output of the global average pooling layer through the 1×1 convolutional layer; perform batch normalization operation and activation operation through the batch normalization layer and the Softmax activation function layer in sequence; fold the output of the Softmax activation function layer into a dynamic convolutional kernel through the folding layer;

[0011] Perform a convolution operation on the feature of the input to the dynamic frequency refinement layer with the dynamic convolutional kernel output by the feature extraction branch through the dynamic convolutional layer, and output the low-frequency feature;

[0012] Subtract the low-frequency feature from the feature of the input to the dynamic frequency refinement layer through the separation layer to obtain the high-frequency feature.

[0013] Further, the encoder includes a first encoder, a second encoder, and a third encoder connected in sequence, and the first encoder, the second encoder, and the third encoder have the same structure; among them, the input of the first encoder is the low-quality image to be restored; after performing the first downsampling operation on the low-quality image to be restored, fuse it with the output of the first encoder, and use the fusion result as the input of the second encoder; after performing the second downsampling operation on the low-quality image to be restored, fuse it with the output of the second encoder, and use the fusion result as the input of the third encoder.

[0014] Further, the decoder includes a first decoder, a second decoder, and a third decoder connected in sequence, and the first decoder, the second decoder, and the third decoder have the same structure, and each includes a gating residual block and a convolutional layer connected in sequence;

[0015] Feature decoding operation is performed on the output of the third encoder through a gated residual block, and the obtained result and the output of the third encoder are jointly used as the input of the gated residual block in the first decoder; the output of the gated residual block in the first decoder and the result obtained after the second downsampling operation on the low-quality image to be restored are fused and used as the input of the convolutional layer in the first decoder;

[0016] The output of the gated residual block in the first decoder and the output of the gated residual block in the second encoder are jointly used as the input of the gated residual block in the second decoder; the output of the gated residual block in the second decoder and the result obtained after the first downsampling operation on the low-quality image to be restored are fused and used as the input of the convolutional layer in the second decoder;

[0017] The output of the gated residual block in the second decoder and the output of the gated residual block in the first encoder are used as the input of the gated residual block in the third decoder; the output of the gated residual block in the third decoder and the low-quality image to be restored are fused and used as the input of the convolutional layer in the third decoder.

[0018] Furthermore, the several parallel frequency branches include parallel first, second, and third frequency branches. Each frequency branch includes a global average pooling layer, a dynamic frequency refinement layer, and a bilinear interpolation upsampling layer connected in sequence. Among them, the downsampling multiples of the global pooling layers in the first, second, and third frequency branches are 2, 4, and 8 respectively.

[0019] Furthermore, the improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module. The multi-scale dynamic frequency refinement module is embedded between the two convolutional layers of the original residual module. The obtained improved residual module includes an input layer, a normalization layer, a first 3×3 convolutional layer, a multi-scale dynamic frequency refinement module, and a second 3×3 convolutional layer connected in sequence; among them, the input end of the input layer and the output end of the second 3×3 convolutional layer are connected through a residual connection.

[0020] The present invention provides an image restoration device, including:

[0021] A network construction module for constructing a neural network including an encoder and a decoder. The encoder includes a convolutional layer and a gated residual block connected in sequence; the gated residual block includes an improved residual module and a gating module; the improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module. The multi-scale dynamic frequency refinement module includes several parallel frequency branches and a fusion layer connected to the output ends of the several parallel frequency branches at the same time;

[0022] A network training module, configured to collect low-quality - high-quality image pairs to construct a dataset, and use the dataset to train the constructed neural network to obtain a restoration model for restoring low-quality images to high-quality images;

[0023] An image restoration module, configured to input a low-quality image to be restored into the restoration model, extract features through convolutional layers in the encoder to obtain an initial feature map; perform feature extraction of different frequencies on the initial feature map through several parallel frequency branches in the multi-scale dynamic frequency refinement module, and fuse the outputs of each frequency branch through the fusion layer in the multi-scale dynamic frequency refinement module to obtain a multi-frequency feature map; suppress redundant features and enhance non-redundant features in the multi-frequency feature map through a gating module to obtain an encoded feature map; decode the encoded feature map through a decoder to obtain a high-quality image corresponding to the low-quality image to be restored.

[0024] The present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above image restoration method is implemented.

[0025] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above image restoration method is implemented.

[0026] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:

[0027] In the image restoration method provided by the present invention, by constructing a neural network combining convolutional layers, gated residual blocks, multi-scale dynamic frequency refinement modules, and gating modules, the ill-posedness and information loss problems in image restoration can be effectively solved. The multi-scale dynamic frequency refinement module in this solution can capture multi-scale information in the image through parallel frequency branches and a fusion layer, avoid information loss in the deep network of a single convolutional structure, and at the same time, the gating module can adaptively suppress redundant features and enhance non-redundant features, improving the utilization efficiency of features. In addition, through the introduction of gated residual blocks, the network can better learn the internal structure of the image and reduce the over-smoothing phenomenon. Overall, this structure improves the network's reconstruction ability for complex degraded images. Especially when processing images with spatially-varying degradation kernels, higher-quality reconstruction results can be obtained without the computational complexity being too high due to the increase in image size, thus effectively solving the limitations of traditional methods and existing convolutional networks in image restoration tasks. Description of the Drawings

[0028] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0029] Figure 1 is a schematic diagram of the overall network architecture provided by the present invention;

[0030] Figure 2 is a schematic diagram of the gated residual block structure provided by the present invention;

[0031] Figure 3 is a schematic diagram of the multi-scale dynamic frequency refinement module structure provided by the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] The research on mainstream image restoration tasks can be divided into the following categories. (1) Model-based methods: Some traditional methods perform image restoration based on various assumptions, optimization algorithms, or manually crafted features. Due to the extremely complex and time-consuming nature of manually crafted features, traditional methods perform poorly on highly degraded images. (2) Deep learning-based methods: The blurred and clear image pairs are fed into a neural network, and relevant loss functions are used to control the gradient update direction of the network. The powerful data fitting ability of the neural network is utilized to learn the mapping relationship between the blurred and clear image pairs. Early convolutional networks based on the encoder-decoder structure were limited by local characteristics, suppressing the high-frequency information of the image and resulting in overly smooth results. The self-attention network can well preserve the texture details of the image by calculating the correlations of global pixels, but due to its high computational complexity, it is difficult to be used for high-resolution images. (3) Methods combining models and deep learning: Such methods use deep networks to learn image prior information and combine optimization algorithms to iteratively restore the process. However, such methods are based on known degradation patterns and are difficult to handle real-world degraded images with unknown degradation patterns.

[0034] Based on this, the present invention proposes a new image restoration network based on an encoder-decoder structure. In the network, the encoder-decoder block is designed as a cascade of a feature extraction module and a gating module, called the gated residual module (GRBlock). This module not only enhances the feature extraction ability but also promotes the flow of multi-scale information in the network, enabling the network to pay more attention to useful information and filter out useless information. In addition, in order to further refine the features, a multi-scale dynamic frequency refinement module (MDFR) is proposed. This module is embedded in the last GRBlock of each layer of the encoder-decoder. It not only enhances the reconstruction of high-frequency information in the image features but also the introduced multi-scale features can better handle spatially varying degradation kernels. This method achieves high-quality image reconstruction with lower computational complexity.

[0035] Embodiment 1

[0036] The following details the image restoration method of this embodiment, which specifically includes the following steps:

[0037] S1: Construct a neural network.

[0038] Image restoration aims to remove internal noise, environmental noise (rain, fog, and snow), and blur artifacts (motion blur and defocus blur) generated during the acquisition, transmission, and storage of images, so as to obtain a reconstructed high-quality image. Mainstream methods use convolutional networks, but limited by the locality and static weights of convolutions, they perform poorly in image detail restoration. The present invention proposes a multi-scale dynamic frequency refinement module, which can achieve the dynamic reconstruction of high-frequency and low-frequency features and combine a backbone network with a gating mechanism to control the transmission of information flow in each layer of the network to achieve high-quality image restoration. The present invention proposes an image restoration network based on an encoder-decoder structure, and the overall framework of the network is as Figure 1As shown in Figure 1. The network consists of three layers of encoders and three layers of decoders. Each layer of encoders passes information to the decoder of the same layer through skip connections. Each layer of encoders is followed by a strided convolution layer as a downsampling operation, and each layer of decoders is followed by a transposed convolution layer as an upsampling operation. Specifically, the input of the network is a degraded image of size H×W×3. First, a 3×3 convolution is used to map the features of the image to a high-dimensional space. Then, the GRBlock module encodes the high-dimensional features. In addition, in order to learn the multi-scale representation of the features, the original degraded image is downsampled twice and four times using linear interpolation, and is concatenated with the features of the first and second downsampling, respectively, as the input of the next layer of encoder. In the process of encoding and downsampling layer by layer, the spatial dimension of the feature is halved and the channel dimension is doubled. The decoder adopts a design structure symmetrical to the encoder, and gradually restores the original size and number of channels of the image during the layer-by-layer decoding process. In order to enable the network to learn features better, the decoded features of each layer are residually connected to the input of the encoder of that layer, and the reconstructed image is restored through 3×3 convolution, and the reconstructed multi-scale image is used for model training.

[0039] S101: Gated residual block (GRBlock).

[0040] When feature information is transmitted in a deep convolutional network, the network tends to retain low-frequency information and filter high-frequency information. In order to encourage the network to reconstruct high-frequency detail information of the image, an improved gating submodule is cascaded after the feature extraction module. Figure 2 As shown, the proposed GRBlock consists of two parts. On the left side of the dotted line is an improved residual module, which, in addition to the layer normalization layer and the convolution layer, also embeds the multi-scale dynamic frequency refinement module proposed in the present invention. The function of this submodule is to extract features. On the right side of the dotted line is the gating module, which does not use the traditional activation function to introduce nonlinearity, but adopts an element-by-element multiplication operation between features. This structure can not only introduce nonlinearity, but also amplify the useful information in each layer of the network and reduce the proportion of useless information, thereby controlling the transmission of information flow in each layer. Since the stable training of deep networks is inseparable from regularization, in the training of small batch samples, the performance of layer normalization (LN) is better than batch normalization (BN), so layer normalization is applied to the residual submodule and the gating submodule respectively. The overall process of the gated residual block can be described as follows:

[0041] RES=X+W3(MDFR(W3(LN(X))))

[0042]

[0043] Among them, X and Y represent the input and output features respectively, LN(·) represents the layer normalization operation, W represents the weight with a convolution kernel size of (·), W represents the grouped convolution, · represents the element-wise dot product, and MDFR is the multi-scale dynamic frequency refinement module proposed in the present invention, which will be introduced in the next subsection. The input features extract high-frequency detail features through the residual module embedded with the MDFR module, and the extracted features use the gating module to screen out useful information, promote the transmission of features in each layer in the network, and retain the high-frequency detail features. To ensure efficiency, the number of gating residual blocks in each layer is set to 16, and the MDFR module is added only in the last block.

[0044] S102: Multi-scale dynamic frequency refinement module (MDFR).

[0045] In the convolutional network, the convolutional block usually acts as a low-pass filter with static weights, which suppresses the high-frequency information of the image and is not conducive to restoring high-quality images. To retain the high-frequency detail features of the image, this method first dynamically separates the low-frequency and high-frequency feature maps in the original features, and then assigns learnable weights to them respectively, so as to LNConv3×3Conv3×3MDFRLNConv1×1Conv1×1Conv3×3Conv3×3Conv1×1 strengthen the restoration of high-frequency detail information. As Figure 3 shown, GAP×N represents downsampling the original feature map by N times using global average pooling, DFR represents the dynamic frequency refinement module, and BLI represents the upsampling operation using bilinear interpolation. As Figure 3 shown in the dashed box, in the DFR module, first use global average pooling to compress the resolution of the original input features into 1×1, and then use 1×1 convolution for dimension expansion. For backpropagation, batch normalization and Softmax operations are also performed on it. To align with the original feature size, use the Fold operation to fold the feature with a size of 1×1×9C after Softmax processing into a convolution kernel with a size of 3×3×C. This convolution kernel is dynamically obtained from the original input features and can adapt to different degradation modes. To separate the frequency information, first use this dynamic convolution kernel to convolve with the original features to obtain the low-frequency feature map, and then subtract the low-frequency feature map from the original features to obtain the high-frequency feature map. Finally, assign learnable weights H and L to them respectively and then perform element-wise addition to obtain the weighted output feature map Y rich in high-frequency detail information. Figure 3 The overall process of can be described as follows:

[0046] W L = Fold(Softmax(BN(W1(GAP(X)))))

[0047]

[0048] Among them, X and Y represent the input and output features respectively, BN represents batch normalization, Softmax is a normalization function, Fold represents the feature folding operation, and L and H are the weights of the low-frequency and high-frequency feature maps respectively. represents the convolution operation, GAPk represents k-fold downsampling using global average pooling, and BLI represents the upsampling operation using bilinear interpolation. In the present invention, the MDFR module is only embedded in the last GRBlock of each layer of the encoder-decoder, which ensures the efficiency of the model.

[0049] S2: Train the neural network.

[0050] The present invention mainly conducts experiments on three image restoration tasks, namely image motion deblurring, image desnowing, and image defogging. Separate models are trained for different tasks. The parameter settings of the model are as follows: The batch size is set to 4. The input data is randomly cropped on the original dataset, with a patch size of 256×256, and each patch is randomly horizontally flipped for data augmentation. The initial learning rate is 1e-4 and gradually decreases to 1e-6 with cosine annealing. The Adam (β1 = 0.9, β2 = 0.999) optimizer is used for training. The number of GRBlock blocks in each encoder-decoder block is set to 16, and the proposed MSFR module is embedded in the last block. We use PyTorch to train on the NVIDIA 4090 GPU.

[0051] S3: Use the trained neural network for image restoration.

[0052] Input the low-quality image to be restored into the restoration model, extract features through the convolutional layer in the encoder to obtain the initial feature map; extract features of different frequencies from the initial feature map through several parallel frequency branches in the multi-scale dynamic frequency refinement module, and fuse the outputs of each frequency branch through the fusion layer in the multi-scale dynamic frequency refinement module to obtain the multi-frequency feature map; suppress the redundant features and enhance the non-redundant features in the multi-frequency feature map through the gating module to obtain the encoded feature map; decode the encoded feature map through the decoder to obtain the high-quality image corresponding to the low-quality image to be restored.

[0053] The above image restoration method proposes a multi-scale dynamic frequency refinement module for the problem that traditional convolutional networks suppress high-frequency detail information, resulting in blurred reconstructed images. This module dynamically obtains a low-pass filter from the original input features through average pooling and feature folding operations, separates the low-frequency and high-frequency feature maps through convolutional operations, and obtains a concatenated feature rich in high-frequency information after weighting. For degraded images with unknown and spatially varying degradation patterns, different from traditional networks that apply multi-scale to the overall input and output of the network, the present invention also applies multi-scale to the proposed dynamic frequency refinement module. This module uses global average pooling operations to downsample the input features to three different scales respectively. After the same frequency refinement operations are performed between different scales, the features of each scale are restored to the input scale through bilinear interpolation, and the output features are obtained after element-wise addition. For the problem that useful information is lost during the transmission process in the deep network, an improved gated residual module is proposed. Among them, the residual sub-module embeds the proposed multi-scale frequency refinement module, and the gated sub-module groups the normalized features using depth convolution, and while introducing non-linearity to the network through element-wise dot product operations, controls the transmission of information flow in each layer of the network.

[0054] The above is the image restoration method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding image restoration device, including:

[0055] A network construction module, configured to construct a neural network including an encoder and a decoder. The encoder includes a convolutional layer and a gated residual block connected in sequence; the gated residual block includes an improved residual module and a gated module; the improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module, and the multi-scale dynamic frequency refinement module includes a plurality of parallel frequency branches and a fusion layer connected to the output ends of the plurality of parallel frequency branches at the same time.

[0056] A network training module, configured to collect low-quality - high-quality image pairs to construct a data set, and use the data set to train the constructed neural network to obtain a restoration model for restoring low-quality images to high-quality images.

[0057] An image restoration module, configured to input a low-quality image to be restored into the restoration model, extract features through the convolutional layer in the encoder to obtain an initial feature map; extract features of different frequencies from the initial feature map through a plurality of parallel frequency branches in the multi-scale dynamic frequency refinement module, and fuse the outputs of each frequency branch through the fusion layer in the multi-scale dynamic frequency refinement module to obtain a multi-frequency feature map; suppress redundant features and enhance non-redundant features in the multi-frequency feature map through the gated module to obtain an encoded feature map; decode the encoded feature map through the decoder to obtain a high-quality image corresponding to the low-quality image to be restored.

[0058] For the specific limitations of the image restoration device, reference can be made to the limitations of the image restoration method in the above text, which will not be elaborated here. Each module in the above image restoration device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0059] The present invention also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above-provided image restoration method.

[0060] The present invention also provides the structure of a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-provided image restoration method.

[0061] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. The non-volatile memory can include a read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc.

[0062] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present invention.

Claims

1. An image restoration method, characterized in that, Including: Construct a neural network including an encoder and a decoder. The encoder includes a convolutional layer and a gated residual block connected in sequence. The gated residual block includes an improved residual module and a gating module. The improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module. The multi-scale dynamic frequency refinement module includes several parallel frequency branches and a fusion layer connected to the output ends of the several parallel frequency branches at the same time. Collect low-quality - high-quality image pairs to construct a dataset, and use the dataset to train the constructed neural network to obtain a restoration model for restoring low-quality images to high-quality images. Input the low-quality image to be restored into the restoration model. Extract features through the convolutional layer in the encoder to obtain an initial feature map. Extract features of different frequencies from the initial feature map through several parallel frequency branches in the multi-scale dynamic frequency refinement module, and fuse the outputs of each frequency branch through the fusion layer in the multi-scale dynamic frequency refinement module to obtain a multi-frequency feature map. Suppress redundant features and enhance non-redundant features in the multi-frequency feature map through the gating module to obtain an encoded feature map. Decode the encoded feature map through the decoder to obtain a high-quality image corresponding to the low-quality image to be restored.

2. The image restoration method according to claim 1, characterized in that, The dynamic frequency refinement layer includes an input layer, a feature extraction branch, a dynamic convolutional layer, and a separation layer connected in parallel to the output end of the input layer. The feature extraction branch includes a global average pooling layer, a 1×1 convolutional layer, a batch normalization layer, a Softmax activation function layer, and a folding layer connected in sequence. The dynamic frequency refinement layer extracts features of different frequencies from the initial feature map, specifically including: Compress the feature dimension of the input to the dynamic frequency refinement layer by using the global average pooling layer; expand the dimension of the output of the global average pooling layer through the 1×1 convolutional layer; perform batch normalization operation and activation operation in sequence through the batch normalization layer and the Softmax activation function layer; fold the output of the Softmax activation function layer into a dynamic convolutional kernel through the folding layer. Perform a convolution operation on the features of the input to the dynamic frequency refinement layer with the dynamic convolutional kernel output by the feature extraction branch through the dynamic convolutional layer, and output low-frequency features. Subtract the low-frequency features from the features of the input to the dynamic frequency refinement layer through the separation layer to obtain high-frequency features.

3. The image restoration method according to claim 1, wherein The encoder includes a first encoder, a second encoder, and a third encoder connected in sequence. The first encoder, the second encoder, and the third encoder have the same structure. Among them, the input of the first encoder is the low-quality image to be restored; after performing the first downsampling operation on the low-quality image to be restored, fuse it with the output of the first encoder, and use the fusion result as the input of the second encoder; after performing the second downsampling operation on the low-quality image to be restored, fuse it with the output of the second encoder, and use the fusion result as the input of the third encoder.

4. The image restoration method according to claim 3, wherein The decoder includes a first decoder, a second decoder, and a third decoder connected in sequence. The first decoder, the second decoder, and the third decoder have the same structure, and each includes a gated residual block and a convolutional layer connected in sequence. Perform feature decoding operation on the output of the third encoder through a gated residual block, and use the obtained result and the output of the third encoder as the input of the gated residual block in the first decoder; fuse the output of the gated residual block in the first decoder and the result obtained after the second downsampling operation on the low-quality image to be restored, and use it as the input of the convolutional layer in the first decoder; Use the output of the gated residual block in the first decoder and the output of the gated residual block in the second encoder as the input of the gated residual block in the second decoder; Fuse the output of the gated residual block in the second decoder and the result obtained after the first downsampling operation on the low-quality image to be restored, and use it as the input of the convolutional layer in the second decoder; Use the output of the gated residual block in the second decoder and the output of the gated residual block in the first encoder as the input of the gated residual block in the third decoder; Fuse the output of the gated residual block in the third decoder and the low-quality image to be restored, and use it as the input of the convolutional layer in the third decoder.

5. The image restoration method according to claim 1, wherein The several parallel frequency branches include parallel first, second, and third frequency branches. Each frequency branch includes a global average pooling layer, a dynamic frequency refinement layer, and a bilinear interpolation upsampling layer connected in sequence. Among them, the downsampling multiples of the global pooling layers in the first, second, and third frequency branches are 2, 4, and 8 respectively.

6. The image restoration method according to claim 1, wherein The improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module. The multi-scale dynamic frequency refinement module is embedded between two convolutional layers of the original residual module. The obtained improved residual module includes an input layer, a normalization layer, a first 3×3 convolutional layer, a multi-scale dynamic frequency refinement module, and a second 3×3 convolutional layer connected in sequence; among them, the input end of the input layer and the output end of the second 3×3 convolutional layer are connected through a residual connection.

7. An image restoration device, characterized in that, Including: A network construction module for constructing a neural network including an encoder and a decoder. The encoder includes a convolutional layer and a gated residual block connected in sequence; the gated residual block includes an improved residual module and a gating module; the improved residual module is obtained by embedding a multi-scale dynamic frequency refinement module in the original residual module, and the multi-scale dynamic frequency refinement module includes several parallel frequency branches and a fusion layer connected to the output ends of the several parallel frequency branches at the same time; A network training module for collecting low-quality - high-quality image pairs to construct a data set, and using the data set to train the constructed neural network to obtain a restoration model for restoring low-quality images to high-quality images; An image restoration module, configured to input a low-quality image to be restored into a restoration model, extract features through convolutional layers in an encoder to obtain an initial feature map; extract features of different frequencies from the initial feature map through a plurality of parallel frequency branches in a multi-scale dynamic frequency refinement module, and fuse the outputs of each frequency branch through a fusion layer in the multi-scale dynamic frequency refinement module to obtain a multi-frequency feature map; suppress redundant features and enhance non-redundant features in the multi-frequency feature map through a gating module to obtain an encoded feature map; decode the encoded feature map through a decoder to obtain a high-quality image corresponding to the low-quality image to be restored.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 above is implemented.

9. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method according to any one of claims 1 to 6 above is implemented.

Citation Information

Patent Citations

  • Underwater image restoration model based on multi-branch gating fusion and restoration method thereof

    CN111754438A

  • Heterocyclic compound, organic light emitting device and composition for organic material layer of organic light emitting device

    KR1020250034723A

Cited By

  • Self-adaptive image restoration method based on degradation type and degree joint perception

    CN120953136A

  • Self-adaptive gating deblurring system and method fusing multi-scale and multi-direction blurring features

    CN122048719A

  • An adaptive gated defuzzification system and method that integrates multi-scale and multi-directional fuzzy features

    CN122048719B