Two-stage image restoration method combining wavelet convolution and Transform

By combining wavelet convolution and Transformer's two-stage image repair method, multi-scale edge features of the image are extracted and global context modeled, the problem of difficulty in recovering edge details and overall consistency in the prior art is solved, and high-quality image repair effect is achieved.

CN119991509APending Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510056349.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to ensure accurate recovery of edge details and overall image consistency in image restoration, especially when dealing with complex structures and textures.

Method used

A two-stage image repair method combining wavelet convolution and Transformer is used. First, multi-scale edge features of the image are extracted through a wavelet convolution network to generate edge information. Then, based on the Transformer network based on the TransUnet architecture, uses missing images and edge information as inputs, and performs global context modeling to generate the repaired image.

Benefits of technology

It significantly improves the quality and efficiency of image repair, and maintains a high level of detail recovery and overall consistency, especially when dealing with complex textures and large-area defect areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991509A_ABST
    Figure CN119991509A_ABST
Patent Text Reader

Abstract

The invention relates to a two-stage image restoration method combining wavelet convolution and Transform, belongs to the field of computer vision, and is used for effectively restoring detail information of a lost area in an image restoration task. The method comprises the following two stages: in the first stage, wavelet convolution is adopted as an encoder and a decoder, a residual structure is introduced into the middle part, and an edge repair network is constructed; multi-scale decomposition is carried out on the image through wavelet convolution, detail features of the image are effectively extracted, and the expression ability of the features is enhanced by using a residual structure, so that accurate edge information is generated; in the second stage, the edge information generated in the first stage is combined with a missing image, the combined image serves as input and is transmitted to an image inpainting network, global context modeling is conducted through a self-attention mechanism, and high-quality image inpainting is achieved. Through the method, the precision and the quality of image restoration can be improved, and meanwhile, the structural consistency and the visual naturalness of the image are kept under the condition of not depending on complex prior information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image processing, and more specifically to a two-stage image restoration method that combines wavelet convolution and Transformer. This method combines edge information restoration with overall image reconstruction to achieve high-quality restoration of missing parts of an image. Background Art

[0002] Image inpainting is a key research area in computer vision, aiming to restore missing or damaged parts of an image, visually integrating the restored image seamlessly with the original. Traditional image inpainting methods rely primarily on interpolation techniques or neighborhood-based information filling. While these methods are effective for simple missing regions, they often struggle to maintain overall image consistency and detail when faced with complex structures and textures.

[0003] In recent years, deep learning technology has made significant progress in the field of image restoration. Convolutional neural networks (CNNs) have been widely used in image restoration tasks due to their powerful feature extraction capabilities. However, CNNs are limited in capturing long-range dependencies and global contextual information. The Transformer model, with its self-attention mechanism, can effectively model global image relationships, compensating for the shortcomings of CNNs. However, using the Transformer alone for image restoration may not perform well in detail recovery and local information processing. Therefore, how to combine the advantages of CNNs and Transformers to achieve high-quality image restoration has become a hot topic of current research.

[0004] Although existing technologies have achieved certain results in image restoration, they still have shortcomings in coordinating the processing of image edge information and global context information, making it difficult to simultaneously ensure the accurate restoration of edge details and the overall consistency of the image. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this paper proposes a two-stage image restoration method that combines wavelet convolution and Transformer. This method uses a two-stage process: first, edge information is generated using a wavelet convolutional network. Then, the missing image and edge information are used as input to perform overall image restoration using an improved TransUnet network.

[0006] In order to achieve the above technical objectives, the specific technical solutions are as follows:

[0007] A two-stage image restoration method combining wavelet convolution and Transformer, the method includes the following steps:

[0008] S1: Input image preprocessing step, normalize the image to be repaired, and generate a mask image of the defect area;

[0009] S2: Construct an edge restoration network, use wavelet convolution as the encoder and decoder, and introduce a residual structure in the middle part to generate edge information;

[0010] S3: Build an image restoration network based on the TransUnet architecture. It uses a wavelet convolutional structure to replace the traditional encoder and decoder. It takes the edge information generated in the first stage and the missing image as input, uses the Transformer to perform global context modeling, and generates the restored image.

[0011] S4: Image reconstruction step, which gradually restores the spatial resolution of the image through the deconvolution layer to generate the final repaired image;

[0012] S5: Loss function and training strategy step, using multiple loss functions and a staged training strategy to optimize the model.

[0013] Furthermore, the edge restoration network of step S2 includes multiple wavelet convolution layers and nonlinear activation functions, which are used to extract multi-scale edge features of the image, wherein the feature extraction process can be expressed as:

[0014] F edge =WTConv(I pre )+Residual(F edge )

[0015] Where, I pre represents the preprocessed image, F edge Represents the extracted edge features, WTConv represents the wavelet convolution operation, and Residual represents the residual structure.

[0016] Furthermore, the Transformer network of step S3 includes multiple Transformer encoder layers, which use the self-attention mechanism to model the global context information of the image, wherein the feature modeling process can be expressed as:

[0017] F global = TransformerEncoder(F input )

[0018] Where, F input F is the input feature that combines the missing image and edge information. global is the global context feature, and TransformerEncoder represents the Transformer encoder operation.

[0019] Furthermore, the loss function of step S5 includes:

[0020] 1) Pixel-level loss, used to ensure the consistency of the perceptual features of the restored image, which is expressed as:

[0021] L pixel =||I true -I restored ||1

[0022] Where, I true represents the original image, I restired represents the restored image;

[0023] 2) Perceptual loss, which is used to ensure the consistency of the perceptual features of the restored image, is expressed as:

[0024] L perceptual =||φ(I true )-φ(I restored )||2

[0025] Where φ represents the features extracted by the pre-trained VGG network;

[0026] 3) Adversarial loss, used to improve the authenticity and naturalness of the restored image, which is expressed as:

[0027]

[0028] Where D represents the discriminator.

[0029] Furthermore, the training strategy of step S5 includes:

[0030] 1) In the first stage, the edge restoration network is trained separately to enable it to accurately generate edge information;

[0031] 2) In the second stage, the Transformer image restoration network and the edge restoration network are jointly trained to enable the overall model to be collaboratively optimized and improve the restoration effect.

[0032] Furthermore, the Transformer encoder layer of step S3 includes a multi-head self-attention mechanism and a feedforward neural network, and its mathematical expression is:

[0033] MultiHead(Q,K,V)=Concat(head1,head2,…head h )W O

[0034]

[0035] Where Q, K, and V represent query, key, and value respectively. is the weight matrix of the i-th head, W O is the output weight matrix and h is the number of heads.

[0036] Furthermore, the image reconstruction step of step S4 gradually restores the spatial resolution of the image through the deconvolution layer, and its mathematical expression is:

[0037] I reconstructed =DeConv(F global )

[0038] Where, I reconstructed Represents the reconstructed image, and DeConv represents the deconvolution operation.

[0039] Furthermore, the defect area mask generation step of step S1 includes detecting the defect area in the image and generating a corresponding binary mask image, the mathematical expression of which is:

[0040]

[0041] Where M(x,y) represents the mask value at position (x,y).

[0042] The present invention has at least the following beneficial effects

[0043] Compared with the existing technology, the two-stage image restoration method of the present invention significantly improves the quality and efficiency of image restoration by combining wavelet convolution and Transformer technology. First, wavelet convolution is used as the encoder and decoder, and the residual structure is introduced to effectively extract the multi-scale edge features of the image and enhance the ability to capture details and structures. Secondly, the Transformer network based on the TransUnet architecture performs global context modeling of the image through the self-attention mechanism, ensuring the consistency and naturalness of the restored image in the overall structure and texture. In addition, the multi-scale feature extraction capability of the wavelet convolution network enables this method to maintain a high level of detail restoration effect when processing complex textures and large defective areas. The two-stage processing flow not only achieves a good balance between details and overall consistency, but also improves the training efficiency and stability of the model through a staged training strategy.

[0044] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0046] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0047] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0048] See also Figure 1 , a two-stage image restoration method combining wavelet convolution and Transformer, the method includes the following steps:

[0049] S1: Input image preprocessing step, normalize the image to be repaired, and generate a mask image of the defect area;

[0050] S2: Construct an edge restoration network, use wavelet convolution as the encoder and decoder, and introduce a residual structure in the middle part to generate edge information;

[0051] S3: Build an image restoration network based on the TransUnet architecture. It uses a wavelet convolutional structure to replace the traditional encoder and decoder. It takes the edge information generated in the first stage and the missing image as input, uses the Transformer to perform global context modeling, and generates the restored image.

[0052] S4: Image reconstruction step, which gradually restores the spatial resolution of the image through the deconvolution layer to generate the final repaired image;

[0053] S5: Loss function and training strategy step, using multiple loss functions and a staged training strategy to optimize the model.

[0054] Furthermore, the edge restoration network of step S2 includes multiple wavelet convolution layers and nonlinear activation functions, which are used to extract multi-scale edge features of the image, wherein the feature extraction process can be expressed as:

[0055] F edge =WTConv(I pre )+Residual(F edge )

[0056] Where, I pre represents the preprocessed image, F edge Represents the extracted edge features, WTConv represents the wavelet convolution operation, and Residual represents the residual structure.

[0057] Furthermore, the Transformer network of step S3 includes multiple Transformer encoder layers, which use the self-attention mechanism to model the global context information of the image, wherein the feature modeling process can be expressed as:

[0058] F global = TransformerEncoder(F input )

[0059] Where, F input F is the input feature that combines the missing image and edge information. global is the global context feature, and TransformerEncoder represents the Transformer encoder operation.

[0060] Furthermore, the loss function of step S5 includes:

[0061] 1) Pixel-level loss, used to ensure the consistency of the perceptual features of the restored image, which is expressed as:

[0062] L pixel =||I true -I restored ||1

[0063] Where, I true represents the original image, I restored represents the restored image;

[0064] 2) Perceptual loss, which is used to ensure the consistency of the perceptual features of the restored image, is expressed as:

[0065] L perceptual =||φ(I true )-φ(I restored )||2

[0066] Where φ represents the features extracted by the pre-trained VGG network;

[0067] 3) Adversarial loss, used to improve the authenticity and naturalness of the restored image, which is expressed as:

[0068]

[0069] Where D represents the discriminator.

[0070] Furthermore, the training strategy of step S5 includes:

[0071] 1) In the first stage, the edge restoration network is trained separately to enable it to accurately generate edge information;

[0072] 2) In the second stage, the Transformer image restoration network and the edge restoration network are jointly trained to enable the overall model to be collaboratively optimized and improve the restoration effect.

[0073] Furthermore, the Transformer encoder layer of step S3 includes a multi-head self-attention mechanism and a feedforward neural network, and its mathematical expression is:

[0074] MultiHead(Q,K,V)=Concat(head1,head2,…head h )W O

[0075]

[0076] Where Q, K, and V represent query, key, and value respectively. is the weight matrix of the i-th head, W O is the output weight matrix and h is the number of heads.

[0077] Furthermore, the image reconstruction step of step S4 gradually restores the spatial resolution of the image through the deconvolution layer, and its mathematical expression is:

[0078] I reconstructed =DeConv(F global )

[0079] Where, I reconstructed Represents the reconstructed image, and DeConv represents the deconvolution operation.

[0080] Furthermore, the defect area mask generation step of step S1 includes detecting the defect area in the image and generating a corresponding binary mask image, the mathematical expression of which is:

[0081]

[0082] Where M(x,y) represents the mask value at position (x,y).

[0083] Figure 1 Flow chart of the method of the present invention

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A two-stage image restoration method combining wavelet convolution and Transformer, characterized in that: The following steps are involved: S1: Input image preprocessing step, standardize the image to be repaired, and generate a mask image of the defect area; S2: Construct an edge restoration network, use wavelet convolution as the encoder and decoder, and introduce a residual structure in the middle part to generate edge information; S3: Build an image restoration network based on the TransUnet architecture, use a wavelet convolution structure to replace the traditional encoder and decoder, take the edge information and missing image generated in the first stage as input, use Transformer to perform global context modeling, and generate the restored image; S4: Image reconstruction step, which gradually restores the spatial resolution of the image through the deconvolution layer to generate the final repaired image; S5: Loss function and training strategy step, using multiple loss functions and a staged training strategy to optimize the model.

2. The method according to claim 1, wherein the edge restoration network of step S2 comprises a plurality of wavelet convolution layers and a nonlinear activation function for extracting multi-scale edge features of the image, wherein the features The extraction process can be expressed as: F edge =WTConv(I pre )+Residual(F edge ) In the formula, I pre represents the preprocessed image, F edge Represents the extracted edge features, WTConv represents the wavelet convolution operation, and Residual represents the residual structure.

3. The method according to claim 1, wherein the Transformer network of step S3 comprises a plurality of Transformer encoder layers, using a self-attention mechanism to model the global context information of the image, wherein the feature modeling The process can be expressed as: F global =TransformerEncoder(F input ) In the formula, F input F is the input feature that combines the missing image and edge information. global is the global context feature, and TransformerEncoder represents the Transformer encoder operation.

4. The method according to claim 1, wherein the loss function of step S5 comprises: 1) Pixel-level loss, which is used to ensure the consistency of the perceptual features of the restored image, is expressed as: L pixel =||I true -I restored ||1 In the formula, I true represents the original image, I restored represents the restored image; 2) Perceptual loss, which is used to ensure the consistency of the perceptual features of the restored image, is expressed as: L perceptual =||φ(I true )-φ(I restored )||2 In the formula, φ represents the features extracted by the pre-trained VGG network; 3) Adversarial loss, which is used to improve the authenticity and naturalness of the restored image, is expressed as: In the formula, D represents the discriminator.

5. The method according to claim 1, wherein the training strategy of step S5 comprises: 1) In the first stage, the wavelet convolution edge restoration network is trained separately to enable it to accurately generate edge information; 2) In the second stage, the Transformer image restoration network and the wavelet convolution edge restoration network are jointly trained so that the overall model can be optimized collaboratively to improve the restoration effect.

6. The method according to claim 1, wherein the Transformer encoder layer of step S3 includes a multi-head self-attention mechanism and a feedforward neural network, and its mathematical expression is: MultiHead(Q,K,V)=Concat(head1,head2,…head h )W O In the formula, Q, K, and V represent query, key, and value respectively. is the weight matrix of the i-th head, W O is the output weight matrix and h is the number of heads.

7. The method according to claim 1, wherein the image reconstruction step of step S4 gradually restores the spatial resolution of the image through a deconvolution layer, and its mathematical expression is: I reconstructed =DeConv(F global ) In the formula, I reconstructed represents the reconstructed image, and DeConv represents the deconvolution operation.

8. The method according to claim 1, wherein the defect area mask generation step of step S1 comprises detecting the defect area in the image and generating a corresponding binary mask image, and the mathematical expression thereof is: Where M(x,y) represents the mask value at position (x,y).

Citation Information

Patent Citations

  • Image restoration method and system based on wavelet prior attention

    CN113034390A

  • Transform and convolutional neural network-based image restoration method

    CN115731138A

  • Multi-scale Transform-Unet low-light image enhancement method based on wavelet transform

    CN118864329A