Image blind restoration method based on self-information and prediction mask enhancement

By adopting an enhancement method based on self-information and prediction masks in image blind repair, using Transformer Block and information enhancement features, the problems of insufficient guidance of damaged areas and susceptible to prediction results in the prior art are solved, and better image repair effect and robustness are achieved.

CN120182141AActive Publication Date: 2025-06-20LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510241414.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing blind image repair methods are insufficiently guided in damaged areas and are susceptible to the prediction results of the front network, resulting in poor repair results.

Method used

The image blind repair method based on self-information and prediction mask enhancement is adopted, and the damaged image is encoded and information enhanced through the encoding module composed of Transformer Block. The prediction mask and damaged image are used to generate information enhancement features, and the enhanced information is gradually optimized to guide the encoding process.

Benefits of technology

Improves the robustness and visual effect of blind repair of images, reduces the possibility of repair failure, and provides better image repair results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182141A_ABST
    Figure CN120182141A_ABST
Patent Text Reader

Abstract

The invention discloses an image blind restoration method based on self-information and prediction mask enhancement, and belongs to the technical field of computer vision. The method comprises the following steps: converting a damaged image into a two-dimensional feature; encoding the two-dimensional features by using a TSBlock encoding module to obtain encoded features; carrying out damaged area mask prediction on the damaged image, and generating information enhancement features which only retain undamaged areas by utilizing a prediction mask; encoding the encoding feature and the information enhancement feature by using an information enhancement IETSBlock encoding module to obtain an information enhancement encoding feature; re-coding the information enhancement coding features by using a TSBlock coding module; the recoding features are recovered into multi-channel features consistent with the damaged image in size through a scale recovery and refinement module; and channel filtering is carried out on the multi-channel features through channel recovery to obtain a restored image. According to the invention, the problem that image restoration is easy to fail is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more particularly to an image blind restoration method based on self-information and prediction mask enhancement. Background Art

[0002] With the development of current artificial intelligence and computer vision technologies, image processing technologies such as Artificial Intelligence Generated Content (AIGC) have gradually shown great development potential and have been applied to various fields such as education and medical care. The development of such technologies relies on a large number of image datasets with complete content. However, the images collected from various sources inevitably have degradation phenomena such as damage, which in turn has a negative impact on the performance of the model. At the same time, the poor visual effect of damaged images also affects the practicality of the images themselves. (For example: damaged medical images, damaged traffic surveillance capture images, etc.).

[0003] Traditional image restoration methods mainly rely on specific patterns to restore images, with weak generalization ability and unable to adapt to the restoration of complex damaged images. Therefore, in recent years, with the development of deep learning, using digital image restoration technology to restore damaged images has become a current research hotspot and has been applied to a certain extent. Digital image restoration technology can quickly restore relatively complex damaged images such as natural landscapes, mural monuments, and old photos.

[0004] Currently, the mainstream digital image restoration methods are mainly divided into two categories: non-blind restoration methods and blind restoration methods. The non-blind damaged image restoration method defaults that the mask is known and has achieved good image restoration effects in damaged image restoration. However, the non-blind restoration method relies heavily on the mask. Therefore, such methods need to first annotate the damaged areas of the image during testing, and the application is limited.

[0005] Based on the above problems, in recent years, researchers have begun to focus on testing image blind restoration methods without masks. The current blind restoration methods mainly include two paradigms: direct blind restoration methods and prediction-based multi-stage blind restoration methods. The direct blind restoration method is usually a single-stage network. Due to the lack of guidance for damaged areas and undamaged areas, such methods are easily affected by the content of damaged areas and there are color artifact residues. The prediction-based multi-stage blind restoration method is mainly composed of several sub-networks trained in multiple stages connected in series. Such methods are easily affected by the prediction results of the previous network, resulting in the failure of subsequent network restoration.

[0006] Therefore, how to solve the problems such as the lack of guidance for damaged areas in direct blind restoration methods and the susceptibility to the prediction results of sub-networks in multi-stage blind restoration methods is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides an image blind restoration method and system based on self-information and prediction mask enhancement, which are used to solve at least some of the technical problems in the background technology.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention first discloses an image blind restoration method based on self-information and prediction mask enhancement, including the following steps:

[0010] Convert the damaged image into feature tokens;

[0011] Encode the feature tokens by using a first encoding module composed of Transformer Blocks to obtain encoded features;

[0012] Perform mask prediction on the damaged image, and use the prediction mask and the damaged image to generate an information-enhanced feature that only retains the pixels in the undamaged area;

[0013] Encode the encoded features and the information-enhanced features by using a second encoding module composed of information-enhanced Transformer Blocks to obtain information-enhanced encoded features;

[0014] Recode the information-enhanced encoded features by using a third encoding module composed of Transformer Blocks to obtain recoded features;

[0015] Restore the recoded features to multi-channel features with the same size as the damaged image through a scale restoration and refinement module;

[0016] Perform channel filtering on the multi-channel features through channel restoration to obtain the restored image.

[0017] In a specific embodiment, converting the damaged image into feature tokens specifically includes:

[0018] Divide the damaged image I with a size of h*w*c d into a number of image blocks I' with the same size d , where the size of each image block is p*p*c, h represents the height of the damaged image I d , w represents the width of the damaged image I d ; p represents the height or width of the image block; c represents the number of channels;

[0019] Flatten each image block I' d into a one-dimensional feature vector of p 2 *c;

[0020] Project the flattened one-dimensional feature vectors into a set D-dimensional space using a linear layer to obtain each image patch I'. d The corresponding Patch Embedding sequence;

[0021] Combine the Patch Embedding sequences of all image patches I'. d to obtain a Patch Embeddings sequence, which is the embedding sequence of the damaged image I d ;

[0022] Add a class token to the embedding sequence of the damaged image I d ;

[0023] Perform positional encoding on the embedding sequence with the added class token to obtain the feature tokens of the damaged image I d ;

[0024] Furthermore, in the step of encoding the feature tokens using the first encoding module composed of TransformerBlocks in the present invention, the first encoding module includes two TransformerBlock units with the same structure and stacked on top of each other, denoted as the first TransformerBlock unit and the second Transformer Block unit;

[0025] wherein, the first TransformerBlock unit is used to perform primary encoding on the feature tokens.

[0026] Furthermore, in the embodiment of the present invention, the first TransformerBlock unit performs primary encoding on the feature tokens, which specifically includes the following encoding process:

[0027] F T1 =TSBlock(F embedding )

[0028] =MLP(Norm(MHA(Norm(F embedding ))+F embedding ))

[0029] wherein, F T1 represents the primary encoding result obtained by the first TransformerBlock performing primary encoding on the feature tokens, TSBlock() represents the TransformerBlock unit processing operation, F embedding represents the feature tokens obtained by converting the damaged image, Norm() represents the normalization operation, MHA() represents the multi-head attention operation, and MLP() represents the multi-layer perceptron operation.

[0030] Furthermore, in the embodiments of the present invention, for predicting the mask of the damaged area of the damaged image, and generating an information-enhanced feature that only retains the undamaged area by using the predicted mask and the damaged image, specifically includes:

[0031] Using the trained deep learning model to perform mask prediction on the damaged image to obtain the predicted mask of the damaged area;

[0032] Obtaining the damaged area and the undamaged area in the damaged image according to the predicted mask of the damaged area;

[0033] Retaining the pixel values of the undamaged area in the damaged image, and replacing the pixel values of the damaged area with fixed values to obtain an information-enhanced feature that only retains the pixel values of the undamaged area in the damaged image.

[0034] Furthermore, in the embodiments of the present invention, in the encoding step of encoding the encoded feature and the information-enhanced feature by using the second encoding module composed of information-enhanced TransformerBlocks, the second encoding module includes two information-enhanced TransformerBlock units with the same structure and stacked on each other, denoted as the first information-enhanced TransformerBlock unit and the second information-enhanced TransformerBlock unit;

[0035] Among them, the first information-enhanced TransformerBlock unit is used for initially encoding the encoded feature and the information-enhanced feature, and the second information-enhanced TransformerBlock unit is used for re-encoding the initial encoding result of the first information-enhanced TransformerBlock unit and the information-enhanced feature.

[0036] Furthermore, in the embodiments of the present invention, the first information-enhanced TransformerBlock unit initially encodes the encoded feature and the information-enhanced feature, specifically including the following encoding process:

[0037] F temp1 = MHA(Norm(F CT1 )) + F CT1 ;

[0038] F IET1 = Linear(Cat(Norm(MLP(Norm(F temp1 )) + F temp1 ), F enhanced ));

[0039] Among them, F temp1 represents the intermediate encoded feature, F IET1 represents the initial encoding result, F CT1It is indicated that the first encoding module obtains the encoded feature, F enhanced It represents the information enhancement feature, Norm() represents the normalization operation, MHA() represents the multi-head attention operation, MLP() represents the multi-layer perceptron operation, Cat() represents the concatenation operation, and Linear() represents the linear mapping.

[0040] Furthermore, in the embodiment of the present invention, in the step of re-encoding the information enhancement encoded feature by using the third encoding module composed of Transformer Blocks, the third encoding module includes five Transformer Block units with the same structure and stacked on each other.

[0041] Furthermore, in the embodiment of the present invention, in the step of restoring the re-encoded feature to a multi-channel feature with the same size as the damaged image through the scale recovery and refinement module, the scale recovery and refinement module includes two scale recovery and refinement units with the same structure and stacked on each other, denoted as the first scale recovery and refinement unit and the second scale recovery and refinement unit;

[0042] Among them, the first scale recovery and refinement unit is used to perform preliminary dimension elevation and refinement on the re-encoded feature output by the third encoding module to obtain the preliminary dimension elevation and refinement result.

[0043] Furthermore, in the embodiment of the present invention, the first scale recovery and refinement unit performs preliminary dimension elevation and refinement on the re-encoded feature output by the third encoding module, which specifically includes the following operation process:

[0044] F recovery1 = Relu(BatchNorm(Conv(Upsampling(F encoder ))))

[0045] Among them, F recovery1 represents the preliminary dimension elevation and refinement result, F encoder represents the re-encoded feature obtained by the third encoding module, Upsampling() represents the upsampling operation, Conv() represents the convolution operation, BatchNorm() represents the batch normalization, and Relu() represents the Relu activation operation.

[0046] The present invention also discloses an image blind repair system based on self-information and prediction mask enhancement, including a computer system. When the computer system executes, it can implement the image blind repair method based on self-information and prediction mask enhancement described in any one of the present invention.

[0047] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an image blind repair method and system based on self-information and prediction mask enhancement, which has the following beneficial effects:

[0048] The present invention proposes an image blind repair method based on self-information and prediction mask enhancement. This method uses the prediction mask and the effective information of the damaged image itself as enhancement information, continuously optimizes the enhancement information in a dynamic form during the training process, and provides guidance for the encoding process, solving the problem that the phased blind repair method is highly sensitive to the results of each stage, resulting in easy repair failure.

[0049] The image blind repair method proposed by the present invention has better robustness, and the repair result of the damaged image has better visual effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0051] Figure 1 Schematic diagram of the overall architecture process of the image blind repair method based on self-information and prediction mask enhancement provided by the invention.

[0052] Figure 2 Schematic diagram of the structure of the first Transformer Block (TSBlock) unit provided by the embodiment of the invention.

[0053] Figure 3 Schematic diagram of the structure of the second encoding module composed of the Information Enhancement Transformer Block (IETSBlock) provided by the embodiment of the present invention.

[0054] Figure 4 Schematic diagram of the structure of the first scale recovery and refinement unit in the scale recovery and refinement module provided by the embodiment of the invention.

[0055] Figure 5 Schematic diagram of the comparison effect of various different repair methods provided by the embodiment of the invention; among them, (a) is the damaged mural image to be repaired, (b) is the repaired image obtained by the TransCNN-HAE method, (c) is the repaired image obtained by the BOWN method, (d) is the repaired image obtained by the DT method, and (e) is the repaired image obtained by the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] An embodiment of the present invention discloses an image blind repair method based on self-information and prediction mask enhancement. As Figure 1 shown, it specifically includes the following steps:

[0058] Convert the damaged image into feature tokens;

[0059] Use the first encoding module composed of Transformer Blocks to encode the feature tokens to obtain encoded features;

[0060] Perform mask prediction on the damaged image, and use the prediction mask and the damaged image to generate an information-enhanced feature that only retains the pixels in the undamaged area;

[0061] Use the second encoding module composed of information-enhanced Transformer Blocks to encode the encoded features and the information-enhanced features to obtain information-enhanced encoded features;

[0062] Use the third encoding module composed of Transformer Blocks to re-encode the information-enhanced encoded features to obtain re-encoded features;

[0063] Restore the re-encoded features to multi-channel features of the same size as the damaged image through a scale recovery and refinement module;

[0064] Perform channel filtering on the multi-channel features through channel recovery to obtain the repaired image.

[0065] Next, each step of the present invention will be described in detail.

[0066] (1) First, in the embodiment of the present invention, it is necessary to perform feature conversion on the damaged image to be repaired, and convert the damaged image to be repaired into two-dimensional feature tokens that can be directly used by the model.

[0067] Assume the damaged image is where c represents the number of channels, and h and w represent the height and width of the image respectively. In this step, the image is dimensionally transformed to obtain Then, through steps such as linear mapping, encoded features are obtained, where e represents the encoding dimension, The specific process of this step is as follows:

[0068] The damaged image I with size h*w*c d is divided into a number of image patches I' of the same size d , and the size of each image patch is p*p*c. Here, h represents the height of the damaged image I d , w represents the width of the damaged image I d ; p represents the height or width of the image patch; c represents the number of channels.

[0069] Each image patch I' d is flattened into a one-dimensional feature vector of p 2 *c. In a specific embodiment, the damaged image can be segmented into image patches with a size of 4*4*3. Here, 3 represents that the number of channels c is 3. Then, the one-dimensional feature vector after flattening each image patch is a one-dimensional feature vector of 16*3.

[0070] The flattened one-dimensional feature vector is projected into a set D-dimensional space by a linear layer to obtain the Patch Embedding sequence corresponding to each image patch I' d .

[0071] The Patch Embedding sequences of all image patches I' d are combined to obtain the Patch Embeddings sequence, that is, the embedding sequence of the damaged image I d .

[0072] A class token is added to the embedding sequence of the damaged image I d . Position encoding is performed on the embedding sequence with the class token added to obtain the feature token of the damaged image I d .

[0073] (2) Next, using the two-dimensional feature token obtained in the previous step as the input, the first encoding module composed of Transformer Blocks is used to encode the two-dimensional feature token to obtain the encoded feature.

[0074] In the embodiment of the present invention, the first encoding module is composed of two TransformerBlock units with the same structure and stacked on each other, which can be respectively denoted as the first TransformerBlock unit and the second TransformerBlock unit. Among them, the first TransformerBlock unit is used to perform primary encoding on the feature token, and the second TransformerBlock unit is used to perform secondary encoding on the primary encoding result of the first TransformerBlock unit to obtain the encoded feature of the first encoding module.

[0075] The structures of the TransformerBlock units in the present invention are the same. As Figure 2 shown, the encoding process of the first TransformerBlock (TSBlock) unit for initially encoding the feature tokens is as follows:

[0076] F T1 = TSBlock(F embedding )

[0077] = MLP(Norm(MHA(Norm(F embedding )) + F embedding ))

[0078] where F T1 represents the initial encoding result obtained by the first TransformerBlock for initially encoding the feature tokens, TSBlock() represents the processing operation of the TransformerBlock unit, F embedding represents the feature tokens obtained by converting the damaged image, Norm() represents the normalization operation, MHA() represents the multi-head attention operation, and MLP() represents the multi-layer perceptron operation.

[0079] In the embodiments of the present invention, the number of attention heads in the multi-head attention mechanism in the Transformer Block unit can specifically be two or more.

[0080] (3) While encoding the feature tokens using the first encoding module, taking the damaged image to be repaired as the input, performing prediction of the damaged area mask, and generating an information-enhanced feature that only retains the non-damaged area using the predicted mask and the damaged image.

[0081] The implementation manner of this process is as follows:

[0082] Using a trained deep learning model such as Mask-RCNN, U-Net, ConvMAE, etc. to perform mask prediction on the damaged image to obtain the predicted mask of the damaged area; obtaining the damaged area and the non-damaged area in the damaged image according to the predicted mask of the damaged area; retaining the pixel values of the non-damaged area in the damaged image, and replacing the pixel values of the damaged area with fixed values to obtain an information-enhanced feature that only retains the pixel values of the non-damaged area in the damaged image.

[0083] Specifically, taking as the input, the deep learning model can first expand the number of input channels through channel convolution, secondly encode the input features through four cascaded downsampling operations, then perform mask prediction feature recovery through four upsampling operations, further reduce the number of channels through channel convolution, and finally obtain the predicted mask M preFinally, the information enhancement feature F composed of the features (pixel values) of only the undamaged areas of the image itself is obtained. enhanced The acquisition formula is as follows:

[0084] F enhanced =(1 - M pre )×I d +M pre (4)

[0085] (4) Use the second encoding module composed of information enhancement TransformerBlock to encode the encoding feature F obtained by the first encoding module in step (2) CT1 and the information enhancement feature F obtained in step (3) enhanced to obtain the information enhancement encoding feature.

[0086] In the embodiment of the present invention, the second encoding module composed of information enhancement TransformerBlock is composed of two information enhancement TransformerBlock units with the same structure and stacked on each other, which are respectively denoted as the first information enhancement TransformerBlock unit and the second information enhancement TransformerBlock unit; referring to Figure 3 as described above, the first information enhancement TransformerBlock (TSBlock) unit performs primary encoding on the encoding feature F output by the first encoding module CT1 and the information enhancement feature F that only retains the pixel values of the undamaged areas enhanced to obtain the primary encoding result F IET1 , and then the second information enhancement TransformerBlock (TSBlock) unit performs secondary encoding on the primary encoding result F obtained by the first information enhancement TransformerBlock unit IET1 and the information enhancement feature F enhanced to obtain the output result of the second encoding module, that is, the information enhancement encoding feature F CT2 .

[0087] Referring to Figure 3 , the first information enhancement TransformerBlock unit performs primary encoding on the encoding feature F CT1 and the information enhancement feature F enhanced , which may specifically include the following encoding process:

[0088] F temp1 =MHA(Norm(F CT1 ))+F CT1 ;

[0089] F IET1= Linear(Cat(Norm(MLP(Norm(F temp1 )) + F temp1 ), F enhanced ));

[0090] Among them, F temp1 represents the intermediate encoded feature, F IET1 represents the initial encoding result, F CT1 represents the encoded feature obtained by the first encoding module, F enhanced represents the information enhancement feature, Norm() represents the normalization operation, MHA() represents the multi-head attention operation, MLP() represents the multi-layer perceptron operation, Cat() represents the concatenation operation, and Linear() represents the linear mapping.

[0091] (5) After obtaining the information-enhanced encoded feature F CT2 output by the second encoding module in step (4), use the third encoding module composed of TransformerBlock to re-encode the encoded feature with enhanced information to obtain the re-encoded feature F encoder .

[0092] In the embodiment of the present invention, the third encoding module is composed of five TransformerBlock units with the same structure and stacked on each other; specifically, the first Transformer Block unit receives the encoded feature F CT2 with enhanced information output by the second encoding module; the output result of the last Transformer Block unit is used as the re-encoded feature F encoder output by the third encoding module; in the middle Transformer Block units, the output result of the previous Transformer Block unit is input to the next Transformer Block unit. The specific structure and encoding process of each Transformer Block unit can refer to the first Transformer Block unit in the first encoding module, only the input and output are different.

[0093] (6) After obtaining the re-encoded feature F encoder , use the scale recovery and refinement module to restore the re-encoded feature to a multi-channel feature with the same size as the damaged image to be repaired.

[0094] In the embodiment of the present invention, the scale restoration and refinement module package is composed of two scale restoration and refinement units with the same structure and superimposed on each other, which are respectively denoted as the first scale restoration and refinement unit and the second scale restoration and refinement unit; among them, the first scale restoration and refinement unit is used to perform preliminary dimension elevation and refinement on the recoded features output by the third coding module to obtain the preliminary dimension elevation and refinement result, and then the second scale restoration and refinement unit performs secondary dimension elevation and refinement on the preliminary dimension elevation and refinement result output by the first scale restoration and refinement unit to obtain multi-channel features consistent with the size of the damaged mural image.

[0095] In the embodiment of the present invention, as Figure 4 shown, the first scale restoration and refinement unit performs preliminary dimension elevation and refinement on the recoded features output by the third coding module, which specifically includes the following operation process:

[0096] F recovery1 = Relu(BatchNorm(Conv(Upsampling(F encoder ))))

[0097] where F recovery1 represents the preliminary dimension elevation and refinement result, F encoder represents the recoded features obtained by the third coding module, Upsampling() represents the upsampling operation, Conv() represents the convolution operation, BatchNorm() represents the batch normalization, and Relu() represents the Relu activation operation.

[0098] Then, the second scale restoration and refinement unit with the same structure is used to perform secondary dimension elevation and refinement on the preliminary dimension elevation and refinement result F recovery1 to obtain multi-channel features F recovery consistent with the size of the damaged mural image.

[0099] In this step, first, the coded features F encoder obtained in step (5) are restored in scale by double upsampling, the channels are recycled and locally refined through the convolution operation, and finally, the gradient disappearance is alleviated and the network generalization ability is improved through the BatchNorm normalization and the Relu activation operation.

[0100] (7) Finally, channel filtering is performed on the multi-channel features through channel restoration to obtain the restored image. Specifically, in this step, the multi-channel features F recovery can be restored to the restored image through channel convolution

[0101] In specific applications, the method described in the present invention can be implemented by a technical machine system including the above model.

[0102] To further demonstrate the superiority of the method proposed in the present invention, in the specific embodiments of the present invention, the blind image restoration method disclosed in the present invention was also experimentally compared with three other leading blind restoration methods in recent years for restoration on a certain mural dataset.

[0103] The objective experimental results are shown in Table 1, where PSNR represents peak signal-to-noise ratio, SSIM represents structural similarity, and the higher the values of the above two indicators, the better the quality of the reconstructed image. L1 represents the difference between the pixels of the restored image and the original image, and the smaller the value, the better. By analyzing the data in the table, it can be seen that the present invention has obtained the best values in all three indicators and has a better restoration effect.

[0104] Table 1 Objective Results of the Comparative Experiment of the Present Invention

[0105]

[0106] The subjective experimental results are as Figure 5 shown, where Figure 5 (a) is the damaged mural image, Figure 5 (b), (c), and (d) are the restored images of the TransCNN-HAE, BOWN, and DT methods respectively. Figure 5 (e) is the restored image of the method of the present invention. By comparison, it can be seen that the method proposed in this paper has fewer artifacts and the restoration results have better visual consistency.

[0107] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0108] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A blind image restoration method based on self-information and prediction mask enhancement, characterized in that: The following steps are involved: Convert the damaged image into feature tokens; Encode the feature token using the first encoding module composed of TransformerBlock to obtain an encoded feature; Perform mask prediction on the damaged image, and use the predicted mask and the damaged image to generate information enhancement features that only retain pixels in the undamaged area; The second encoding module composed of information enhancement TransformerBlock is used to encode the encoding features and information enhancement features to obtain information enhancement encoding features; The information enhancement coding features are re-encoded using the third coding module composed of TransformerBlock to obtain the re-encoded features; The re-encoded features are restored to multi-channel features of the same size as the damaged image through the scale recovery and refinement module; Channel filtering is performed on the multi-channel features through channel restoration to obtain a restored image.

2. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 1, characterized in that: Convert the damaged image into feature tokens, including: The damaged image I with the size of h*w*c d Divide into several image blocks of the same size I' d , the size of each image block is p*p*c, h represents the damaged image I d The height of w represents the damaged image I d width; p represents the height or width of the image block; c represents the number of channels; Each image block I' d Flatten to p 2 * One-dimensional eigenvector of c; Use the linear layer to project the flattened one-dimensional feature vector into the set D-dimensional space to obtain each image block I' d The corresponding Patch Embedding sequence; All image blocks I' d The Patch Embedding sequence is combined to obtain the Patch Embeddings sequence, i.e., the damaged image I d The embedding sequence of For damaged image I d Add category token to the embedding sequence; Perform position encoding on the embedded sequence with added category tokens to obtain the damaged image I d The feature token.

3. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 1, characterized in that: In the step of encoding the feature token using a first encoding module composed of TransformerBlock, the first encoding module includes two TransformerBlock units with the same structure and superimposed on each other, denoted as a first TransformerBlock unit and a second Transformer Block unit; The first TransformerBlock unit is used to perform initial encoding on the feature token.

4. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 3, characterized in that: The first TransformerBlock unit performs initial encoding on the feature token, which specifically includes the following encoding process: F T1 =TSBlock(F embedding ) =MLP(Norm(MHA(Norm(F embedding ))+F embedding )) Among them, F T1 Indicates the initial encoding result, TSBlock() indicates the TransformerBlock unit processing operation, F embedding It represents the feature token obtained by converting the damaged image, Norm() represents the normalization operation, MHA() represents the multi-head attention operation, and MLP() represents the multi-layer perceptron operation.

5. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 1, characterized in that: The damaged area mask is predicted for the damaged image, and the predicted mask and the damaged image are used to generate information enhancement features that only retain the undamaged area, including: Use the trained deep learning model to predict the mask of the damaged image and obtain the predicted mask of the damaged area; Obtaining damaged areas and undamaged areas in the damaged image according to the predicted mask of the damaged area; The pixel values ​​of the undamaged area in the damaged image are retained, and the pixel values ​​of the damaged area are replaced with fixed values, so as to obtain the information enhancement feature that only retains the pixel values ​​of the undamaged area in the damaged image.

6. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 1, characterized in that: In the step of encoding the encoding features and the information enhancement features using a second encoding module composed of an information enhancement TransformerBlock, the second encoding module includes two information enhancement TransformerBlock units having the same structure and superimposed on each other, which are denoted as a first information enhancement TransformerBlock unit and a second information enhancement TransformerBlock unit; Among them, the first information enhancement TransformerBlock unit is used to initially encode the coding features and information enhancement features, and the second information enhancement TransformerBlock unit is used to re-encode the initial encoding result and information enhancement features of the first information enhancement TransformerBlock unit.

7. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 6, characterized in that: The first information enhancement TransformerBlock unit performs initial encoding on the encoding features and the information enhancement features, specifically including the following encoding process: F temp1 =MHA(Norm(F CT1 ))+F CT1 ; F IET1 =Linear(Cat(Norm(MLP(Norm(F temp1 ))+F temp1 ),F enhanced )); Among them, F temp1 represents the intermediate encoding feature, F IET1 Indicates the initial encoding result, F CT1 Indicates that the first encoding module obtains the encoding feature, F enhanced Represents information enhancement features, Norm() represents normalization operation, MHA() represents multi-head attention operation, MLP() represents multi-layer perceptron operation, Cat() represents connection operation, and Linear() represents linear mapping.

8. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 1, characterized in that: In the step of re-encoding the information enhancement coding feature using a third coding module composed of Transformer Block, the third coding module includes five Transformer Block units with the same structure and superimposed on each other.

9. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 1, characterized in that: In the step of restoring the re-encoded features to multi-channel features of the same size as the damaged image through a scale restoration and refinement module, the scale restoration and refinement module includes two scale restoration and refinement units with the same structure and superimposed on each other, which are denoted as a first scale restoration and refinement unit and a second scale restoration and refinement unit; The first scale recovery and refinement unit is used to perform preliminary dimensionality enhancement and refinement on the re-encoded features output by the third encoding module to obtain preliminary dimensionality enhancement and refinement results.

10. The method for blind image restoration based on self-information and prediction mask enhancement according to claim 9, characterized in that: The first scale recovery and refinement unit performs preliminary dimension enhancement and refinement on the re-encoded features output by the third encoding module, specifically including the following operation processes: F recovery1 =Relu(BatchNorm(Conv(Upsampling(F encoder )))) Among them, F recovery1 represents the preliminary dimension enhancement and refinement results, F encoder It represents the re-encoded features obtained by the third encoding module, Upsampling() represents the upsampling operation, Conv() represents the convolution operation, BatchNorm() represents batch normalization, and Relu() represents the Relu activation operation.

Citation Information

Patent Citations

  • Image blind restoration method based on semantic inconsistency detection

    CN114897738A

  • Image blind restoration method and device

    CN117333397A