Image restoration method and device, electronic equipment and storage medium
The image is extracted, self-attention coding and decoding through pre-trained repair models, and the problem of manual hypothesis degradation process is solved in the prior art, achieving better image repair effects.
Patent Information
- Application Number
- CN202411774657.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-06
Smart Images

Figure CN119941574A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an image restoration method, device, electronic equipment and storage medium. Background Art
[0002] Image restoration refers to the process of reconstructing lost or damaged parts of images and videos. It can usually be used in image beautification, image generation, video editing, video generation and other fields. It is a common image processing method.
[0003] Existing image restoration methods are usually implemented based on the type of image degradation. Specifically, first obtain the image to be restored. Then, artificially assume the image degradation process and establish a corresponding mathematical model. For example, assuming that the image to be restored is an image obtained after the original image is contaminated by noise, a Gaussian denoising model can be artificially constructed. After that, use an optimization algorithm to solve the model parameters to obtain an adjusted mathematical model. Finally, use the adjusted mathematical model to restore the image to be restored.
[0004] However, when using artificially constructed mathematical models to repair images, it is necessary to manually estimate the causes of degradation, which makes it difficult to use the information of the image itself to repair it, and thus makes it impossible to repair and obtain a more ideal target image. Summary of the invention
[0005] In view of this, the embodiments of the present invention are committed to providing an image restoration method, device, electronic device and storage medium, which use a pre-trained restoration model to repair the image to be restored, so as to solve the problem in the prior art that artificial assumptions about the image degradation process are needed to construct a mathematical model, resulting in poor restoration effect.
[0006] The present invention provides an image restoration method, comprising:
[0007] Determine the image to be repaired;
[0008] Input the image to be repaired into the feature extraction layer of the pre-trained repair model to obtain the image features output by the feature extraction layer;
[0009] A self-attention encoding operation is performed based on the image features through the encoder of the restoration model to obtain a deep encoding feature of the image to be restored, wherein the self-attention encoding operation is: for each element in the first feature to be encoded, a second feature to be encoded is obtained based on other elements belonging to the same row as the element, and for each element in the second feature to be encoded, an encoded feature is obtained based on other elements belonging to the same column as the element in the second feature to be encoded;
[0010] Through the decoder of the repair model, a self-attention decoding operation is performed on the deep coding features to obtain the decoding features of the image to be repaired, and based on the decoding features, a target image is obtained, wherein the target image is the image after the image to be repaired is repaired.
[0011] Optionally, the repair model is trained in the following manner:
[0012] Obtain sample images corresponding to a plurality of degradation types, and determine the original images corresponding to the sample images as the annotations corresponding to the sample images;
[0013] Taking each sample image as input, inputting it into the feature extraction layer of the restoration model to be trained, and obtaining each sample feature output by the feature extraction layer;
[0014] Input each sample feature into the encoder of the restoration model to obtain the deep coding features corresponding to each sample feature output by the encoder;
[0015] The decoder of the restoration model decodes the depth coding features corresponding to each sample feature, and obtains the restoration result of each sample image according to the decoded features;
[0016] The restoration model is trained based on the gap between the restoration results of each sample image and the annotations corresponding to each sample image.
[0017] Optionally, a self-attention encoding operation is performed based on the image features through the encoder of the restoration model to obtain deep encoding features of the image to be restored, including:
[0018] Perform a self-attention encoding operation on the image feature to obtain a first feature, and downsample the first feature to obtain a downsampled first feature;
[0019] Perform self-attention encoding on the downsampled first feature to obtain the second feature;
[0020] By repairing the decoder of the model, self-attention decoding is performed on the deep encoding features to obtain the decoded features of the target image, including:
[0021] Performing a self-attention decoding operation based on the second feature to obtain a third feature, performing upsampling based on the third feature to obtain an upsampled third feature, and performing a fusion operation based on the upsampled third feature and the first feature to obtain a fused third feature;
[0022] A self-attention decoding operation is performed on the fused third feature to obtain the fourth feature, which is used as the decoding feature of the image.
[0023] Optionally, the encoder and the decoder include a high-dimensional feature mapping structure, the high-dimensional feature mapping structure includes a first convolution branch and a second convolution branch;
[0024] Through the encoder of the repair model, a self-attention encoding operation is performed based on the image features to obtain the deep encoding features of the image to be repaired, including:
[0025] Obtaining a first convolution feature of the image feature through a first convolution branch of the high-dimensional feature mapping structure;
[0026] Obtaining a second convolution feature of the image feature through a second convolution branch of the high-dimensional feature mapping structure, wherein the convolution depths of the first convolution branch and the second convolution branch are different;
[0027] Perform element-wise multiplication on the first convolution feature and the second convolution feature to obtain a high-dimensional feature;
[0028] Through the attention network, self-attention encoding operations are performed on high-dimensional features to obtain the deep encoding features of the image to be repaired.
[0029] Optionally, a self-attention encoding operation is performed based on the image features through the encoder of the restoration model to obtain deep encoding features of the image to be restored, including:
[0030] Determine a first feature to be encoded according to the image feature, and for each element in the first feature to be encoded, determine a first row vector centered on the element and including elements matching the window length according to a preset window length;
[0031] Determine the attention weights corresponding to the elements in the first row vector according to the elements corresponding to the channels in the first feature to be encoded;
[0032] Determine the first encoding of the element according to the first row vector and the attention weights corresponding to each element in the first row vector, and obtain the intermediate feature of the first feature to be encoded according to the first encoding corresponding to each element in the first feature to be encoded;
[0033] For each element in the intermediate feature, determine the sampling interval according to the window length, perform equal-interval sampling with the element as the center, and determine a second row vector with the element as the center and the number of elements matching the window length;
[0034] According to the elements corresponding to each channel in the intermediate feature, determine the attention weights corresponding to each element in the second row vector;
[0035] Determine the second encoding of the element according to the second row vector and the attention weights corresponding to each element in the second row vector, and obtain the second feature to be encoded according to the second encoding corresponding to each element in the intermediate feature;
[0036] A vertical attention encoding operation is performed on the second feature to be encoded to obtain a deep encoding feature of the image to be repaired.
[0037] Optionally, the sampling interval is determined according to the window length, including:
[0038] According to the window length, the distance from the element at the center of the first row vector to the element at the edge of the first row vector is determined as the sampling interval.
[0039] Optionally, training a restoration model according to the gap between the restoration result of each sample image and the annotation corresponding to each sample image includes:
[0040] Determine a pixel value difference between the restoration result of each sample image and the annotation corresponding to each sample image as a first difference, and determine a frequency domain difference between the restoration result of each sample image and the annotation corresponding to each sample image as a second difference;
[0041] According to the first gap and the second gap, the loss is determined, and the repair model is trained with minimizing the loss as the optimization goal.
[0042] The present invention provides an image restoration device, comprising:
[0043] A determination module, used for determining an image to be repaired;
[0044] A feature extraction module is used to input the image to be repaired into the feature extraction layer of the pre-trained repair model to obtain the image features output by the feature extraction layer;
[0045] The encoding module is used to perform a self-attention encoding operation based on the image features through the encoder of the restoration model to obtain the encoding features of the image to be restored, wherein the self-attention encoding operation is: for each element in the first feature to be encoded, based on other elements belonging to the same row as the element, a second feature to be encoded is obtained, and for each element in the second feature to be encoded, based on other elements belonging to the same column as the element in the second feature to be encoded, an encoded deep encoding feature is obtained;
[0046] The decoding module is used to perform a self-attention decoding operation on the deep coding features through the decoder of the repair model to obtain the decoding features of the target image, and obtain the target image based on the decoding features, wherein the target image is the image after the image to be repaired is repaired.
[0047] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned image restoration method is implemented.
[0048] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the image restoration method is implemented when the processor executes the program.
[0049] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0050] In the image restoration method provided in this specification, feature extraction is performed on the image to be restored through the feature extraction layer of the pre-trained restoration model to obtain image features, and then a self-attention encoding operation is performed based on the image features through the encoder and decoder of the restoration model to obtain deep coding features that further encode the features of the image to be restored, and the deep coding features are decoded to obtain decoded features, and finally the restored image, i.e., the target image, is obtained based on the decoded features.
[0051] The image restoration method can perform a self-attention encoding operation based on the image's own features in the encoder to obtain a deep coding feature that further encodes the image's own features to be restored, and perform a self-attention decoding operation on the deep coding feature in the decoder to obtain a decoding feature, and then obtain a target image based on the decoding feature. In other words, the image restoration method can restore an ideal target image based on the image's own features to be restored. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 Shown is a flow chart of the image restoration method provided in this specification;
[0053] Figure 2 A schematic diagram of the network architecture of the repair model provided in this manual;
[0054] Figure 3 A schematic diagram of the structure of the high-dimensional feature mapping structure provided in this specification;
[0055] Figure 4 A schematic diagram of a process for determining a second feature to be encoded provided in this specification;
[0056] Figure 5 The figure shows a schematic diagram of the structure of the image restoration device provided in this specification;
[0057] Figure 6 The corresponding Figure 1 electronic equipment. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] It should be noted that all actions to obtain signals, information or data in this manual are performed in compliance with the relevant local data protection laws and policies and with the authorization given by the owner of the corresponding device.
[0060] An embodiment of the present invention provides an image restoration method, which uses a feature extraction layer of a pre-trained restoration model to extract features of an image to be restored to obtain image features, and then uses an encoder and a decoder of the restoration model to perform self-attention encoding operations based on the image features to obtain deep coding features that further encode the features of the image to be restored, and decodes the deep coding features to obtain decoded features, and finally obtains a target image based on the decoded features.
[0061] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0062] Figure 1 The following is a flowchart of an image restoration method in this specification, which specifically includes the following steps:
[0063] S100: Determine an image to be restored.
[0064] In one or more embodiments provided in this specification, the image restoration method can be applied to a server with a restoration model deployed thereon, and can also be applied to electronic devices such as smart terminals that can interact with a server with a restoration model deployed thereon. For the convenience of description, the following description will be based on the execution process of the image restoration method executed by the server as an example.
[0065] Generally, the server may first determine an image to be restored.
[0066] In an example, the server may receive an image restoration request sent by other devices, and parse the image restoration request to determine an image carried in the image restoration request as an image to be restored.
[0067] In one example, the server may pre-store images to be repaired. When the repair condition is triggered, the server may select any image to be repaired from the pre-stored images to be repaired as the image to be repaired for which the image repair method described in this specification needs to be executed.
[0068] In one example, the image to be repaired may be an image that is overexposed due to limitations of photographing conditions, or an image whose quality is reduced due to limitations of transmission bandwidth.
[0069] In one example, the image to be repaired may be an image that is partially damaged (eg, pixels are set to 0).
[0070] The specific degradation type corresponding to the image to be repaired can be set as needed, and this specification does not limit this.
[0071] S102: Inputting the image to be repaired into the feature extraction layer of a pre-trained repair model to obtain image features output by the feature extraction layer, wherein the repair model is trained by using sample images corresponding to a plurality of degradation types and annotations corresponding to the current image.
[0072] In one or more embodiments provided in this specification, the image restoration method can project the information of the image to be restored into a high-dimensional space through a pre-trained restoration model to obtain image features, and then perform self-attention encoding on the image features to further amplify the features of the image to be restored itself, so as to restore the image based on its own features.
[0073] Therefore, the server can determine the image features corresponding to the image to be repaired through the repair model.
[0074] In one example, the restoration model includes a feature extraction layer, an encoder and a decoder, so the server can take the image to be restored as input and input it into the feature extraction layer of the restoration model to obtain the image features output by the feature extraction layer.
[0075] In one example, the restoration model is trained by a server for training the model according to sample images corresponding to a plurality of degradation types and annotations corresponding to the sample images.
[0076] S104: Perform a self-attention encoding operation based on the image features through the encoder of the repair model to obtain the deep coding features of the image to be repaired, wherein the self-attention encoding operation is: for each element in the first feature to be encoded, obtain the second feature to be encoded based on other elements belonging to the same row as the element, and for each element in the second feature to be encoded, obtain the encoded feature based on other elements belonging to the same column as the element in the second feature to be encoded.
[0077] In one or more embodiments provided in this specification, as mentioned above, the model result of the repair model is a codec structure. Therefore, after obtaining the image features, the server can perform self-attention encoding operations based on the image features through the encoder of the repair model, thereby further amplifying the features of the image to be repaired itself and obtaining deep coding features. Among them, amplifying the features of the image to be repaired itself is not an intuitive operation such as increasing the size of the image to be repaired or increasing the brightness of the image to be repaired in a physical sense, but refers to enhancing or highlighting certain features in the image through a self-attention mechanism. In the image repair scenario used in this specification, the further amplification of the features of the image to be repaired itself can specifically be to enhance or highlight the features of the image to be repaired that facilitate image repair through a self-attention mechanism.
[0078] Therefore, the server can take the image features as input and input them into the encoder of the pre-trained restoration model, in which the image features are self-attention encoded to obtain deep encoding features.
[0079] Specifically, first, the server may take the image feature as input and input it into the encoder of the restoration model.
[0080] Then, in the encoder, the server may use the image feature as the first feature to be encoded, and for each element in the first feature to be encoded, perform self-attention encoding on the element according to other elements in the first feature to be encoded that belong to the same row as the element, to obtain a second feature to be encoded. The second encoded feature is a feature obtained by the server performing a horizontal self-attention encoding operation on the first feature to be encoded.
[0081] Afterwards, the server may perform self-attention encoding on each element in the second feature to be encoded according to other elements in the second feature to be encoded that belong to the same column as the element, to obtain a third encoding feature. The third encoding feature is a feature obtained by the server performing vertical attention encoding on the second feature to be encoded.
[0082] Finally, the server may use the third coding feature as the depth coding feature of the image to be restored output by the encoder.
[0083] Optionally, in one example, the server may first perform vertical attention coding on the image features to obtain a second feature to be coded, then perform horizontal attention coding on the second feature to be coded to obtain a third coding feature, and finally use the third coding feature as the depth coding feature of the image to be repaired. The vertical attention coding operation is: for each element, self-attention coding is performed on the element according to other elements in the same column as the element. The horizontal attention coding operation is: for each element, self-attention coding is performed on the element according to other elements in the same row as the element.
[0084] In one example, the horizontal attention encoding operation is: for each element, the element is attention-encoded according to the element and at least one other element in the same row as the element. The vertical attention encoding operation is: for each element, the element is attention-encoded according to the element and at least one other element in the same column as the element.
[0085] In one example, the network structure inside the encoder may be a star network structure.
[0086] S106: Perform a self-attention decoding operation on the deep coding features through the decoding layer of the repair model to obtain the decoding features of the image to be repaired, and obtain a target image based on the decoding features, wherein the target image is the image after the image to be repaired is repaired.
[0087] In one or more embodiments provided in the present specification, as described above, the model result of the repair model is a codec structure. Therefore, after obtaining the depth coding feature, the server can decode the depth coding feature through the decoder of the repair model to obtain the decoding feature, and then obtain the target image.
[0088] Specifically, the server can take the deep coding feature as input and input it into the decoder of the pre-trained repair model. In the decoder, the server can perform a self-attention decoding operation based on the deep coding feature to obtain a decoded feature.
[0089] Afterwards, the server can obtain the target image according to the decoded features.
[0090] In one example, the server may input the deep coding features into the decoder, in which a self-attention decoding operation is performed based on the deep coding features to obtain decoding features, and based on the decoding features, a target image is obtained as output data of the decoder.
[0091] In one example, the self-attention decoding operation is: for each element in the first feature to be decoded, a second feature to be decoded is obtained based on other elements belonging to the same row as the element, and for each element in the second feature to be decoded, a decoded feature is obtained based on other elements belonging to the same column as the element in the second feature to be decoded.
[0092] In one example, the server may directly use the deep coding feature as the first feature to be decoded, and perform a horizontal attention coding operation on the first feature to be decoded to obtain a second feature to be decoded, and then perform a vertical attention coding operation on the second feature to be decoded to obtain a third decoded feature, which is output as a decoded feature.
[0093] It should be noted that the internal network structures of the encoder and decoder can be the same or similar network structures. During the training process of the repair model, the internal parameters of the encoder and decoder can be adjusted to different parameters to achieve different functions.
[0094] based on Figure 1 The image restoration method extracts features of the image to be restored through the feature extraction layer of the pre-trained restoration model to obtain image features, and then performs self-attention encoding operations based on the image features through the encoder and decoder of the restoration model, thereby obtaining deep coding features that further encode the features of the image to be restored, and decoding the deep coding features to obtain decoding features, and finally obtaining the target image based on the decoding features. The method can perform self-attention encoding operations based on image features to obtain deep coding features that further encode the features of the image to be restored, and perform self-attention decoding operations on the deep coding features to obtain decoding features, and then obtain the target image based on the decoding features. Without the need to construct a mathematical model, the ideal target image can be restored based only on the features of the image itself.
[0095] At the same time, compared with the preset square convolution kernel, for each element in the image feature, a matrix with the same size as the convolution kernel and centered on the element is determined, and then the element is self-attention encoded by the matrix dot product of the matrix and the convolution kernel. In the case of the same window, since the present application only needs to self-attention encode the element according to the elements in the same row and the same column in the window, fewer computing resources are required.
[0096] Furthermore, the repair model in the aforementioned steps S100 to S106 can be trained in the following manner:
[0097] First, the server used for training the model can obtain sample images corresponding to each degradation type, and determine the annotations corresponding to each sample image. For each sample image, the annotation of the sample image is the original image of the sample image, and the sample image is the degradation result obtained by performing degradation processing on the original image of the sample image.
[0098] In one example, the degradation processing may be compressing the image quality, adding noise to the original image, deleting local pixels in the original image, etc.
[0099] In an example, for each sample image, the sample image corresponds to at least one degradation type.
[0100] Secondly, the server for training the model can take each sample image as input and input the sample image into the feature extraction layer of the restoration model to be trained to obtain the sample features corresponding to the sample image output by the feature extraction layer.
[0101] Then, the server for training the model can input the sample features into the encoder of the repair model to obtain the deep coding features of the sample features output by the encoder.
[0102] Afterwards, the server used to train the model can input the deep coding features into the decoder of the repair model, and the decoder decodes the deep coding features of the sample features to obtain the repair result of the sample image.
[0103] Finally, according to the difference between the restoration results of each sample image and its annotation, the server for training the model can determine the loss and train the restoration model according to the loss. The trained restoration model can be used to restore the image to be restored.
[0104] It should be noted that the server used for training the model and the server for executing the image restoration method may be the same electronic device or different electronic devices.
[0105] In addition, according to the gap between the restoration results of each sample image and its annotation, the server for training the model can determine the loss, and when training the restoration model according to the loss, the server can determine the pixel value gap between the restoration results of each sample image and its annotation as the first gap, and determine the gap between the spectrum graph of the restoration results of each sample image and its annotation as the second gap. The server then determines the loss based on the first gap and the second gap, and trains the restoration model with minimizing the loss as the optimization goal. In this way, not only the texture features of the image can be learned, but also the frequency domain features of the image can be learned, and the image to be restored can be restored based on the frequency domain features, with a better restoration effect.
[0106] Furthermore, in the embodiments of the present specification, the repair model can be constructed by the U-Net architecture, taking the need to process two features of different sizes in the repair model as an example. Then step S104 can be: performing a self-attention encoding operation on the image feature to obtain a first feature, and downsampling the first feature to obtain a downsampled first feature; performing a self-attention encoding operation on the downsampled first feature to obtain a second feature; performing a self-attention decoding operation on the second feature to obtain a third feature, and upsampling the third feature to obtain an upsampled third feature, and performing a fusion operation based on the upsampled third feature and the first feature to obtain a fused third feature; performing a self-attention decoding operation on the fused third feature to obtain a fourth feature. The fourth feature is used as a decoding feature of the image feature.
[0107] In one example, the server may use the image feature as an input to an encoder of the restoration model, in which a self-attention encoding operation is performed on the image feature to obtain a first feature, and the first feature is downsampled to obtain a downsampled first feature. And a self-attention encoding operation is performed on the downsampled first feature to obtain a second feature.
[0108] In an example, the server may use at least one of the first feature and the second feature as a depth coding feature.
[0109] In one example, the server may input the second feature as a deep coding feature into the decoder of the repair model, in which a self-attention decoding operation is performed on the second feature to obtain a third feature, and the third feature is upsampled to obtain an upsampled third feature, and a fusion operation is performed based on the upsampled third feature and the first feature to obtain a fused third feature; a self-attention encoding operation is performed on the fused third feature to obtain a fourth feature, and the fourth feature is used as the deep coding feature of the image feature.
[0110] In an example, the first feature, the fourth feature, and the upsampled third feature have the same dimension, and the second feature, the third feature, and the downsampled first feature have the same dimension.
[0111] In one example, the smaller the feature size, the more global features, such as texture features, can be learned when the attention operation is performed on it. The larger the feature size, the more detailed features can be learned when the attention operation is performed on it. Therefore, the deep coding features obtained based on the above features of different sizes can contain richer global features and detailed features, thereby ensuring the accuracy of the target image determined based on the deep coding features.
[0112] It should be noted that the above-mentioned first feature is a feature obtained by performing a self-attention encoding operation on the image feature, which is not the same feature as the first feature to be encoded. Similarly, the second feature is a feature obtained by performing a self-attention encoding operation on the downsampled first feature, which is not the same feature as the second feature to be encoded.
[0113] In one example, the above description only takes the example that the encoder and decoder of the repair model contain branches for processing features of two sizes respectively. However, in practice, the encoder and decoder of the repair model may also contain branches for processing features of multiple sizes respectively. Taking the number of branches as 3 as an example, the corresponding network structure of the repair model is as follows: Figure 2 shown.
[0114] Figure 2 A schematic diagram of the structure of the network architecture of the repair model provided in this specification. It can be seen that the repair model includes a feature extraction layer, an encoder and a decoder, and the encoder and decoder include the same number of downsampling sublayers and upsampling sublayers. The downsampling sublayer will reduce the resolution of the feature, and the upsampling sublayer will increase the resolution of the feature. Multiple downsampling sublayers are arranged in the order of decreasing output resolution, and multiple upsampling sublayers are arranged in the order of increasing output resolution.
[0115] In addition to the upsampling sublayer and the downsampling sublayer, the codec may also include several attention sublayers. The figure takes the codec containing a total of six attention sublayers as an example for illustration.
[0116] Therefore, the above step S104 can also be implemented by the following implementation:
[0117] First, the server may input the image feature into the first attention sublayer to obtain a first feature output by the first attention sublayer, and downsample the first feature to obtain a downsampled first feature. The dimension of the image feature and the dimension of the first feature may be C×H×W.
[0118] Secondly, the server may use the downsampled first feature as the input of the second attention sublayer, determine the second feature through the second attention sublayer, and determine the downsampled second feature. The dimensions of the second feature and the downsampled first feature may be
[0119] Then, the server may continue to use the downsampled second feature as the input of the fifth attention sublayer, and determine the fifth feature through the fifth attention sublayer. The dimensions of the fifth feature and the downsampled second feature may be
[0120] Afterwards, the server may use the fifth feature as the input of the sixth attention sub-layer, and determine the sixth feature through the sixth attention sub-layer. The dimension of the sixth feature may be
[0121] Next, the server may upsample the sixth feature, and input the upsampled sixth feature and the second feature into the fusion sublayer to obtain the fused sixth feature, and use the fused sixth feature as the input of the third attention sublayer to obtain the third feature. The dimensions of the third feature, the upsampled sixth feature, and the fused sixth feature may be
[0122] Finally, the server can continue to upsample the third feature, and obtain a fused third feature based on the upsampled third feature and the first feature, and then obtain the fourth feature through the fourth attention sublayer and the fused third feature, and the fourth feature serves as the deep coding feature corresponding to the image to be repaired.
[0123] Among them, multiple downsampling sublayers constitute a contraction path, and multiple upsampling sublayers constitute an expansion path. Therefore, through the features obtained by downsampling, the encoder can capture higher-level features. At the same time, based on upsampling, the features are restored and the upsampling results are fused with the features before downsampling. The fusion result can also retain the original spatial information. Based on the above content, the image restoration model can learn higher-dimensional information, thereby ensuring the amount of information of the determined deep coding features, so that the accurate target image can be determined based on the deep coding features, further ensuring the success rate of the image restoration.
[0124] It should be noted that: for each self-attention sublayer in the codec, the server can perform a self-attention encoding operation on the features input into the self-attention sublayer in the self-attention sublayer. The figure takes the codec including two downsampling sublayers as an example for explanation, but in practice, the codec may include only one downsampling sublayer, or may include multiple downsampling sublayers. The number of downsampling sublayers, upsampling sublayers and attention sublayers that the codec may include can be set as needed, and this specification does not limit this.
[0125] At the same time, the numbers corresponding to the names of the attention sub-layers in the above-mentioned figures do not represent the execution order of the attention sub-layers. The numbers corresponding to the names of the attention sub-layers can be set as needed, and this specification does not limit this.
[0126] In addition, generally speaking, higher-dimensional features mean more knowledge, that is, richer information. Therefore, the accuracy of the target image obtained based on the deep coding features with higher dimensions is usually higher than that of the target image obtained based on the deep coding features with lower dimensions. At present, when mapping image features to high dimensions, stacked linear convolutional layers are usually chosen to achieve this. That is to say, multiple linear convolutional layers are stacked, and each layer extracts new features based on the previous layer, thereby gradually mapping the features to a higher-dimensional space. However, in the case where the receptive field of the linear convolutional layer is limited and stacking linear convolutional layers is prone to feature redundancy, it is obvious that mapping the features to high dimensions requires more computing resources, and the model structure will also increase accordingly. On this basis, the present specification provides a feature mapping sublayer based on an improved star network, which reduces the stacking of linear network layers and obtains high-dimensional features through element-wise multiplication. Therefore, the aforementioned step S104 can be specifically the following steps:
[0127] Specifically, the encoder may include a high-dimensional feature mapping structure and an attention network, and the high-dimensional feature mapping structure includes a first convolution branch and a second convolution branch.
[0128] Therefore, the server can use the image features as inputs to the first convolution branch and the second convolution branch of the high-dimensional feature mapping structure, respectively, to obtain the first convolution feature output by the first convolution branch and the second convolution feature output by the second convolution branch. The convolution depths of the first convolution branch and the second convolution branch are different. In other words, the area of the input feature map covered by the convolution kernel is also different.
[0129] Afterwards, the server may perform element-wise multiplication on the first convolution feature and the second convolution feature to obtain a high-dimensional feature. The element-wise multiplication refers to multiplying the corresponding elements of the feature. Taking the image feature X as an example, assuming that the first convolution feature is W1X and the second convolution feature is W2X, where the dimensions of W1 and W2 are both m×n, and the dimension of the image feature X is n×p. Then the high-dimensional feature can be in, is the element in row i and column j in W1X, is the element in the i-th row and j-th column of W2X. The dimension of this high-dimensional feature is the same as that of W1X and W2X, both of which are m×p. However, the implicit feature dimension in the high-dimensional feature is larger than that of the first convolution feature and the second convolution feature, so it can represent richer information.
[0130] Finally, through the high-dimensional feature, the attention network can perform self-attention encoding operation on the high-dimensional feature to obtain deep encoding features, such as Figure 3 shown.
[0131] Figure 3 This is a schematic diagram of the structure of the high-dimensional feature mapping structure provided in this specification. In the figure, the server can normalize the image features first, taking the network layer for normalizing the image features as the layer normalization (LayerNormalization, LN) layer as an example, the normalized image features are X LN =LN(X input ).
[0132] In one example, when the high-dimensional feature mapping structure includes an LN layer, X input is the input feature of the high-dimensional feature mapping structure.
[0133] In one example, when the high-dimensional feature mapping structure does not include an LN layer, X LN The input features of this high-dimensional feature map structure.
[0134] In one example, the server may not perform normalization on the image features, but directly process the image features through a high-dimensional feature mapping structure.
[0135] Afterwards, the server can use the image features or the normalized image features as input and input them into the first convolution branch and the second convolution branch respectively to obtain the first convolution features output by the first convolution branch and the second convolution features output by the second convolution branch.
[0136] Finally, the server may perform element-wise multiplication on the first convolution feature and the second convolution feature to obtain a high-dimensional feature.
[0137] In one example, the server may perform convolution and residual connection operations on the element multiplication operation to obtain high-dimensional features.
[0138] In this way, the server can obtain that although the feature size is the same as the input feature size of the high-dimensional feature mapping structure, the elements contained therein in the feature space are more high-dimensional, and are high-dimensional features with larger implicit feature dimensions. The high-dimensional features contain richer information, and the image is restored based on the high-dimensional features, and the accuracy of the restored target image is higher.
[0139] Furthermore, in order to avoid the situation where the high-dimensional features only contain higher-dimensional features and gradually forget the low-dimensional features, the server can also fuse the operation results with the input features of the high-dimensional feature mapping structure after obtaining the operation results of the element-wise multiplication of the first convolution feature and the second convolution feature, and use the fusion results as the high-dimensional features.
[0140] In one example, the high-dimensional feature mapping structure can be an Efficient Residual Star Module (ERSM), that is, the high-dimensional feature mapping structure is a basic residual connection network structure with a context-element-aware star operation added. The star operation is an element-wise multiplication operation of the corresponding positions of the features. Compared with the addition operation (the addition operation of the corresponding positions of the features), the star operation can map the features to a higher-dimensional space. In addition, we designed context-element perception in the second convolution branch and innovatively proposed a context-element-aware star operation to achieve the purpose of quickly and efficiently mapping the features to high dimensions.
[0141] Furthermore, the high-dimensional feature mapping structure included in the attention layer can be multi-layered, and multiple high-dimensional feature mapping structures are serially connected. In other words, multiple high-dimensional feature mapping structures are stacked. Through the stacked high-dimensional feature mapping structure, it is possible to map features to a higher dimension.
[0142] In this specification, under normal circumstances, the server can, for each element in the image feature, convolve other elements in the fixed window corresponding to the element based on the fixed window to determine the code corresponding to the element, and then determine the deep coding features of the image to be repaired. However, when the window size is limited, the amount of information contained in the determined deep coding features of the image to be repaired is still not rich enough. In order to improve this problem, in this specification, when determining the code corresponding to an element, in addition to the elements contained in the fixed window, the server can also sample other elements from the elements outside the fixed window as one of the elements for self-attention encoding of the element.
[0143] Specifically, taking directly performing a horizontal self-attention encoding operation on an image feature as an example, in step S104, the server may use the image feature as the first feature to be encoded.
[0144] For each element in the first feature to be encoded, according to a preset window length, a first row vector centered on the element and containing elements whose number matches the window length is determined.
[0145] Secondly, the server may determine the attention weight corresponding to each element in the first row vector according to the elements corresponding to each channel in the first feature to be encoded.
[0146] Then, according to the attention weights corresponding to the first row vector and each element in the first row vector, the server can perform weighted summation of the attention weights corresponding to the first row vector and each element in the first row vector to obtain the first code corresponding to the element.
[0147] After obtaining the first codes corresponding to the elements, the server may determine the intermediate features of the first feature to be encoded according to the first codes corresponding to the elements and the positions of the elements.
[0148] Afterwards, the server may determine the sampling interval for each element in the intermediate feature according to the window length, and perform equal-interval sampling with the element as the center to obtain a second row vector, wherein the second row vector is a vector with the element as the center and the number of elements included matches the window length.
[0149] Therefore, after obtaining the second row vector, the server can determine the attention weights corresponding to each element in the second row vector according to the elements corresponding to each channel in the intermediate feature.
[0150] In one example, the server may perform global average pooling on each channel in the first feature to be encoded, and determine the channel codes corresponding to each channel according to the pooling result. After that, by performing convolution and other operations on the vectors composed of the channel codes corresponding to each channel, attention weights that can be used to characterize the element relationship between each channel in the first feature to be encoded are obtained. The attention weights and the first row vector are weighted and summed to obtain a first encoding that contains both information about the element relationship between channels and the relationship between the element and other elements around the element. The description of the second encoding can refer to the aforementioned description of the first encoding, and this specification will not repeat it.
[0151] Finally, according to the second row vector and the attention weights corresponding to each element in the second row vector, the server can determine the second encoding of the element in the intermediate feature. Then, according to the second encoding corresponding to each element in the intermediate feature, the server can obtain the second feature to be encoded of the first feature to be encoded. Figure 4 shown.
[0152] Figure 4 A flow chart of determining the second feature to be encoded provided in this specification. In the figure, the black squares are the elements that need to be determined for encoding this time, and the first feature to be encoded and the gray squares in the first row vector are the elements that are in the same row as the black squares and are close to the black squares. The window length in the figure is 5, and the number of elements contained in the first row vector matches the window length.
[0153] For each element in the first feature to be encoded, the server may determine a first row vector from the first feature to be encoded, and determine the attention weights corresponding to each element in the first row vector according to the first feature to be encoded, that is, Figure 4The attention weight in is 1. Therefore, the server can perform weighted summation on the first row vector and the attention weights corresponding to each element in the first row vector to obtain a first code, and the first code corresponding to each element in the first feature to be encoded can constitute the intermediate feature of the image to be repaired.
[0154] After obtaining the intermediate feature, the server can determine the sampling interval for each element in the intermediate feature according to the window length. Therefore, the server can sample the element to the left and right of the element at the preset sampling interval, with the element as the center, to obtain the second row vector. The intermediate feature and the gray squares in the second row vector are the sampled elements. The number of elements contained in the second row vector matches the window length.
[0155] Afterwards, the server can determine the attention weights corresponding to each element in the second row vector according to the intermediate features, that is, Figure 4 The attention weight in 2.
[0156] Finally, the server may perform a weighted summation of the second row vector and the attention weights corresponding to each element in the second row vector to obtain a second code, and the second code corresponding to each element in the intermediate feature may constitute a second feature to be encoded of the image to be repaired.
[0157] In one example, the above is explained by taking the horizontal self-attention encoding of the first feature to be encoded as an example. The process of horizontal self-attention encoding of the second feature to be encoded, vertical self-attention encoding of the first feature to be encoded, and vertical self-attention encoding of the second feature to be encoded can be referred to the above description, and this application will not go into details.
[0158] In this way, without increasing the size of the convolution kernel, the server can perform self-attention encoding based on more elements in the image features. In other words, without increasing the size of the convolution kernel, the present application expands the receptive field, so that the extracted deep coding features of the image to be repaired contain richer information, thereby ensuring the accuracy of the image repair.
[0159] Furthermore, the server may also determine the distance between the element at the center of the first row vector and the element at the edge of the first row vector according to the window length as the sampling interval. Assuming that the window length is K, the distance between the element at the center of the first row vector and the element at the edge of the first row vector may be Therefore, the sampling interval can be Figure 4 If K is 5, the sampling interval is 2.
[0160] In one example, the server may also be greater than Any integer is used as the sampling interval.
[0161] In this way, it can be ensured that, except for the elements that need to be self-attention encoded (i.e., the elements located at the center of the vector), the receptive fields corresponding to the elements in the second row of the vector determined have no overlap with the receptive fields corresponding to the elements in the first row of the vector, further ensuring the richness of the information content of the learned features.
[0162] It should be noted that, when the repair model is a U-net structure, for each attention sub-layer in the attention layer, the structure of the attention sub-layer can be a structure composed of a high-dimensional feature mapping structure + an attention network. The attention network can be specifically described in the above description of step S104. In other words, except for the content of the network structure, the structure of the attention layer described in the above step S104 can be the structure of the attention network. For details, please refer to the above description of step S104, and this specification does not limit this.
[0163] At the same time, as mentioned above, the internal network structure of the encoder and the decoder can be the same or similar network structure. During the training process of the repair model, the internal parameters of the encoder and the decoder can be adjusted to different parameters to achieve different functions. Similarly, the self-attention encoding operation and the self-attention decoding operation can be similar operations, the difference is only that the modules they belong to are different. Therefore, the specific steps of the self-attention decoding operation can refer to the above description of the self-attention encoding operation, and this specification will not repeat them here.
[0164] The above is an image restoration method provided by one or more embodiments of this specification. Based on the same idea, this specification also provides an image restoration device, such as Figure 5 shown.
[0165] Figure 5 The structural schematic diagram of the image restoration device provided in this specification, wherein the image restoration device may include:
[0166] The determination module 200 is used to determine the image to be restored.
[0167] The feature extraction module 202 is used to input the image to be repaired into the feature extraction layer of the pre-trained repair model to obtain the image features output by the feature extraction layer.
[0168] The encoding module 204 is used to perform a self-attention encoding operation based on the image features through the encoder of the repair model to obtain the deep coding features of the image to be repaired, wherein the self-attention encoding operation is: for each element in the first feature to be encoded, a second feature to be encoded is obtained based on other elements belonging to the same row as the element, and for each element in the second feature to be encoded, an encoded feature is obtained based on other elements belonging to the same column as the element in the second feature to be encoded.
[0169] The decoding module 206 is used to perform a self-attention decoding operation on the deep coding features through the decoder of the repair model to obtain the decoded features of the image to be repaired, and obtain the target image based on the decoded features, wherein the target image is the image after the image to be repaired is repaired.
[0170] Optionally, the device further comprises:
[0171] The training module 208 is used to train the restoration model in the following manner: obtain sample images corresponding to a plurality of degradation types, and determine the original images corresponding to each sample image as the annotations corresponding to each sample image; use each sample image as input, input it into the feature extraction layer of the restoration model to be trained, and obtain each sample feature output by the feature extraction layer; input each sample feature into the encoder of the restoration model, and obtain the depth coding features corresponding to each sample feature output by the encoder; decode the depth coding features corresponding to each sample feature through the decoder of the restoration model, and obtain the restoration results of each sample image according to the decoded features; train the restoration model according to the gap between the restoration results of each sample image and the annotations corresponding to each sample image.
[0172] Optionally, the encoding module 204 is used to: perform a self-attention encoding operation on the image feature to obtain a first feature, and downsample the first feature to obtain a downsampled first feature; perform a self-attention encoding operation on the downsampled first feature to obtain a second feature.
[0173] Optionally, the decoding module 206 is used to: perform a self-attention decoding operation based on the second feature to obtain a third feature, perform upsampling based on the third feature to obtain an upsampled third feature, and perform a fusion operation based on the upsampled third feature and the first feature to obtain a fused third feature; perform a self-attention decoding operation on the fused third feature to obtain a fourth feature, and use the fourth feature as a decoding feature of the image.
[0174] Optionally, the encoder and the decoder include a high-dimensional feature mapping structure, and the high-dimensional feature mapping structure includes a first convolution branch and a second convolution branch; the encoding module 204 is used to: obtain a first convolution feature of the image feature through the first convolution branch of the high-dimensional feature mapping structure; obtain a second convolution feature of the image feature through the second convolution branch of the high-dimensional feature mapping structure, wherein the convolution depths of the first convolution branch and the second convolution branch are different; perform element-wise multiplication operations on the first convolution feature and the second convolution feature to obtain a high-dimensional feature; and perform self-attention encoding operations on the high-dimensional feature through an attention network to obtain a deep encoding feature of the image to be repaired.
[0175] Optionally, the encoding module 204 is used to: determine a first feature to be encoded according to the image feature, and for each element in the first feature to be encoded, determine a first row vector centered on the element and containing elements whose number matches the window length according to a preset window length; determine an attention weight corresponding to each element in the first row vector according to the elements corresponding to each channel in the first feature to be encoded; determine a first encoding of the element according to the first row vector and the attention weight corresponding to each element in the first row vector, and obtain an intermediate feature of the first feature to be encoded according to the first encoding corresponding to each element in the first feature to be encoded; and for the intermediate For each element in the feature, the sampling interval is determined according to the window length, and sampling is performed at equal intervals with the element as the center to determine a second row vector with the element as the center and the number of elements matching the window length; the attention weights corresponding to each element in the second row vector are determined according to the elements corresponding to each channel in the intermediate feature; the second encoding of the element is determined according to the second row vector and the attention weights corresponding to each element in the second row vector, and the second feature to be encoded is obtained according to the second encoding corresponding to each element in the intermediate feature; a vertical attention encoding operation is performed on the second feature to be encoded to obtain a deep encoding feature of the image to be repaired.
[0176] Optionally, the encoding module 204 is used to determine, according to the window length, a distance from an element located at the center of the first row vector to an element at the edge of the first row vector as a sampling interval.
[0177] Optionally, the training module 208 is used to: determine the pixel value difference between the restoration result of each sample image and the annotation corresponding to each sample image as the first difference, and determine the frequency domain difference between the restoration result of each sample image and the annotation corresponding to each sample image as the second difference; determine the loss based on the first difference and the second difference, and train the restoration model with minimizing the loss as the optimization goal.
[0178] The specification also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above-mentioned image restoration method.
[0179] This manual also provides Figure 6 The schematic structure diagram of the electronic device shown in FIG. Figure 6As described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned image restoration method. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0180] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0181] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0182] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0183] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0184] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0186] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0188] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0189] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0190] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0191] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0192] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0193] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing nodes connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage nodes.
[0194] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0195] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. An image restoration method, characterized in that: include: Determine the image to be repaired; Inputting the image to be repaired into a feature extraction layer of a pre-trained repair model to obtain image features output by the feature extraction layer; By means of the encoder of the inpainting model, a self-attention encoding operation is performed based on the image feature to obtain a deep coding feature of the image to be inpainted, wherein the self-attention encoding operation is: for each element in the first feature to be encoded, a second feature to be encoded is obtained based on other elements belonging to the same row as the element, and for each element in the second feature to be encoded, an encoded feature is obtained based on other elements belonging to the same column as the element in the second feature to be encoded; The decoder of the restoration model performs a self-attention decoding operation on the deep coding features to obtain the decoding features of the image to be restored, and obtains a target image based on the decoding features, wherein the target image is the image after the image to be restored is restored.
2. The image restoration method according to claim 1, characterized in that: The repair model is trained in the following way: Acquire sample images corresponding to a plurality of degradation types, and determine original images corresponding to the sample images as annotations corresponding to the sample images; Taking each sample image as input, inputting it into the feature extraction layer of the restoration model to be trained, and obtaining each sample feature output by the feature extraction layer; Inputting each of the sample features into the encoder of the restoration model to obtain the depth coding features respectively corresponding to each of the sample features output by the encoder; Decoding the depth coding features respectively corresponding to the sample features through the decoder of the restoration model, and obtaining the restoration results of the sample images according to the decoded features; The restoration model is trained according to the gap between the restoration results of the sample images and the annotations respectively corresponding to the sample images.
3. The image restoration method according to claim 1, characterized in that: The encoder of the restoration model performs a self-attention encoding operation based on the image features to obtain a deep encoding feature of the image to be restored, including: Performing the self-attention encoding operation on the image feature to obtain a first feature, and downsampling the first feature to obtain a downsampled first feature; Performing the self-attention encoding operation on the downsampled first feature to obtain a second feature; The process of performing a self-attention decoding operation on the deep coding feature through the decoder of the repair model to obtain a decoded feature of the target image includes: Performing the self-attention decoding operation based on the second feature to obtain a third feature, performing upsampling based on the third feature to obtain an upsampled third feature, and performing a fusion operation based on the upsampled third feature and the first feature to obtain a fused third feature; The self-attention decoding operation is performed on the fused third feature to obtain a fourth feature, and the fourth feature is used as a decoding feature of the image.
4. The image restoration method according to claim 1, characterized in that: The encoder and the decoder include a high-dimensional feature mapping structure, wherein the high-dimensional feature mapping structure includes a first convolution branch and a second convolution branch; The encoder of the restoration model performs a self-attention encoding operation based on the image features to obtain a deep encoding feature of the image to be restored, including: Obtaining a first convolution feature of the image feature through a first convolution branch of the high-dimensional feature mapping structure; Obtaining a second convolution feature of the image feature through a second convolution branch of the high-dimensional feature mapping structure, wherein the convolution depths of the first convolution branch and the second convolution branch are different; Performing element-wise multiplication on the first convolution feature and the second convolution feature to obtain a high-dimensional feature; The self-attention encoding operation is performed on the high-dimensional features through the attention network to obtain the deep encoding features of the image to be repaired.
5. The image restoration method according to claim 1, characterized in that: The encoder of the restoration model performs a self-attention encoding operation based on the image features to obtain a deep encoding feature of the image to be restored, including: Determine a first feature to be encoded according to the image feature, and for each element in the first feature to be encoded, determine a first row vector centered on the element and including elements matching the window length according to a preset window length; Determine, according to the elements corresponding to the channels in the first feature to be encoded, the attention weights corresponding to the elements in the first row vector; Determine a first code of the element according to the first row vector and the attention weights corresponding to each element in the first row vector, and obtain an intermediate feature of the first feature to be encoded according to the first codes corresponding to each element in the first feature to be encoded; For each element in the intermediate feature, a sampling interval is determined according to the window length, sampling is performed at equal intervals with the element as the center, and a second row vector with the element as the center and the number of elements matching the window length is determined; Determine, according to the elements corresponding to the channels in the intermediate features, the attention weights corresponding to the elements in the second row vector; Determine a second encoding of the element according to the second row vector and the attention weights corresponding to each element in the second row vector, and obtain a second feature to be encoded according to the second encoding corresponding to each element in the intermediate feature; A vertical attention encoding operation is performed on the second feature to be encoded to obtain a deep encoding feature of the image to be restored.
6. The image restoration method according to claim 5, characterized in that: The step of determining the sampling interval according to the window length includes: According to the window length, a distance from an element at the center of the first row vector to an element at the edge of the first row vector is determined as a sampling interval.
7. The image restoration method according to claim 2, characterized in that: The training of the restoration model according to the gap between the restoration results of each sample image and the annotations respectively corresponding to each sample image comprises: Determine a pixel value difference between the restoration result of each sample image and the annotations respectively corresponding to each sample image as a first difference, and determine a frequency domain difference between the restoration result of each sample image and the annotations respectively corresponding to each sample image as a second difference; The loss is determined according to the first gap and the second gap, and the repair model is trained with minimization of the loss as an optimization goal.
8. An image restoration device, characterized in that: include: A determination module, used for determining an image to be repaired; A feature extraction module, used for inputting the image to be repaired into a feature extraction layer of a pre-trained repair model to obtain image features output by the feature extraction layer; An encoding module, configured to perform a self-attention encoding operation based on the image features through an encoder of the restoration model to obtain a coded feature of the image to be restored, wherein the self-attention encoding operation is: for each element in the first feature to be encoded, a second feature to be encoded is obtained based on other elements belonging to the same row as the element, and for each element in the second feature to be encoded, an encoded deep coding feature is obtained based on other elements belonging to the same column as the element in the second feature to be encoded; A decoding module is used to perform a self-attention decoding operation on the deep coding features through a decoder of the repair model to obtain a decoding feature of a target image, and obtain a target image based on the decoding feature, wherein the target image is an image after the image to be repaired is repaired.
9. A computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.
Citation Information
Patent Citations
Incremental image restoration method based on wireframe and edge structure
CN114399436A
Multi-channel HE staining pathological image segmentation method introducing gated axial self-attention
CN114693675A
Image restoration method, model and device based on convolution and converter hybrid network
CN116309155A
Wavelet-based double-flow network mural restoration method
CN118195965A
CT image noise reduction method and system based on three-dimensional axial attention mechanism
CN118505555A