Training method of degraded region-aware image restoration model and image restoration method

Through the DAPIR model, the feature alignment and fusion of DAPM and DAIM are used to improve the effect and accuracy of image restoration, and the problem of inaccurate image restoration in complex scenarios is solved.

CN120182118BActive Publication Date: 2025-08-12HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510671836.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-12
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The image restoration results in complex scenarios are inaccurate, which affects the accuracy of downstream visual tasks.

Method used

Through the degraded area-aware image restoration model DAPIR, the degraded area features are obtained using DAPM's encoder, and aligned and fused with the intermediate restoration features of IR through DAIM to improve the image restoration effect.

Benefits of technology

Improves the effect and accuracy of the image restoration model in the image degradation situation in multiple complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182118B_ABST
    Figure CN120182118B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for DAPIR, which includes an encoder of a degraded area perception model DAPM and an image restoration model IR, wherein the IR includes a degraded area interaction module DAIM, an encoder, and a decoder. The training method includes: inputting the original image into the encoder and IR of the DAPM respectively; obtaining degradation features of the degraded area of the original image at different scales through the encoder of the DAPM; obtaining intermediate restoration features of the original image through the encoder of the IR and / or the decoder of the IR; aligning the scale of the degradation feature with the intermediate restoration feature through the encoder of the DAIM, converting the degradation feature into a spatial perception feature, fusing the spatial perception feature with the aligned intermediate restoration feature to obtain a regional perception restoration feature, and transmitting the regional perception restoration feature back to the DAPIR for training the DAPIR to improve the quality of image restoration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a training method for a degraded region-aware image restoration model and an image restoration method. Background Art

[0002] In daily life, captured images are often affected by various factors and degraded due to the influence of the shooting environment and the limitations of the camera's hardware configuration. For example, in inclement weather such as rain, snow, fog, and haze, images captured often have a large amount of noise degradation, seriously affecting the subsequent use of the images and adversely affecting downstream vision tasks such as object detection and autonomous driving.

[0003] To address this situation, images can be restored before performing downstream visual tasks to improve accuracy. However, image restoration results in complex scenarios such as severe weather are often inaccurate, and poor image restoration results can also affect the accuracy of downstream tasks. Summary of the Invention

[0004] The purpose of this application is to provide a training method, image restoration method and electronic equipment for the degraded area perception image restoration model DAPIR to solve the problem of inaccurate image restoration results in complex scenes.

[0005] In a first aspect, the present application provides a training method for a degraded region-aware image restoration model DAPIR, wherein the DAPIR includes an encoder of a degraded region-aware model DAPM and an image restoration model IR, wherein the IR includes a degraded region interaction module DAIM, an encoder, and a decoder. The training method includes:

[0006] Input the original image into the encoder of the DAPM and the IR respectively;

[0007] Obtaining degradation features of the degraded region of the original image at different scales through the encoder of the DAPM;

[0008] Obtaining intermediate restored features of the original image through the IR encoder and / or the IR decoder;

[0009] The scale of the degraded feature is aligned with the intermediate restoration feature through the encoder of the DAIM, the degraded feature is converted into a spatial perception feature, the spatial perception feature is fused with the aligned intermediate restoration feature to obtain a regional perception restoration feature, and the regional perception restoration feature is fed back to the DAPIR for training the DAPIR.

[0010] Optionally, the training method further includes:

[0011] Before the original picture is input into the encoder of the DAPM and the IR respectively,

[0012] Training the DAPM using a training set and a test set with degraded region mask information, wherein the DAPM includes an encoder, a segmentation head, and a supervised loss module;

[0013] Optimizing the DAPM by gradient calculation based on the true value of the degraded region mask information;

[0014] After optimizing the DAPM, the encoder of the DAPM is segmented, and the segmentation head and the supervision loss function of the DAPM are discarded;

[0015] The DAPIR is constructed based on the encoder of the DAPM and the IR.

[0016] Optionally, the number of encoders of the IR or the number of decoders of the IR is consistent with the number of encoders of the DAPM, converting the degraded features into spatial perception features, and fusing the spatial perception features with the aligned intermediate restoration features of the IR to obtain region-aware restoration features, including:

[0017] For the degradation features output by each encoder of the DAPM, determine the intermediate restoration features of the coding layer and / or decoding layer in the IR corresponding to the encoder:

[0018] The DAIM is based on the Scale module to align the resolution of each degraded feature with the resolution of the corresponding intermediate restored feature;

[0019] The aligned degraded features are converted into transformed restoration features adapted to the image restoration task of the IR through the Conv convolution layer;

[0020] Performing Softmax activation function calculation and channel separation operation on the features output by the convolution layer of the transformed and restored features to obtain mask information of the key features and mask information of the value features respectively;

[0021] Performing probability distribution calculation on the features output by the convolution layer of the transformed and restored features through an activation function to obtain spatial perception features;

[0022] The weight mask of the key feature and the weight mask of the value feature are superimposed on the basis of the spatial perception feature to obtain the region-aware restoration feature.

[0023] Optionally, fusing the spatial perception feature with the intermediate restoration feature of the IR to obtain a region perception restoration feature includes:

[0024] In the DAIM, the intermediate restored features output by the encoder of the IR and / or the decoder of the IR are input into a first branch of the DAIM, and the intermediate restored features are processed by a convolution layer and a grouped convolution layer to obtain the query features;

[0025] In the DAIM, the intermediate restored features output by the IR are input into the second branch of the DAIM, the intermediate restored features are processed by a convolution layer and a grouped convolution layer to obtain intermediate features, and the intermediate features are subjected to a channel separation operation to obtain key features and value features;

[0026] By element-wise multiplication, the mask information of the key feature and the mask information of the value feature are added to the key feature and the value feature respectively, and are fused through a channel connection operation to obtain a feature with degraded region attention information;

[0027] The features with the degraded regional attention information are converted into K-key features and V-key features in the regional perception restoration features through convolution, grouped convolution and channel separation operations;

[0028] Obtaining the region-aware restoration feature based on the query feature and the K key feature and the V key feature in the region-aware restoration feature;

[0029] Use the standard Self-Attention module to adjust the attention of the intermediate restoration features to obtain the restoration features that initially integrate the degraded area information;

[0030] The restored features and the spatial perception features are subjected to degraded region information fusion using element-wise multiplication to obtain the region perception restored features.

[0031] Optionally, the above method further includes:

[0032] The IR adopts a DACLIP structure, including at least four encoder layers, at least four decoder layers and at least one intermediate layer;

[0033] In the first coding layer of the IR, the original image is input into two layers of base blocks, and the coding output results of the two layers of base blocks are adjusted by the at least one attention layer to obtain intermediate restoration features of the current coding layer; based on the DAIM, the first layer of degraded area features obtained by the DAPM are embedded into the intermediate restoration features and down sampled to obtain output features of the current coding layer;

[0034] In the coding layer after the first coding layer of the IR, the output features of the previous coding layer are used as input, and the output features of the current coding layer are obtained by performing the same operation as that of the previous coding layer;

[0035] In the middle layer of the IR, the output of the last encoding layer is used as the input of the middle layer, and the output of the middle layer is obtained after passing through a base block layer, an attention layer, and a base block layer.

[0036] In the first decoding layer of the IR, the output of the intermediate layer is input into two layers of Base Block and one layer of attention layer for feature extraction to obtain intermediate restoration features. Based on the DAIM, the first layer degradation features obtained by the DAPM are embedded into the intermediate restoration features and upsampled by Up Sample to obtain the output features of the first encoding layer;

[0037] In any decoding layer after the first decoding layer of the IR, the output of the previous decoding layer and the output of the encoding layer of the corresponding layer are fused, input into two layers of base blocks and one layer of attention layer for feature extraction, and intermediate restoration features are obtained. Based on the DAIM, the degradation features of the corresponding scale of the corresponding layer obtained by the DAPM are embedded into the intermediate restoration features and up-sampled to obtain the output features of the current encoding layer;

[0038] The output of the last decoding layer of the IR is used as the restored image of the IR.

[0039] Optionally, training the DAPM using a training set and a test set with degraded region mask information includes:

[0040] The DAPM constructs the true mask value of the degraded area based on the sample image for training. When the degradation task corresponding to the sample image is global degradation, the true mask value of the degraded area of the sample image is all 1s; when the degradation task corresponding to the sample image is local degradation, the true mask value of the degraded area is obtained based on the semi-supervised annotation method of OpenCV;

[0041] The DAPM constructs a training set with degraded region mask information based on the sample image and the true mask value, and performs degraded region perception training on the DAPM using the data set with degraded region mask information.

[0042] Optionally, transmitting the region-aware restoration feature back to the DAPIR for training the DAPIR, further comprising:

[0043] The DAPIR is trained based on the true value of the mask of the degraded area corresponding to the original image and the comparison result between the restored image and the original image.

[0044] In a second aspect, an embodiment of the present application provides an image restoration method, comprising:

[0045] The acquired image to be restored is restored by using the degraded region-aware image restoration model DAPIR, wherein the DAPIR is trained by the above-mentioned training method.

[0046] In a third aspect, an embodiment of the present application provides an electronic device, characterized in that it includes: at least one memory and at least one processor,

[0047] The at least one memory stores executable code, and the at least one processor is configured to execute the executable code in the at least one memory to implement the above method.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and wherein the computer program is executed by a processor to implement the above method.

[0049] The embodiment of the present application obtains degradation features of different scales through the DAPM encoder, converts the degradation features into spatial perception features through DAIM, and fuses the degradation features of different scales with the intermediate restoration features of the corresponding IR level, thereby improving the image restoration model's response to image degradation in various complex scenarios, thereby improving the effect and accuracy of image restoration. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] FIG1( a ) is a schematic diagram of the structure of a training model of an image model provided in an embodiment of the present application;

[0051] Figure 1(b) is a schematic diagram of the training process of the image model provided in an embodiment of the present application;

[0052] Figure 2 A schematic diagram of a degraded region perception module provided in an embodiment of the present application;

[0053] Figure 3 A schematic diagram of degradation types provided in an embodiment of the present application;

[0054] Figure 4 A schematic diagram of the DAIM structure provided in an embodiment of the present application;

[0055] Figure 5 Schematic diagram of the IR structure in the embodiment of this application. DETAILED DESCRIPTION

[0056] The present application will be described in detail below in conjunction with the specific embodiments shown in the accompanying drawings, but these embodiments do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in this field based on these embodiments are included in the scope of protection of the present application.

[0057] An embodiment of the present application provides a degraded area-aware image restoration model DAPIR (i.e., a degraded area-aware image restoration integrated model). As shown in Figure 1(a), DAPIR may include an encoder of a degraded area perception module (DAPM) and an image restoration model (IR). The IR includes a degraded area interaction module (DAIM), an encoder, and a decoder. The training process of the DAPIR can be implemented based on different types of degradation tasks. Therefore, the trained DAPIR can embed the features of the image degraded area into the interactive process of image restoration, guide the image restoration work, and improve the quality of image restoration.

[0058] As shown in Figure 1(b), the training process may include the following steps:

[0059] S101: Construct training sets and test sets corresponding to various degradation task types;

[0060] S102: Preprocess the training set and test set constructed in S101, and apply a semi-supervised method based on OpenCV to obtain the degraded region mask corresponding to the degraded image in the training set;

[0061] S103: Construct the DAPM degradation region perception network, including the DAPM encoder, segmentation head, supervision loss and other structures;

[0062] S104: Input the training set and test set with the degraded region mask information obtained in S102 into the DAPM constructed in S103. Relying on S102 to obtain the true value of the degraded region mask information, the model parameters of the DAPM are optimized through gradient calculation to obtain the ability to perceive the degraded region.

[0063] S105: Build the degraded region-aware image restoration network DAPIR, which includes an encoder, decoder, intermediate layers, and training loss.

[0064] S106, relying on the training set with various degradation task types built in S101, trains the image restoration network based on degradation area perception built in S105, supervises through the true values in the training set, and updates the model parameters to obtain the optimal model parameters.

[0065] In the embodiment of the present application, DAPIR includes the encoder and IR of DAPM, and IR includes DAIM, encoder and decoder. The training method includes:

[0066] Input the original image into the DAPM encoder and IR respectively;

[0067] The degradation features of the degraded areas of the original image at different scales are obtained through the DAPM encoder;

[0068] Obtaining intermediate restored features of the original image through the IR encoder and / or the IR decoder;

[0069] The scale of the degraded feature is aligned with the intermediate restoration feature through the DAIM encoder, the degraded feature is converted into a spatial perception feature, the spatial perception feature is fused with the aligned intermediate restoration feature to obtain the region-aware restoration feature, and the region-aware restoration feature is fed back to DAPIR for training the DAPIR.

[0070] Different types of degradation tasks may include at least 10 types, such as dehazing, de-JPEG (Joint Photographic Experts Group) compression, de-motion blur, dark light enhancement, de-noising, de-raining, de-raindropping, de-snowing, de-shadowing, and image completion. Degradation tasks may be divided into, but are not limited to, two types: global degradation tasks and local degradation tasks. Global degradation tasks may include, but are not limited to, five types of degradation tasks, such as dehazing, de-JPEG compression, de-motion blur, dark light enhancement, and de-noising. Local degradation tasks may include, but are not limited to, five types of degradation tasks, such as de-shadowing, de-raindropping, de-raining, de-snowing, and image completion.

[0071] In global degradation tasks, degradation affects all pixels in the sample image. Therefore, in this scenario, a mask of all 1s is used as the true value of the degraded area of the sample image. In local degradation tasks, the true value of the degraded area of the sample image can be obtained through automated annotation and manual correction.

[0072] For example, in a local degradation task, the true value image and the input sample image can be compared to obtain the relative difference area, the relative difference area can be fine-tuned using OpenCV, and then the true value of the degraded area of the sample image can be obtained by manually using the relative difference area fine-tuned by OpenCV.

[0073] Because the true value of the degraded region in the global degradation task uses an all-ones mask, directly restoring the image based on the degraded region mask may lead to incorrect guidance due to the limited accuracy of the degraded region mask itself, and restore objects that do not normally exist. Degradation categories are high-level semantic information. If different degradation types are directly used for image restoration, the guidance for image restoration may also lead to restoration errors because the image restoration model cannot understand such high-level semantic information.

[0074] Therefore, the embodiments of this application refer to Figure 2 As shown, the original picture Input into DAPM and IR respectively: DAPM extracts the original image through the Encoder Characteristics of the degradation region , the process can be expressed as DAIM can use this feature Guide the image restoration work of IR to output the restored image. Since DAPM can perceive the degraded areas in the degraded image, it can help IR focus more attention on the degraded areas, thereby improving the quality of the restored image. Since the training process adapts to different types of image degradation tasks, the IR trained by this scheme can better restore the image.

[0075] For example, DAPM is used to achieve perception of degraded areas, and the Segformer model can be used as the basic architecture, where Segformer includes two parts: Encoder and SegHead. The Encoder is used to extract features from the image, and the SegHead is used to perform region segmentation on the extracted features.

[0076] For example, the encoder may encode and extract features from the image and perceive the degraded area using, but not limited to, the following formula:

[0077]

[0078] in, Represents the degradation features of different scales perceived by the Encoder on the degraded areas of the original image. The features extracted by Encoder at different scales are relative to The downsampling factor.

[0079] Exemplarily, DAIM is used to embed the degradation features perceived by DAPM into IR, guide the image restoration work of IR, and improve the quality of the image restored by IR.

[0080] Since the degradation features perceived by DAPM are at different scales, DAIM can align the scales of the degradation features perceived by DAPM and the features in IR.

[0081] Exemplarily, DAIM adjusts the resolution of the degradation features obtained by the encoder of the DAPM through the Scale module so that the degradation features of each scale are aligned with the intermediate restoration features output by the coding layer corresponding to the IR in terms of resolution. The resolution of the IR is aligned with the resolution of the intermediate restoration feature corresponding to the IR, and the degraded features are converted to Converted into the restored features of the image restoration task adapted to the IR ;

[0082] The transformed features are restored by the Softmax activation function and Chunck channel separation operation. After the features output by the convolution layer, channel separation calculation is performed to obtain key features respectively. Mask information Sum value characteristics Mask information ;

[0083] The transformed features are restored by the Sigmoid activation function The convolution is used to calculate the probability distribution and obtain the spatial perception features. ;

[0084] Perceive features in the space Based on the superposition of the key features The weight mask and value features of The weight mask of , to obtain region-aware restoration features.

[0085] See also Figure 4 , in a DAPIR training process, The convolution data are input into the two branches of DAIM respectively; the intermediate output results of each layer of IR are Enter the two branches of the DAIM separately.

[0086] DAIM performs scale alignment and feature transformation on the data received by the two branches respectively, thereby obtaining the region-aware restoration features as the output of DAIM. .

[0087] The first branch of DAIM can include a sampling module Scale and a convolution layer Conv. Sampling is performed so that The resolution and Align, and then pass the convolution layer Conv to the sampled Perform feature transformation to obtain Part, so obtained It can be more suitable for image restoration tasks, avoiding the problem of only being suitable for the degraded area detection task corresponding to DAPM but not for the image restoration task.

[0088] Optionally, the operation process of the first branch can refer to the following formula:

[0089]

[0090] Among them, the Scale characterization sampling module is Perform sampling operations, Conv represents the feature transformation process of the convolution layer, Characterize the degraded region characteristics of the original image The features output after the convolution operation.

[0091] Then, for the convolution The corresponding key features can be obtained through activation function and channel separation operation Mask information , value feature Mask information , can be calculated by referring to but not limited to the following methods:

[0092]

[0093] Among them, Softmax represents the Softmax activation function, Chunck represents the channel separation operation, Characterizing key features Mask information , Characterization value features Mask information .

[0094] Obtaining spatial perception features through activation function , can be calculated using, but not limited to, the following methods:

[0095]

[0096] Among them, Sigmoid represents the Sigmoid activation function, Characterize the region-aware restoration features, Characterize the degraded region features of the original image The features output after the convolution operation.

[0097] The operation process of the second branch above can be referred to below:

[0098] Will Input two different convolutional layers Conv and grouped convolutional layer GConv to get query features , key features Sum value characteristics , you can use but not limited to the following formula:

[0099]

[0100]

[0101] Among them, GConv represents the group convolution operation, Characterize query features, Characterize the key features, Characterization value features, where Characterizes the intermediate restoration features of the encoding layer or decoding layer in the IR model.

[0102] See formula:

[0103] Cat represents the channel connection operation in DAIM, Characterize the intermediate restored features of the output after channel connection in DAIM.

[0104] In the embodiment of the present application, element multiplication can also be used to convert the degraded region characteristics into and Add to key features separately Sum value characteristics Then, the features with attention information of the degraded region are obtained. Then, the features with attention mechanism are separated to obtain the adaptive self-attention calculation and .

[0105] in, and Represents the K key features in the region-aware restoration features and V key features .

[0106] Through the self-attention mechanism, the features of the degraded areas perceived by the above two branches can be better integrated into the image restoration process, and its calculation can be based on the value feature And query features and The self-attention feature is calculated from the probability distribution of the transposition of , for example, using the following formula:

[0107]

[0108] in, Characterization based on value features And query features and Calculate the self-attention features, where and Characterizing K-key features in region-aware restoration features and V key features .

[0109] DAIM uses the spatial attention mechanism to calculate the spatial perception features calculated by the first branch above The self-attention features calculated by the second branch are fused to obtain the regional perception restoration features, and the perception features of the degraded area are further fused into the restoration features to obtain the output features that integrate the perception information of the degraded area. ,The region perception restoration features are transmitted back to IR for subsequent restoration work.

[0110] The spatial perception features calculated by the first branch above The process of fusing with the self-attention features calculated by the second branch can be achieved as follows:

[0111]

[0112] in, Characterize query features, Characterize the output of DAIM, Characterization based on value features And query features and Calculate the self-attention features, and Characterizing K-key features in region-aware restoration features and V key features .

[0113] Exemplarily, in an optional embodiment of the present application, IR adopts a DACLIP structure, including at least four encoder layers, at least four decoder layers, at least one attention layer, and at least one intermediate layer;

[0114] In the first coding layer of the IR, the original image is input into two layers of base blocks, and the coding output results of the two layers of base blocks are adjusted through the at least one attention layer to obtain the intermediate restoration features of the current coding layer; based on the DAIM, the first layer degradation features obtained by the DAPM are embedded into the intermediate restoration features and down-sampled to obtain the output features of the current coding layer, wherein the self-attention features of the at least one attention layer are returned to the IR through the DAIM;

[0115] In the coding layer after the first coding layer of the IR, the output features of the previous coding layer are used as input, and the output features of the current coding layer are obtained by performing the same operation as that of the previous coding layer;

[0116] In the middle layer of the IR, the output of the last encoding layer is used as the input of the middle layer, and the output of the middle layer is obtained after passing through a layer of Base Block, a layer of Attention Layer and a layer of Base Block. ;

[0117] In the first decoding layer of the IR, the output of the intermediate layer is converted to Input two layers of Base Block and one layer of Attention Layer for feature extraction to obtain intermediate restoration features. Based on the DAIM, the first layer of degradation features obtained by the DAPM are embedded into the intermediate restoration features and upsampled to obtain the output features of the first layer of encoding layer.

[0118] In any decoding layer after the first decoding layer of the IR, the output of the previous decoding layer and the output of the encoding layer of the corresponding layer are fused, input into two layers of base blocks and one layer of attention layer for feature extraction to obtain intermediate restoration features, and the first layer degradation features obtained by the DAPM are embedded into the intermediate restoration features based on the DAIM and upsampled by Up Sample to obtain the output features of the current encoding layer;

[0119] The output of the last decoding layer of the IR is used as the restored image of the IR.

[0120] Reference Figure 5 In the embodiments of the present application, IR can be, but is not limited to, an integrated image restoration model, such as the DACLP model. DAPM can be used to perceive degraded regions of the original image, while DAIM is used to more effectively embed information about the degraded regions perceived by DAPM into IR, thereby helping IR focus more on the degraded regions and improving the quality of the IR-restored image.

[0121] Assuming that IR adopts the DACLIP model, DACLIP can use Encoder-Decoder as the backbone network. The structure of its Encoder-Decoder is as follows: Figure 5 As shown in the figure, it includes: a 4-layer Encoder Block, a 4-layer Decoder Block and a 1-layer Middle Block. In the Encoder-Decoder structure, the input of each layer of Encoder Block and Decoder Block will be encoded by 2 layers of Base Block, and then the encoded features will be adjusted through the attention mechanism. The encoded features are then input into the DAIM module to extract , and then The input is DACLIP, and the scale is transformed through the sampling layer in DACLIP to improve the receptive field of DACLIP.

[0122] An exemplary Encoder Block may include the following modules:

[0123] In the first layer of the Encoder module, the original image is used as input. After the original image is encoded by the two layers of Base Block, the attention mechanism is used to make further adjustments to obtain the output. , then refer to Figure 4 As shown in Figure 2, DAIM is used to embed the degradation region features perceived by DAPM into In this process, the output of DAIM passes through a down sample layer to obtain the output of the first Encoder .

[0124] In the second layer Encoder module, As input, like the first-layer Encoder module, feature encoding is achieved through two layers of BaseBlock, and then the intermediate restoration features are obtained through the attention mechanism adjustment , then refer to Figure 4 As shown in Figure 2, DAIM is used to embed the degradation region features perceived by DAPM into In this process, the output of DAIM passes through a layer of DownSample downsampling layer to obtain the output of the second Encoder .

[0125] Similar to the second layer, the third and fourth layers take the output of the previous layer as input and generate corresponding outputs respectively. and .

[0126] The Middle Block can contain 2 layers of Base Block and one layer of attention layer:

[0127] The Middle Block takes the output of the last encoding layer as input, passes through a Base Block layer, an Attention Layer layer, and another Base Block layer, and obtains the output of the Middle Block.

[0128] Optionally, the Middle Block can be composed of 2 layers of Base Block and one layer of attention mechanism module. As input, it passes through a layer of Base Block for feature extraction, and then connects a layer of attention mechanism module for attention adjustment, and then passes through the Base Block with the same configuration to obtain feature output .

[0129] In the first decoding layer of the IR, the output of the Middle Block can be Input two layers of Base Block and one layer of Attention Layer for feature extraction to obtain intermediate restoration features. Based on the DAIM, the first layer of degradation features obtained by the DAPM are embedded into the intermediate restoration features and upsampled to obtain the output features of the first layer of encoding layer.

[0130] In any decoding layer after the first decoding layer of the IR, the output of the previous decoding layer and the output of the encoding layer of the corresponding layer are fused, input into two layers of base blocks and one layer of attention layer for feature extraction to obtain intermediate restoration features, and the degradation features of the corresponding scale obtained by the DAPM are embedded into the intermediate restoration features based on the DAIM and up-sampled to obtain the output features of the current encoding layer;

[0131] The output of the last decoding layer of the IR is used as the restored image of the IR.

[0132] Reference Figure 5 As shown in FIG, for the third layer decoder, the fourth layer decoder has the same structure and calculation process as the second layer decoder, which respectively generates corresponding features and the final restored image.

[0133] In the embodiment of the present application, DAPM constructs the true mask value of the degraded area based on the original image, wherein, when the degradation task of the IR is global degradation, the true mask value of the degraded area is all 1; when the degradation task of the IR is local degradation, the DAPM obtains the true mask value of the degraded area through a semi-supervised annotation method based on OpenCV;

[0134] Therefore, DAPM can construct a dataset based on the original image and the mask truth value, and perform degraded region-aware training through the dataset.

[0135] The IR model can use the above method to obtain the true value of the mask of the degraded area corresponding to the original image and the comparison result between the restored image and the original image to train the IR.

[0136] Since the degraded area mask cannot be directly obtained automatically, and the manually input mask may have the problem of incorrect degraded area due to various reasons (such as labeling errors, inaccurate predictions, etc.), the embodiment of the present application embeds the degraded area features as implicit features into the intermediate restoration features of the image restoration model, which has stronger anti-interference ability. Moreover, this solution uses the degraded area features as input, which will make the image restoration results more robust.

[0137] Based on the same inventive concept, an embodiment of the present application also provides an electronic device, comprising: at least one memory and at least one processor, the at least one memory storing executable code, the at least one processor being used to execute the executable code in the at least one memory to implement the above-mentioned image restoration model training method, and / or image restoration method.

[0138] The memory may be a random access memory, a read-only memory, a non-volatile memory, a programmable ROM, an erasable PROM, an electrically erasable memory, a flash memory, an optical memory, a register, and the like. The processor may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing a computer program stored in the memory, and the general-purpose processor may use data stored in the memory in the process of performing the steps and / or operations. The general-purpose processor may be a central processing unit, an ASIC, an FPGA, and the like. During implementation, each step of the above method may be completed by an integrated logic circuit of hardware in the processor or by instructions in the form of software. The method disclosed in conjunction with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0139] Exemplarily, the input device includes but is not limited to at least one of a keyboard, a touch panel, a voice input device, and an image sensor.

[0140] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or solid-state drive (SSD).

[0141] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0142] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The above is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application are included in the scope of protection of this application.

Claims

1. A training method for a degraded region-aware image restoration model DAPIR, characterized in that: The DAPIR includes an encoder of a degraded area perception model DAPM and an image restoration model IR, wherein the IR includes a degraded area interaction module DAIM, an encoder and a decoder. The training method includes: Training the DAPM using a training set and a test set with degraded region mask information, wherein the DAPM includes an encoder, a segmentation head, and a supervised loss module; Optimizing the DAPM by gradient calculation based on the true value of the degraded region mask information; After optimizing the DAPM, the encoder of the DAPM is segmented, and the segmentation head and the supervision loss function of the DAPM are discarded; Constructing the DAPIR based on the encoder of the DAPM and the IR; Inputting the original image into the encoder of the DAPM and the IR respectively; Obtaining degradation features of the degraded region of the original image at different scales through the encoder of the DAPM; Obtaining intermediate restored features of the original image through the IR encoder and / or the IR decoder; The scale of the degraded feature is aligned with the intermediate restoration feature through the encoder of the DAIM, the degraded feature is converted into a spatial perception feature, the spatial perception feature is fused with the aligned intermediate restoration feature to obtain a regional perception restoration feature, and the regional perception restoration feature is fed back to the DAPIR for training the DAPIR.

2. The method according to claim 1, wherein The number of coding layers of the IR encoder is consistent with the number of coding layers of the DAPM encoder, and / or, The number of decoding layers of the IR decoder is consistent with the number of encoding layers of the DAPM encoder; Converting the degraded features into spatial perception features, fusing the spatial perception features with the aligned intermediate restoration features of the IR to obtain region-aware restoration features, including: For the degradation features output by each coding layer of the DAPM, determine the intermediate restoration features of the coding layer and / or decoding layer in the IR corresponding to the coding layer: The DAIM is based on the Scale module to align the resolution of each degraded feature with the resolution of the corresponding intermediate restored feature; The aligned degraded features are converted into transformed restoration features suitable for the image restoration task of the IR through the Conv convolution layer; Performing Softmax activation function calculation and channel separation operation on the features output by the convolution layer of the transformed and restored features to obtain mask information of the key features and mask information of the value features respectively; Performing probability distribution calculation on the features output by the convolution layer of the transformed and restored features through a Sigmoid activation function to obtain spatial perception features; On the basis of the spatial perception feature, the weight mask of the key feature and the weight mask of the value feature are superimposed to obtain the region-aware restoration feature.

3. The method according to claim 1, wherein The spatial perception feature is fused with the intermediate restoration feature of the IR to obtain the regional perception restoration feature, including: In the DAIM, the intermediate restored features output by the IR encoder and / or the IR decoder are input into a first branch of the DAIM, and the intermediate restored features are processed by a convolution layer and a grouped convolution layer to obtain a query feature; In the DAIM, the intermediate restored features output by the IR are input into the second branch of the DAIM, the intermediate restored features are processed by a convolution layer and a grouped convolution layer to obtain intermediate features, and the intermediate features are subjected to a channel separation operation to obtain key features and value features; By element-wise multiplication, the mask information of the key feature and the mask information of the value feature are added to the key feature and the value feature respectively, and fused through a channel connection operation to obtain a feature with degraded region attention information; The features with the degraded regional attention information are converted into K-key features and V-key features in the regional perception restoration features through convolution, grouped convolution and channel separation operations; Obtaining the region-aware restoration feature based on the query feature and the K key feature and the V key feature in the region-aware restoration feature; Use the standard Self-Attention module to adjust the attention of the intermediate restoration features to obtain the restoration features that initially integrate the degraded area information; The restored features and the spatial perception features are subjected to degraded region information fusion using element-wise multiplication to obtain the region perception restored features.

4. The method according to claim 1, wherein The method further comprises: The IR adopts a DACLIP structure, including at least four encoding layers, at least four decoding layers, at least one attention layer and at least one intermediate layer; In the first coding layer of the IR, the original image is input into two layers of base blocks, and the coding output results of the two layers of base blocks are adjusted by the at least one attention layer to obtain intermediate restoration features of the current coding layer; based on the DAIM, the first layer of degraded area features obtained by the DAPM are embedded into the intermediate restoration features and down sampled to obtain output features of the current coding layer; In the coding layer after the first coding layer of the IR, the output features of the previous coding layer are used as input, and the output features of the current coding layer are obtained by performing the same operation as that of the previous coding layer; In the middle layer of the IR, the output of the last encoding layer is used as the input of the middle layer, and the output of the middle layer is obtained after passing through a base block layer, an attention layer, and a base block layer. In the first decoding layer of the IR, the output of the intermediate layer is input into two layers of Base Block and one layer of Attention Layer for feature extraction to obtain intermediate restoration features. Based on the DAIM, the first layer degradation features obtained by the DAPM are embedded into the intermediate restoration features and upsampled by Up Sample to obtain the output features of the first encoding layer; In any decoding layer after the first decoding layer of the IR, the output of the previous decoding layer and the output of the encoding layer of the corresponding layer are fused, input into two layers of base blocks and one layer of attention layer for feature extraction, and intermediate restoration features are obtained. Based on the DAIM, the degradation features of the corresponding scale of the corresponding layer obtained by the DAPM are embedded into the intermediate restoration features and up-sampled to obtain the output features of the current encoding layer; The output of the last decoding layer of the IR is used as the restored image of the IR.

5. The method according to claim 1, wherein The DAPM is trained using a training set and a test set with degraded region mask information, including: The DAPM constructs the true mask value of the degraded area based on the sample image for training. When the degradation task corresponding to the sample image is global degradation, the true mask value of the degraded area of the sample image is all 1s; when the degradation task corresponding to the sample image is local degradation, the true mask value of the degraded area is obtained based on the semi-supervised annotation method of OpenCV; The DAPM constructs a training set with degraded region mask information based on the sample image and the true mask value, and performs degraded region perception training on the DAPM using the data set with degraded region mask information.

6. The method according to claim 4, wherein Transmitting the region-aware restoration features back to the DAPIR for training the DAPIR, further comprising: The DAPIR is trained based on the true value of the mask of the degraded area corresponding to the original image and the comparison result between the restored image and the original image.

7. An image restoration method, characterized in that: include: The acquired image to be restored is restored by using a degraded region-aware image restoration model DAPIR, wherein the DAPIR is trained by the method according to any one of claims 1 to 6.

8. An electronic device, characterized in that: include: at least one memory and at least one processor, The at least one memory stores executable code, and the at least one processor is configured to execute the executable code in the at least one memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Model training method, quality evaluation method, device, equipment and product

    CN118736347A

  • Two-stage rain and fog image restoration method based on conditional diffusion model

    CN118864319A