Training method for blind image super-resolution model, blind image super-resolution method, electronic device, storage medium and program product
The blind image super-resolution model is trained by comparative learning and distillation learning methods based on degradation prior constraints, and the problem of high model complexity is solved and efficient image reconstruction on low-computing equipment is realized.
Patent Information
- Application Number
- CN202510183438.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The existing blind image super-resolution method is limited in applications on low-computing equipment, with high model complexity and large calculation amounts, making it difficult to effectively deal with complex degradation in real scenarios.
The degradation feature estimator is trained using a contrast learning method based on degradation prior constraints, combined with the distillation learning method, reduce the amount of model calculation, and reconstruct the high-definition image through the degradation feature estimator and image feature correction module.
Excellent image reconstruction quality is achieved on low computing power equipment, improving the discriminantity of degraded features and overall training efficiency, and is suitable for end-side equipment.
Smart Images

Figure CN119722460B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of image processing technologies, and more specifically, to a method for training a blind image super-resolution model, a blind image super-resolution method, an electronic device, a storage medium, and a program product. Background Art
[0002] Blind Super-Resolution (BSR) is an important image processing task, which mainly refers to the process of restoring a Low-resolution image (LR) to a High-resolution image (HR) without knowing or being uncertain about the specific process of image degradation (such as blur kernel, noise level, downsampling operation, etc.). It has been widely applied in fields such as remote sensing, video surveillance, and medical imaging. At the same time, blind image super-resolution can also be used as a pre-operation for many advanced vision tasks, such as image retrieval, object detection, and scene analysis. The challenge faced by the BSR task is the unknown nature of the degradation process, because the image degradation in the real world is often complex and diverse. Therefore, how to accurately estimate the degradation information and how to use the degradation information to guide the reconstruction process are the key issues of concern in the BSR field.
[0003] Early BSR methods adopted an explicit degradation parameter estimation strategy, which accurately estimated the degradation information of the LR image through a complex explicit degradation estimation network and used it as the guiding information for the BSR network. However, this explicit degradation parameter estimation strategy still cannot handle the complex degradations in real scenarios well, including noise interference and multiple degradation interactions. Therefore, the BSR field has gradually started to study implicit degradation parameter estimation strategies and how to use implicit degradation features to guide the reconstruction of LR images to HR images.
[0004] However, for existing BSR methods based on implicit degradation feature guidance, their model complexity is getting higher and higher, often exceeding 10M, and the computational cost is also relatively large, often exceeding 600 GFLOPs (Giga Floating-Point Operations per Second, indicating one billion floating-point operations per second), which greatly limits the application of these methods on low-computing-power devices. Summary of the Invention
[0005] The present disclosure provides a method for training a blind image super-resolution model, a blind image super-resolution method, an electronic device, a storage medium, and a program product, which are used to solve at least one of the above problems.
[0006] According to the first aspect of the embodiments of the present disclosure, a method for training a blind image super-resolution model is provided. The blind image super-resolution model includes a degradation feature estimator for determining the degradation features of a blind image, an image feature correction module for correcting the image features of the blind image according to the degradation features, and an image reconstruction module for reconstructing a high-definition image based on the corrected image features. Wherein, the training method includes: determining training samples in the first stage based on the first sample blind image, and performing first-stage training on the to-be-trained degradation feature estimator by using a contrastive learning method based on degradation prior constraint to obtain a first degradation feature estimator; determining training samples in the second stage based on the second sample blind image, and performing second-stage training on the blind image super-resolution model including the first degradation feature estimator to obtain a teacher model; determining training samples in the third stage based on the third sample blind image, taking the teacher model as the learning object, and performing third-stage training by using a distillation learning method to obtain the final blind image super-resolution model.
[0007] Optionally, the determining training samples in the first stage based on the first sample blind image, and performing first-stage training on the to-be-trained degradation feature estimator by using a contrastive learning method based on degradation prior constraint to obtain a first degradation feature estimator includes: constructing a backbone branch and a momentum branch, wherein both the backbone branch and the momentum branch include a degradation feature estimator; performing degradation prior constraint processing on the first sample blind image to obtain a training blind image; using the training blind image as a training sample, and performing training on the backbone branch and the momentum branch by using a contrastive learning method, and taking the degradation feature estimator in the trained backbone branch as the first degradation feature estimator.
[0008] Optionally, the performing degradation prior constraint processing on the first sample blind image to obtain a training blind image includes: performing dimensional extension processing on the blur kernel and noise information in the first sample blind image to obtain a degradation prior having the same spatial dimension as the first sample blind image; splicing the degradation prior and the first sample blind image along the channel dimension to obtain the training blind image.
[0009] Optionally, taking the training blind image as a training sample and using the contrastive learning method to train the backbone branch and the momentum branch, and using the degradation feature estimator in the finally obtained backbone branch as the first degradation feature estimator includes: performing the following iterative calculation steps on the backbone branch and the momentum branch until the stop condition is met, and using the degradation feature estimator in the latest backbone branch as the first degradation feature estimator. Here, the number of training blind images is multiple, and the training blind images used in each round of calculation are different: inputting the training blind images for the current round of calculation into the backbone branch and the momentum branch respectively to obtain backbone features and momentum features; obtaining reference features, where the reference features are features obtained by inputting an image different from the training blind image for the current round of calculation into the backbone branch and / or the momentum branch; determining a contrast loss value according to the backbone features, the momentum features, and the reference features; adjusting the parameters of the backbone branch according to the contrast loss value, and using the momentum update strategy of contrastive learning to adjust the parameters of the momentum branch based on the parameter adjustment of the backbone branch.
[0010] Optionally, the image feature correction module includes a degradation feature converter and a degradation feature adaptation module. The degradation feature converter includes a channel domain conversion branch and a spatial domain conversion branch. The channel domain conversion branch is used to convert the degradation features into channel domain degradation features, and the spatial domain conversion branch is used to convert the degradation features into spatial domain degradation features. The degradation feature adaptation module includes multiple degradation feature adaptation groups. Each degradation feature adaptation group includes multiple degradation feature adaptation units. Each degradation feature adaptation unit includes a feature splitter, a channel domain adaptation branch, a spatial domain adaptation branch, and a feature fuser. The feature splitter is used to split the image features input thereto into first image features and second image features. The channel domain adaptation branch is used to incorporate the channel domain degradation features into the first image features to obtain first corrected image features. The spatial domain adaptation branch is used to incorporate the spatial domain degradation features into the second image features to obtain second corrected image features. The feature fuser is used to fuse the first corrected image features and the second corrected image features into new image features.
[0011] Optionally, determining the training samples for the third stage based on the third sample blind image, taking the teacher model as the learning object, and using the distillation learning method to perform the third stage of training to obtain the final blind image super-resolution model includes: copying the teacher model to obtain a student model; taking the third sample blind image as a training sample and using the distillation learning method to perform model learning with the aim of aligning the outputs of the degradation feature converters in the teacher model and the student model, and using the learned student model as the final blind image super-resolution model.
[0012] According to a second aspect of the embodiments of the present disclosure, a blind image super-resolution method is provided, including: using a degradation feature estimator in a blind image super-resolution model to process a blind image to obtain the degradation features of the blind image, where the blind image super-resolution model further includes an image feature correction module and an image reconstruction module; using the image feature correction module to correct the image features of the blind image according to the degradation features to obtain corrected image features; using the image reconstruction module to reconstruct a high-definition image corresponding to the blind image based on the corrected image features, where the blind image super-resolution model is trained by a training method of a blind image super-resolution model according to an exemplary embodiment of the present disclosure.
[0013] According to a third aspect of the embodiments of the present disclosure, a training device for a blind image super-resolution model is provided. The blind image super-resolution model includes a degradation feature estimator for determining the degradation features of a blind image, an image feature correction module for correcting the image features of the blind image according to the degradation features, and an image reconstruction module for reconstructing a high-definition image based on the corrected image features. The training device includes: a first training unit configured to determine training samples in a first stage based on a first sample blind image, and perform first-stage training on a to-be-trained degradation feature estimator by using a contrast learning method based on a degradation prior constraint to obtain a first degradation feature estimator; a second training unit configured to determine training samples in a second stage based on a second sample blind image, and perform second-stage training on the blind image super-resolution model including the first degradation feature estimator to obtain a teacher model; a third training unit configured to determine training samples in a third stage based on a third sample blind image, use the teacher model as a learning object, and perform third-stage training by using a distillation learning method to obtain a final blind image super-resolution model.
[0014] According to a fourth aspect of the embodiments of the present disclosure, a blind image super-resolution device is provided, including: a degradation estimation unit configured to use a degradation feature estimator in a blind image super-resolution model to process a blind image to obtain the degradation features of the blind image, where the blind image super-resolution model further includes an image feature correction module and an image reconstruction module; a correction unit configured to use the image feature correction module to correct the image features of the blind image according to the degradation features to obtain corrected image features; a reconstruction unit configured to use the image reconstruction module to reconstruct a high-definition image corresponding to the blind image based on the corrected image features, where the blind image super-resolution model is trained by a training method of a blind image super-resolution model according to an exemplary embodiment of the present disclosure.
[0015] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, including: at least one processor; at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute a method for training a blind image super-resolution model or a blind image super-resolution method according to an exemplary embodiment of the present disclosure.
[0016] According to a sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute a method for training a blind image super-resolution model or a blind image super-resolution method according to an exemplary embodiment of the present disclosure.
[0017] According to a seventh aspect of the embodiments of the present disclosure, a computer program product is provided, including computer instructions, which when run by at least one processor, cause the at least one processor to execute a method for training a blind image super-resolution model or a blind image super-resolution method according to an exemplary embodiment of the present disclosure.
[0018] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: According to the method for training a blind image super-resolution model, the blind image super-resolution method, the electronic device, the storage medium, and the program product of the present disclosure, a new implicit degradation feature learning scheme is proposed. By adopting a degradation prior constraint on the basis of the contrast learning method, the degradation prior obtained by prior analysis can be utilized to make the degradation feature estimator pay more attention to degradation-related information during the contrast learning process, which can effectively improve the discriminability of the learned implicit degradation features. There is no need to design a very complex blind image super-resolution model to improve the super-resolution performance. Combining with the distillation learning method can further reduce the computational amount of the model. An excellent image reconstruction quality can be achieved using a low-complexity blind image super-resolution model, which is suitable for applications on low-computing-power devices and edge devices. In addition, the degradation feature estimator is used to determine the degradation features. By applying the contrast learning method based on the degradation prior constraint in the first stage to the degradation feature estimator alone, rather than directly applying it to the entire model, the ability to extract degradation features can be specifically improved first, and then the degradation feature estimator trained in the first stage is embedded into the entire model for training, which helps to improve the overall training efficiency.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation to the present disclosure.
[0021] Figure 1 It is a flowchart of a method for training a blind image super-resolution model according to an exemplary embodiment of the present disclosure.
[0022] Figure 2 It is a schematic structural diagram of a degradation feature converter according to a specific embodiment of the present disclosure.
[0023] Figure 3 It is a schematic structural diagram of a blind image super-resolution model according to a specific embodiment of the present disclosure.
[0024] Figure 4 It is a schematic diagram of the structure and training logic of a teacher network according to a specific embodiment of the present disclosure.
[0025] Figure 5 It is a schematic diagram of the training logic of the first and second stages according to a specific embodiment of the present disclosure.
[0026] Figure 6 It is a schematic structural diagram of a degradation feature adaptation unit according to a specific embodiment of the present disclosure.
[0027] Figure 7 It is a schematic diagram of the structure and training logic of a student model according to a specific embodiment of the present disclosure.
[0028] Figure 8 It is a schematic diagram of the training logic of the third stage according to a specific embodiment of the present disclosure.
[0029] Figure 9 It is a flowchart of a blind image super-resolution method according to an exemplary embodiment of the present disclosure.
[0030] Figure 10 It is a block diagram of a training device for a blind image super-resolution model according to an exemplary embodiment of the present disclosure.
[0031] Figure 11 It is a block diagram of a blind image super-resolution device according to an exemplary embodiment of the present disclosure.
[0032] Figure 12 It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Embodiments
[0033] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0034] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0035] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations, including "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "executing at least one of step one and step two" means the following three parallel situations: (1) executing step one; (2) executing step two; (3) executing step one and step two.
[0036] Next, a method for training a blind image super-resolution model, a blind image super-resolution method, an electronic device, a storage medium, and a program product according to an exemplary embodiment of the present disclosure will be described in detail with reference to the drawings.
[0037] Figure 1 is a flowchart of a method for training a blind image super-resolution model according to an exemplary embodiment of the present disclosure. This method can be executed on an electronic device with sufficient computing power.
[0038] The blind image super-resolution model includes a degradation feature estimator, an image feature correction module, and an image reconstruction module. The degradation feature estimator is used to determine the degradation features of the blind image, the image feature correction module is used to correct the image features of the blind image according to the degradation features, and the image reconstruction module is used to reconstruct a high-definition image based on the corrected image features. It should be understood that a blind image is a low-resolution image whose specific degradation process is unknown or uncertain, and a high-definition image is a high-resolution image restored from the low-resolution image. Here, low resolution and high resolution are relative concepts, and super-resolution means that the resolution of the original image (i.e., the low-resolution image) is improved after restoration (i.e., a high-resolution image is obtained). It should be understood that the model may also include other reasonable structures, such as, but not limited to, an image feature extractor for extracting the image features of the blind image.
[0039] Refer to Figure 1, in step S101, based on the first sample blind image, the training samples in the first stage are determined, and a contrast learning method based on degradation prior constraint is adopted to perform the first-stage training on the degradation feature estimator to be trained, obtaining the first degradation feature estimator.
[0040] Specifically, the first sample blind image can be the directly obtained original blind image, or several patches sampled from the directly obtained original blind image according to a specified size. The present disclosure places no restrictions on this. The same applies to the second sample blind image and the third sample blind image in the following text, and will not be repeated one by one. Moreover, the first sample blind image, the second sample blind image, and the third sample blind image can be exactly the same, partially the same, or completely different. The present disclosure places no restrictions on this.
[0041] Optionally, step S101 includes: constructing a backbone branch and a momentum branch, where both the backbone branch and the momentum branch include a degradation feature estimator; performing degradation prior constraint processing on the first sample blind image to obtain a training blind image; using the training blind image as a training sample, and adopting a contrast learning method to train the backbone branch and the momentum branch, and using the degradation feature estimator in the trained backbone branch as the first degradation feature estimator. By constructing a backbone branch and a momentum branch that both include a degradation feature estimator, contrast learning can be achieved by comparing the output results of the two branches, and then using the degradation feature estimator trained in the backbone branch as the training result of the first stage. By using the result of performing degradation prior constraint processing on the first sample blind image as the training sample in this stage, a degradation prior can be made to exist in the training sample, which can guide the degradation feature estimator to pay more attention to degradation-related information during training, thereby effectively improving the discriminability of the learned degradation features. It should be understood that since the degradation prior is incorporated into the training sample in advance and can be used as a learning reference for the degradation feature estimator, the degradation prior can also be called a degradation reference prior. It should be noted that the structures of the degradation feature estimators in the backbone branch and the momentum branch are the same as those in the blind image super-resolution model, but the ways of adjusting the parameters of the backbone branch and the momentum branch are different, resulting in different final trained degradation feature estimators in the two branches. The degradation feature estimator in the backbone branch is directly applied to the blind image super-resolution model, while the degradation feature estimator in the momentum branch is only used to assist in training the degradation feature estimator in the backbone branch. As an example, the backbone branch and the momentum branch further include a projector connected after the degradation feature estimator for performing non-linear transformation of features, which can improve the representation ability of the features.
[0042] Further optionally, the operation of performing degradation prior constraint processing on the first sample blind image in step S101 to obtain a training blind image includes: performing dimensional extension processing on the blur kernel and noise information in the first sample blind image to obtain a degradation prior with the same spatial dimension as the first sample blind image; splicing the degradation prior and the first sample blind image along the channel dimension to obtain a training blind image. The blur kernel is a mathematical model used to describe the process of image blurring and can be regarded as a convolution kernel that convolves a high-resolution image to generate a low-resolution image. In the blind image super-resolution task, the blur kernel is usually unknown and needs to be estimated from the low-resolution image. Specific estimation methods can refer to relevant technologies in this field, such as the literature "Li Gongping, Lu Yao, Wang Zijian, Wu Ziwei, Wang Shunzhou. Blind Image Super-Resolution Neural Network Based on Blur Kernel Estimation. Acta Automatica Sinica, 2023, 49(10): 2109-2121 doi: 10.16383 / j.aas.c200987", which will not be elaborated here. The noise information includes, for example, the noise level, which refers to the intensity or amplitude of the noise in the image and can be estimated by various methods, such as by analyzing the statistical characteristics of the image or using a deep learning model. This is also an existing technology in this field and will not be elaborated. By using existing technologies to extract the blur kernel and noise information in the first sample blind image, information related to degradation can be obtained in advance, and then these information can be extended into a degradation prior through a dimensional extension strategy, making the spatial dimension of the degradation prior consistent with that of the first sample blind image. Then, the degradation prior and the first sample blind image can be spliced along the channel dimension to obtain a blind image incorporating the degradation prior as the training blind image, realizing the constraint based on the degradation prior.
[0043] For the embodiment where the first sample blind image is a sampled patch, as an example, when performing degradation prior constraint processing, specifically, a random sample of patches of size is taken from each original blind image to form a blind patch sequence as the first sample blind image sequence. Each pixel point of each blind patch is represented by three components (red, green, blue) in the RGB color space, so each blind patch can be represented as a third-order matrix . Then, using the dimensional extension strategy, the corresponding degradation prior is generated according to the degradation information (including the blur kernel and noise level) of each blind patch. Specifically, for each blind patch, the first step is to vectorize the -sized blur kernel to size and project it to size using principal component analysis technology; the second step is to -sized noise level Replicate three times and concatenate to the aforementioned blur kernel vector to form The third step is to copy the above vector times, that is, forming a degenerate prior Finally, the obtained degradation prior is integrated into the corresponding blind image block, specifically by splicing the two along the channel dimension to form a training blind image sequence , each training blind image can be represented as a third-order matrix .
[0044] Further optionally, in step S101, the training blind image is used as a training sample, and the contrast learning method is adopted to train the trunk branch and the momentum branch, and the operation of using the degradation feature estimator in the trunk branch obtained by training as the first degradation feature estimator includes: performing the following iterative calculation steps on the trunk branch and the momentum branch until the stop condition is met, and using the degradation feature estimator in the latest trunk branch as the first degradation feature estimator, wherein the number of training blind images is multiple, and the training blind images used in each round of calculation are different: inputting the training blind images calculated in this round into the trunk branch and the momentum branch respectively to obtain trunk features and momentum features; obtaining reference features, wherein the reference features are features obtained by inputting images different from the training blind images calculated in this round into the trunk branch and / or the momentum branch; determining the contrast loss value according to the trunk features, the momentum features and the reference features; adjusting the parameters of the trunk branch according to the contrast loss value, and adjusting the parameters of the momentum branch based on the parameters of the trunk branch using the momentum update strategy of contrast learning. The principle of contrastive learning is to shorten the distance between the features of the same or similar images and to increase the distance between the features of different images. Since both the trunk branch and the momentum branch contain degraded feature estimators, and the same training blind images are input to the two in each round of calculation, it is necessary to shorten the outputs of the two, that is, shorten the trunk features and the momentum features. By introducing the above-mentioned reference features, a reference for shortening can be provided, that is, shortening the distance between the trunk features and the reference features, so that contrastive learning can be achieved based on the contrast loss value composed of the three features. In addition, it has been found in practice that the training efficiency can be improved by adjusting the parameters of the trunk branch only based on the contrast loss value and adjusting the parameters of the momentum branch using the momentum update strategy.
[0045] Still taking the embodiment in which the first sample blind image is a sampled image block as an example, in each round of calculation, a training blind image sequence corresponding to an original blind image may be As the training blind image for this round of calculation, and taking the set of backbone features and / or momentum features output in each previous round of calculation as reference features, the number of reference features will gradually increase as the iterative calculation progresses. It should be noted that the "backbone features and / or momentum features" here means that only the backbone branch can be used, only the momentum branch can be used, or both the backbone branch and the momentum branch can be used at the same time. The present disclosure does not limit this. In addition, for the first round of calculation, since there has been no previous calculation, the reference features can be randomly determined, or an additional original blind image can be selected, and the corresponding training blind image sequence can be determined and input into the untrained backbone branch and / or momentum branch to obtain the reference features. Of course, the reference features can also be determined by other reasonable methods, and the present disclosure does not limit this either.
[0046] Continue to refer to Figure 1 In step S102, based on the second sample blind image, the training samples for the second stage are determined, and the blind image super-resolution model including the first degradation feature estimator is trained in the second stage to obtain the teacher model.
[0047] Any reasonable training method can be used for the training in the second stage. For example, in the supervised training method, the second sample blind image is also accompanied by the corresponding real high-definition image, which is used to calculate the loss function value to realize the training of the overall model. The present disclosure does not limit this.
[0048] It should be noted that when training the entire model at this stage, the specific degradation feature estimator being trained is the first degradation feature estimator, that is, the degradation feature estimator of the backbone branch. For the degradation feature estimator of the momentum branch, since it is no longer used, there is no need to train it. Of course, it can also be trained, and the present disclosure does not limit this. For the embodiment of jointly training the degradation feature estimator in the momentum branch, the parameter adjustment method can follow the method adopted in the first stage.
[0049] In step S103, based on the third sample blind image, the training samples for the third stage are determined, and the teacher model is used as the learning object, and the third stage of training is carried out using the distillation learning method to obtain the final blind image super-resolution model.
[0050] By adopting the distillation learning method in the third stage, the degradation-related knowledge learned by the teacher model can be transferred to the student model with a simple structure, and the student model is used as the final blind image super-resolution model, which helps to reduce the computational amount of the model and realize the lightweight model deployment in the inference stage.
[0051] Before specifically introducing step S103, the image feature correction module will be further introduced.
[0052] Optionally, the image feature correction module includes a degradation feature converter and a degradation feature adaptation module. Among them, the degradation feature converter includes a channel domain conversion branch and a spatial domain conversion branch. The channel domain conversion branch is used to convert the degradation features into channel domain degradation features, and the spatial domain conversion branch is used to convert the degradation features into spatial domain degradation features. The degradation feature adaptation module includes multiple degradation feature adaptation groups. Each degradation feature adaptation group includes multiple degradation feature adaptation units. Each degradation feature adaptation unit includes a feature splitter, a channel domain adaptation branch, a spatial domain adaptation branch, and a feature fuser. The feature splitter is used to split the input image features into first image features and second image features. The channel domain adaptation branch is used to incorporate the channel domain degradation features into the first image features to obtain first corrected image features. The spatial domain adaptation branch is used to incorporate the spatial domain degradation features into the second image features to obtain second corrected image features. The feature fuser is used to fuse the first corrected image features and the second corrected image features into new image features. By using the degradation feature converter to convert the degradation features into channel domain degradation features and spatial domain degradation features, and using the degradation feature adaptation module to perform feature adaptation and fusion repeatedly for multiple times, that is, first splitting the image features, then performing feature fusion for the channel domain and spatial domain respectively, and then merging and fusing the two corrected image features after fusion, the characteristics of both the channel domain and the spatial domain can be considered simultaneously, realizing more efficient fusion of the degradation features and the image features, thereby achieving efficient correction of the image features. It should be understood that the structures of the channel domain conversion branch and the spatial domain conversion branch in the degradation feature converter can be reasonably designed as required. As an example, as Figure 2 shown, the channel domain conversion branch may include a global average pooling layer and a fully connected layer, and the spatial domain conversion branch may include a pixel reorganization layer and a convolutional layer. The present disclosure does not limit this. In addition to the multiple degradation feature adaptation groups in the degradation feature adaptation module, other structural layers may also be included. In addition to the multiple degradation feature adaptation units in each degradation feature adaptation group, other structural layers may also be included. As an example, as Figure 3 shown, the degradation feature adaptation module further includes a convolutional layer connected after the multiple degradation feature adaptation groups. Each degradation feature adaptation group further includes a ConvNeXt (AConvNet for the 2020s) block and a convolutional layer connected after the multiple degradation feature adaptation units. The present disclosure also does not limit this.
[0053] On this basis, optionally, step S103 includes: copying the teacher model to obtain the student model; using the third sample blind image as the training sample, and adopting the distillation learning method to perform model learning for the purpose of aligning the outputs of the degradation feature transformers in the teacher model and the student model, and using the learned student model as the final blind image super-resolution model. By first copying the teacher model to the student model, it can be ensured that the student model has the same model structure as the teacher model, ensuring that the model structure remains unchanged. In addition, the operations related to the degradation features are mainly concentrated in the degradation feature estimator and the degradation feature transformer. After that, only the channel-domain degradation features and spatial-domain degradation features obtained by conversion need to be adaptively fused with the image features extracted from the training samples. Based on this, by aligning the outputs of the degradation feature transformers in the teacher model and the student model during the distillation learning, the degradation-related knowledge learned in the teacher model can be concentrated and transferred to the student model, thereby improving the distillation learning efficiency. It should be understood that for the purpose of aligning the outputs of the degradation feature transformers in the teacher model and the student model means using the outputs of the degradation feature transformers in the two models as the basis for calculating the loss value, and then adjusting the parameters of the student model according to the loss value.
[0054] Next, a blind image super-resolution model and its training method according to a specific embodiment of the present disclosure will be introduced. In this specific embodiment, a set of low-resolution images (i.e., the aforementioned original blind images) for training are given and divided into three groups, which are respectively used for the training of three stages. The training process is as follows. Among them, step S1 is the basic model construction step, steps S2 to S6 correspond to the training of the first stage, steps S7 to S12 correspond to the training of the second stage, and steps S13 to S16 correspond to the training of the third stage.
[0055] Step S1: Construct a blind image super-resolution model based on implicit degradation feature guidance as Figure 3 shown. This model includes 1 image feature extractor, 1 degradation feature estimator, 1 degradation feature transformer, 1 degradation feature adaptation module, and 1 image reconstruction module. Among them, 1 degradation feature transformer and 1 degradation feature adaptation module form an image feature correction module.
[0056] The image feature extractor uses 1 3×3 convolutional layer.
[0057] The degradation feature estimator uses 1 six-layer convolutional network.
[0058] The degradation feature transformer includes 1 channel-domain conversion branch and 1 spatial-domain conversion branch. As Figure 2 shown, the former includes 1 global average pooling layer and 1 fully connected layer, and the latter includes 1 pixel recombination layer and 1 3×3 convolutional layer.
[0059] The structure of the degradation feature adaptation module is as Figure 3 shown. Specifically, the basic component of the degradation feature adaptation module is the degradation feature adaptation unit. A series of degradation feature adaptation units and convolutional blocks form a degradation feature adaptation group, and then a series of degradation feature adaptation groups and convolutional blocks form the degradation feature adaptation module. The structure of the degradation feature adaptation unit will be introduced below.
[0060] The image reconstruction module can adopt an upsampling module based on a convolutional neural network.
[0061] Step S2: In the first-stage training, first, construct a teacher network structure based on contrastive learning with degradation prior constraints. It should be noted that the teacher network is modified based on the above-mentioned blind image super-resolution model according to the requirements of the "contrastive learning method with degradation prior constraints". As Figure 4 shown, the teacher network includes two branches: the backbone branch and the momentum branch, both of which are composed of a degradation feature estimator and a projector. The degradation feature estimator is a six-layer convolutional network with a GELU (Gaussian Error Linear Unit) activation function, and the projector is a two-layer fully connected network with a GELU activation function.
[0062] Note that the degradation feature estimator in the backbone branch of the teacher network is Figure 3 physically the same as the degradation feature estimator in the Figure 4 blind image super-resolution model in
[0063] The degradation feature estimator in the momentum branch of the teacher network is consistent with the degradation feature estimator in the above-mentioned blind image super-resolution model in structure, but due to the different way of adjusting its parameters from the backbone branch, the parameters are different. The degradation feature estimator in the momentum branch is only used to assist in training the degradation feature estimator in the backbone branch. It should be noted that as Figure 5 shown in the upper row of the process in
[0064] Step S3: In the data preprocessing stage, first, for each low-resolution image in the first group of low-resolution images, randomly sample in this low-resolution image Patches of a certain size are used to form a sequence of low-resolution patches corresponding to the low-resolution image. Then, a dimensionality extension strategy can be adopted to extract the degradation prior (also known as the degradation reference prior) of each low-resolution patch. Finally, the degradation prior is incorporated into the low-resolution patches to form a sequence of training samples. .
[0065] Step S4: The sequence of training samples is fed into the aforementioned teacher network for training the backbone branch and the momentum branch. The backbone branch and the momentum branch respectively encode the training samples, where the outputs of the degradation feature estimators of the two branches are and respectively. Here, is due to two downsamplings in the degradation feature estimator, and 256 represents the feature dimension. The final outputs of the two branches are and respectively, where 128 represents the feature dimension.
[0066] In the teacher network, the parameter update of the backbone branch and the momentum branch in each round is carried out in two steps. The first step is to update the parameters of the backbone branch by using the contrastive loss between the outputs of the backbone branch and the momentum branch, and combining with the negative sample queue. The negative sample queue refers to the queue formed by the output features obtained by inputting other low-resolution patches different from the sequence of low-resolution patches calculated in this round into the backbone branch and / or the momentum branch. Specifically, the backbone features output by the backbone branch and / or the momentum features output by the momentum branch in the previous rounds of calculations can be used. Its loss function is defined as follows.
[0067]
[0068] Among them, represents the index of the output of the backbone branch, represents the index of the output of the momentum branch, and the similarity function adopts the InfoNCE (Information Noise-Contrastive Estimation) loss, which is defined as follows.
[0069]
[0070] Among them, represents the temperature coefficient, represents the negative sample queue, represents the length of, represents the th negative sample. Here, is a superscript rather than representing The calculation of the power is not marked as a subscript because it belongs to a different marking level from the subscripts i and j, so different marking positions are adopted to show the distinction.
[0071] In the first step, the random gradient descent algorithm is used as an optimizer to optimize the parameters of the backbone branch. Among them, the learning rate is set to 1e-3 until the loss value is less than or equal to the preset value.
[0072] Step S6: The second step is to update the parameters of the momentum branch using the momentum update strategy of contrastive learning, that is, after updating the parameters of the backbone branch, update the parameters of the momentum branch based on the parameters of the backbone branch, as follows.
[0073]
[0074] Among them, and are the parameters of the backbone branch and the momentum branch respectively, is the momentum coefficient.
[0075] After the first stage of training, the degenerate feature estimator in the backbone branch can be obtained as the first degenerate feature estimator, and it is embedded into the blind image super-resolution model as the training object of the second stage.
[0076] Step S7: In the second stage of training, for the second group of low-resolution images, the same method as in step S3 is used to randomly sample the low-resolution patch sequence corresponding to each low-resolution image , as a training sample sequence. As Figure 5 shown, in each round of calculation, a training sample sequence is fed into the blind image super-resolution model obtained in the first stage. First, the image feature extractor is used to extract the image features of each low-resolution patch (hereinafter referred to as low-resolution patch features) .
[0077] Step S8: Use the first degenerate feature estimator obtained in the first stage of training to extract the degenerate features , and use the degenerate feature converter to expand the output by the first degenerate feature estimator at the channel level and the spatial level. Specifically, as Figure 2 shown, the channel domain conversion branch performs global average pooling operation on , and then sends the pooling result into the fully connected layer for non-linear transformation to obtain the channel domain degenerate feature ; the spatial domain conversion branch uses the pixel recombination layer to upsample to the same size as the low-resolution patch, and then sends the upsampled feature map into a 3×3 convolutional layer to obtain the spatial domain degenerate feature .
[0078] Step S9: Incorporate the degradation features into the low-resolution patch features using the degradation feature adaptation module.
[0079] Specifically, for the degradation feature adaptation unit, as Figure 6 shown, first divide the input low-resolution patch features into two parts along the channel dimension in a 1:3 ratio. In the spatial domain adaptation branch, first, concatenate with along the channel dimension, and then modulate through two 3×3 convolutional layers and a Sigmoid function. Next, multiply the modulated spatial domain degradation features by the original , and then follow with a 5×5 convolutional layer for further fusion to obtain the low-resolution patch features after spatial domain degradation adaptation. In the channel domain adaptation branch, perform GAP (Global Average Pooling) on the remaining , then concatenate the pooling result with along the channel dimension, and then modulate through two fully connected layers and a Sigmoid function. Next, multiply the modulated channel domain degradation features by the original to obtain the low-resolution patch features after channel domain degradation adaptation. Finally, concatenate and along the channel dimension and feed them into a ConvNeXt block for further fusion. Finally, once again introduce the low-resolution patch features using skip connections to restore high-frequency details. This process can be expressed as:
[0080]
[0081] where is a ConvNeXt block containing 1 depthwise separable convolutional layer of 7×7 and 2 convolutional layers of 1×1. represents the image features processed by the current degradation feature adaptation unit.
[0082] For the degradation feature adaptation group, it consists of 4 degradation feature adaptation units, 1 ConvNeXt block, and 1 3×3 convolutional layer. Within the degradation feature adaptation group, the low-resolution patch features and the degradation features are sequentially fed into each component for processing, and skip connections are added between the input and output of the degradation feature adaptation group.
[0083] For the degradation feature adaptation module, it is composed of 3 degradation feature adaptation groups connected in series with a 3×3 convolutional layer. Within the degradation feature adaptation module, the low-resolution patch features and the degradation features are sequentially fed into each component for processing, and a skip connection is added between the input and output of the degradation feature adaptation module. After being processed by the degradation feature adaptation module, the low-resolution patch features that fully integrate the degradation information can be obtained , that is, the corrected image features.
[0084] Step S10: Use the image reconstruction module with the foregoing as the input to reconstruct the high-resolution patch (or high-resolution block) corresponding to the foregoing low-resolution patch. The image reconstruction process can be expressed as:
[0085]
[0086] wherein, represents the super-resolution module based on the convolutional neural network, represents the reconstructed high-definition image.
[0087] Step S11: Calculate the L1 loss between the real high-resolution patch and the reconstructed high-resolution patch as follows:
[0088]
[0089] Step S12: Use the stochastic gradient descent algorithm as the optimizer to optimize the parameters of the entire blind image super-resolution model. Among them, the learning rate is gradually adjusted from 2e-4 to 1e-6 using the cosine annealing adjustment strategy, and the number of iterations is set to 600. The trained blind image super-resolution model is used as the teacher model. It should be understood that the teacher model here and the foregoing teacher network are different concepts.
[0090] Step S13: In the training of the student model in the third stage, first construct a student model as shown in Figure 7 . Note that the student model can be obtained by copying the teacher model. As shown in Figure 7 , the student model mainly trains a degradation feature estimator and a degradation feature converter.
[0091] Step S14: For each low-resolution image in the third group of low-resolution images, still use the method introduced in Step S3 to obtain the corresponding sequence of low-resolution patches, which is fed into the student model as a training sample sequence for training the degradation feature estimator and the degradation feature transformer. For easy distinction, the degradation feature transformer in the student model is called the student degradation feature transformer, and the degradation feature transformer in the teacher model is called the teacher degradation feature transformer. As mentioned above, the outputs of the spatial domain transformation branch and the channel domain transformation branch in the teacher degradation feature transformer are respectively defined as and respectively. Correspondingly, the outputs of the spatial domain transformation branch and the channel domain transformation branch in the student degradation feature transformer are respectively defined as and .
[0092] Step S15: As shown in Figure 8 , knowledge distillation is achieved by aligning the outputs between the teacher degradation feature transformer and the student degradation feature transformer. For the distillation of spatial domain degradation features, the L2 loss is used to perform the knowledge transfer from to :
[0093]
[0094] For the distillation of channel domain degradation features, the L1 loss is used to minimize the absolute difference between and :
[0095]
[0096] The total loss of knowledge distillation is:
[0097]
[0098] Use the stochastic gradient descent algorithm as the optimizer to optimize the parameters of the degradation feature estimator and the degradation feature transformer of the student model, where the learning rate is set to 1e-3 until the loss value is less than or equal to the preset value.
[0099] Step S16: Save the model parameters at the end of training, and the expected lightweight blind image super-resolution model based on implicit degradation feature guidance can be obtained, whose number of parameters and computational complexity are reduced by more than 50% compared with the current state-of-the-art models.
[0100] Figure 9 is a flowchart of a blind image super-resolution method according to an exemplary embodiment of the present disclosure. This method can be executed on an electronic device with sufficient computing power.
[0101] Refer to Figure 9, in step S901, a degradation feature estimator in the blind image super-resolution model is used to process the blind image to obtain the degradation features of the blind image. The blind image super-resolution model further includes an image feature correction module and an image reconstruction module.
[0102] In step S902, the image feature correction module is used to correct the image features of the blind image according to the degradation features to obtain the corrected image features.
[0103] In step S903, the image reconstruction module is used to reconstruct the high-definition image corresponding to the blind image based on the corrected image features.
[0104] In the blind image super-resolution method according to the exemplary embodiments of the present disclosure, the blind image super-resolution model is trained by the training method of the blind image super-resolution model according to the exemplary embodiments of the present disclosure, and thus has all the beneficial technical effects of this training method, which will not be elaborated herein.
[0105] Figure 10 is a block diagram of a blind image super-resolution device according to the exemplary embodiments of the present disclosure.
[0106] The blind image super-resolution model includes a degradation feature estimator for determining the degradation features of the blind image, an image feature correction module for correcting the image features of the blind image according to the degradation features, and an image reconstruction module for reconstructing the high-definition image based on the corrected image features.
[0107] Referring to Figure 10 , the training device 1000 of the blind image super-resolution model includes a first training unit 1001, a second training unit 1002, and a third training unit 1003.
[0108] The first training unit 1001 can determine the training samples in the first stage based on the first sample blind image, and perform the first stage training on the to-be-trained degradation feature estimator by using the contrast learning method based on the degradation prior constraint to obtain the first degradation feature estimator.
[0109] The second training unit 1002 can determine the training samples in the second stage based on the second sample blind image, and perform the second stage training on the blind image super-resolution model including the first degradation feature estimator to obtain the teacher model.
[0110] The third training unit 1003 can determine the training samples in the third stage based on the third sample blind image, use the teacher model as the learning object, and perform the third stage training by using the distillation learning method to obtain the final blind image super-resolution model.
[0111] Optionally, the first training unit 1001 may also: construct a backbone branch and a momentum branch, where both the backbone branch and the momentum branch include a degradation feature estimator; perform degradation prior constraint processing on the first sample blind image to obtain a training blind image; use the training blind image as a training sample, and adopt a contrastive learning method to train the backbone branch and the momentum branch, and use the degradation feature estimator in the trained backbone branch as the first degradation feature estimator.
[0112] Optionally, the first training unit 1001 may also: perform dimensional extension processing on the blur kernel and noise information in the first sample blind image to obtain a degradation prior with the same spatial dimension as the first sample blind image; splice the degradation prior and the first sample blind image along the channel dimension to obtain a training blind image.
[0113] Optionally, the first training unit 1001 may also: perform the following iterative calculation steps on the backbone branch and the momentum branch until the stop condition is met, and use the degradation feature estimator in the latest backbone branch as the first degradation feature estimator, where the number of training blind images is multiple, and the training blind images used in each round of calculation are different: input the training blind images of this round of calculation into the backbone branch and the momentum branch respectively to obtain backbone features and momentum features; obtain reference features, where the reference features are features obtained by inputting an image different from the training blind image of this round of calculation into the backbone branch and / or the momentum branch; determine a contrast loss value according to the backbone features, the momentum features and the reference features; adjust the parameters of the backbone branch according to the contrast loss value, and use the momentum update strategy of contrastive learning to adjust the parameters of the momentum branch based on the parameter adjustment of the backbone branch.
[0114] Optionally, the image feature correction module includes a degradation feature converter and a degradation feature adaptation module. The degradation feature converter includes a channel domain conversion branch and a spatial domain conversion branch. The channel domain conversion branch is used to convert the degradation feature into a channel domain degradation feature, and the spatial domain conversion branch is used to convert the degradation feature into a spatial domain degradation feature. The degradation feature adaptation module includes multiple degradation feature adaptation groups. Each degradation feature adaptation group includes multiple degradation feature adaptation units. Each degradation feature adaptation unit includes a feature splitter, a channel domain adaptation branch, a spatial domain adaptation branch, and a feature fuser. The feature splitter is used to split the input image feature into a first image feature and a second image feature. The channel domain adaptation branch is used to integrate the channel domain degradation feature into the first image feature to obtain a first corrected image feature. The spatial domain adaptation branch is used to integrate the spatial domain degradation feature into the second image feature to obtain a second corrected image feature. The feature fuser is used to fuse the first corrected image feature and the second corrected image feature into a new image feature.
[0115] Optionally, the third training unit 1003 may further: copy the teacher model to obtain a student model; use the third sample blind image as a training sample, and adopt a distillation learning method to perform model learning for the purpose of aligning the outputs of the degradation feature transformers in the teacher model and the student model, and use the learned student model as the final blind image super-resolution model.
[0116] Figure 11 is a block diagram of a blind image super-resolution device according to an exemplary embodiment of the present disclosure. Referring to Figure 11 , the blind image super-resolution device 1100 includes a degradation estimation unit 1101, a correction unit 1102, and a reconstruction unit 1103.
[0117] The degradation estimation unit 1101 may use the degradation feature estimator in the blind image super-resolution model to process the blind image to obtain the degradation features of the blind image. The blind image super-resolution model further includes an image feature correction module and an image reconstruction module.
[0118] The correction unit 1102 may be configured to use the image feature correction module to correct the image features of the blind image according to the degradation features to obtain the corrected image features.
[0119] The reconstruction unit 1103 may be configured to use the image reconstruction module to reconstruct the high-definition image corresponding to the blind image based on the corrected image features. The blind image super-resolution model is trained by the training method of the blind image super-resolution model according to an exemplary embodiment of the present disclosure.
[0120] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0121] Figure 12 shows a block diagram of an electronic device 1200 according to an exemplary embodiment of the present disclosure.
[0122] Referring to Figure 12 , the electronic device 1200 includes: at least one memory 1201 and at least one processor 1202. Computer-executable instructions are stored in the at least one memory 1201. When the computer-executable instructions are run by the at least one processor 1202, the at least one processor is caused to execute the target corresponding method as described in the above exemplary embodiment.
[0123] As an example, the electronic device 1200 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 1200 does not have to be a single electronic device 1200, but can also be a collection of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The electronic device 1200 can also be a part of an integrated control system or system manager, or can be configured as a portable electronic device 1200 that interfaces with a local or remote device (e.g., via wireless transmission).
[0124] In the electronic device 1200, the processor 1202 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 1202 can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
[0125] The processor 1202 can run instructions or code stored in the memory 1201, where the memory 1201 can also store data. The instructions and data can also be sent and received over a network via a network interface device, where the network interface device can use any known transmission protocol.
[0126] The memory 1201 can be integrated with the processor 1202. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 1201 can include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 1201 and the processor 1202 can be operatively coupled or can communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 1202 can read files stored in the memory.
[0127] In addition, the electronic device 1200 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 1200 can be connected to each other via a bus and / or a network.
[0128] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein the instructions, when run by at least one processor, cause the at least one processor to execute the target corresponding method as described in the above exemplary embodiment. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0129] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including computer instructions that, when run by at least one processor, execute the target corresponding method as described in the above exemplary embodiment.
[0130] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0131] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for a blind image super-resolution model, characterized in that The blind image super-resolution model includes a degradation feature estimator for determining the degradation features of a blind image, an image feature correction module for correcting the image features of the blind image according to the degradation features, and an image reconstruction module for reconstructing a high-definition image based on the corrected image features. Wherein, the training method includes: Determine the training samples in the first stage based on the first sample blind image, and use the contrast learning method based on the degradation prior constraint to perform the first stage of training on the to-be-trained degradation feature estimator to obtain the first degradation feature estimator; Determine the training samples in the second stage based on the second sample blind image, and perform the second stage of training on the blind image super-resolution model including the first degradation feature estimator to obtain the teacher model; Determine the training samples in the third stage based on the third sample blind image, use the teacher model as the learning object, and perform the third stage of training using the distillation learning method to obtain the final blind image super-resolution model; Wherein, the determining the training samples in the first stage based on the first sample blind image, using the contrast learning method based on the degradation prior constraint to perform the first stage of training on the to-be-trained degradation feature estimator to obtain the first degradation feature estimator includes: Construct a backbone branch and a momentum branch, wherein both the backbone branch and the momentum branch include a degradation feature estimator; Perform dimensional extension processing on the blur kernel and noise information in the first sample blind image to obtain a degradation prior with the same spatial dimension as the first sample blind image; Concatenate the degradation prior and the first sample blind image along the channel dimension to obtain a training blind image; Use the training blind image as a training sample, and use the contrast learning method to train the backbone branch and the momentum branch, and use the degradation feature estimator in the trained backbone branch as the first degradation feature estimator.
2. The training method according to claim 1, characterized in that The using the training blind image as a training sample, using the contrast learning method to train the backbone branch and the momentum branch, and using the degradation feature estimator in the trained backbone branch as the first degradation feature estimator includes: Perform the following iterative calculation steps on the backbone branch and the momentum branch until the stop condition is met, and use the degradation feature estimator in the latest backbone branch as the first degradation feature estimator. Wherein, the number of the training blind images is multiple, and the training blind images used in each round of calculation are different: Input the training blind image of this round of calculation into the backbone branch and the momentum branch respectively to obtain the backbone feature and the momentum feature; Obtain a reference feature, where the reference feature is a feature obtained by inputting an image different from the training blind image of this round of calculation into the backbone branch and / or the momentum branch; Determine the contrast loss value according to the backbone feature, the momentum feature and the reference feature; Adjust the parameters of the backbone branch according to the contrast loss value, and use the momentum update strategy of contrast learning to adjust the parameters of the momentum branch based on the parameters of the backbone branch.
3. The training method according to claim 1 or 2, wherein The image feature correction module includes a degradation feature converter and a degradation feature adaptation module, wherein, the degradation feature converter includes a channel domain conversion branch and a spatial domain conversion branch. The channel domain conversion branch is used to convert the degradation feature into a channel domain degradation feature, and the spatial domain conversion branch is used to convert the degradation feature into a spatial domain degradation feature. The degradation feature adaptation module includes multiple degradation feature adaptation groups. Each degradation feature adaptation group includes multiple degradation feature adaptation units. Each degradation feature adaptation unit includes a feature splitter, a channel domain adaptation branch, a spatial domain adaptation branch, and a feature fuser. The feature splitter is used to split the input image feature into a first image feature and a second image feature. The channel domain adaptation branch is used to incorporate the channel domain degradation feature into the first image feature to obtain a first corrected image feature. The spatial domain adaptation branch is used to incorporate the spatial domain degradation feature into the second image feature to obtain a second corrected image feature. The feature fuser is used to fuse the first corrected image feature and the second corrected image feature into a new image feature.
4. The training method according to claim 3, wherein, Determining the training samples for the third stage based on the third sample blind image, taking the teacher model as the learning object, and using the distillation learning method for the third stage training to obtain the final blind image super-resolution model, including: Copying the teacher model to obtain a student model; Taking the third sample blind image as the training sample, using the distillation learning method, aiming to align the outputs of the degradation feature converters in the teacher model and the student model, performing model learning, and taking the learned student model as the final blind image super-resolution model.
5. A blind image super-resolution method, characterized in that, Including: Using the degradation feature estimator in the blind image super-resolution model to process the blind image to obtain the degradation feature of the blind image. The blind image super-resolution model further includes an image feature correction module and an image reconstruction module; Using the image feature correction module to correct the image feature of the blind image according to the degradation feature to obtain a corrected image feature; Using the image reconstruction module to reconstruct the high-definition image corresponding to the blind image based on the corrected image feature, wherein, the blind image super-resolution model is trained by the training method of the blind image super-resolution model according to any one of claims 1 to 4.
6. An electronic device, characterized in that, Including: At least one processor; At least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method of the blind image super-resolution model according to any one of claims 1 to 4 or the blind image super-resolution method according to claim 5.
7. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the training method of the blind image super-resolution model according to any one of claims 1 to 4 or the blind image super-resolution method according to claim 5.
8. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by at least one processor, the at least one processor is caused to execute the training method of the blind image super-resolution model according to any one of claims 1 to 4 or the blind image super-resolution method according to claim 5.