Super-division model training method, image super-division method, device, equipment and product
By performing noise processing on the image and extracting the fusion attention feature, the image super-score model is constructed, which solves the problems of low fidelity and introduction of distorted elements of image super-score processing in the prior art, and achieves higher quality image super-score results.
Patent Information
- Application Number
- CN202510192772.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
When the prior art deals with complex degradation conditions, the fidelity of the generated image is low, and distorted elements that do not exist in the original image are easily introduced, and image prior information is not fully utilized.
By acquiring the sample image and the first image, the reference image corresponding to the first image is determined; the sample image is noise-added based on the sample image data to obtain the noise-added image data; the pre-constructed initial model is obtained, and the initial model is fused with attention feature extraction of the first image, reference image and noise-added image data to determine the predicted image noise; the model loss function is constructed based on the predicted image noise and actual image noise, and the initial model is trained to obtain the image super-score model.
The generation quality of the model's image super-score results is improved, the processing ability of low-definition images is enhanced, the possibility of introducing distorted elements is reduced, and the image prior information is fully utilized.
Smart Images

Figure CN120107066A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a training method for an image super-resolution model, an image super-resolution processing method, a training device for an image super-resolution model, an image super-resolution processing device, an electronic device, and a computer program product. Background Art
[0002] Image super-resolution (abbreviated as super-resolution) refers to the process of restoring a perceptually realistic high-resolution image from a low-resolution image. Since images are inevitably affected by degradation such as blur and noise during acquisition and transmission, image super-resolution has become an important task with wide applications in many fields. With the development of deep learning technology, algorithms based on generative adversarial networks (GAN) and diffusion models have become the mainstream choice for image super-resolution. Summary of the invention
[0003] The present disclosure provides a training method for an image super-resolution model, an image super-resolution processing method, an image super-resolution model training device, an image super-resolution processing device, an electronic device, a computer-readable storage medium, and a computer program product, to at least solve the problems in the related art of generating images with low fidelity in complex degradation situations, introducing distortion elements that do not exist in the original image, and not fully utilizing image prior information. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a method for training an image super-resolution model is provided, comprising: acquiring a sample image and a first image, and determining a reference image corresponding to the first image; performing noise processing on the sample image based on the sample image data to obtain noisy image data; acquiring a pre-constructed initial model, and using the initial model to perform fusion attention feature extraction on the first image, the reference image and the noisy image data to determine predicted image noise; determining actual image noise contained in the noisy image data, and constructing a model loss function based on the predicted image noise and the actual image noise; and training the initial model based on the model loss function to obtain an image super-resolution model.
[0005] In an exemplary embodiment of the present disclosure, determining a reference image corresponding to the first image includes: acquiring a pre-constructed image activation model; and performing denoising and optimization processing on the first image using the image activation model to obtain the reference image.
[0006] In an exemplary embodiment of the present disclosure, the denoising process is performed on the sample image based on the sample image data to obtain the noisy image data, including: encoding the sample image to obtain the encoded sample image data; obtaining pre-configured random noise and diffusion steps, and adding the random noise to the sample image data based on the diffusion steps to obtain the noisy image data.
[0007] In an exemplary embodiment of the present disclosure, the initial model includes a control network module and a reference fusion module, and the initial model performs fusion attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise, including: the control network module performs feature extraction processing on the first image and the reference image respectively to obtain first image features and reference image features; the reference fusion module performs fusion attention feature extraction on the first image features, the reference image features and the noisy image data to obtain output image features; the output image features are decoded to obtain the predicted image noise.
[0008] In an exemplary embodiment of the present disclosure, the reference fusion module includes a reference attention module, a self-attention module and a first attention module; the reference fusion module performs fusion attention feature extraction on the first image feature, the reference image feature and the noisy image data to obtain an output image feature, including: determining the noisy image feature corresponding to the noisy image data based on the self-attention module; inputting the first image feature and the reference image feature into the reference attention module to obtain an enhanced sample feature after the first image and the reference image are fused; inputting the enhanced sample feature and the noise image feature into the first attention module to obtain the output image feature.
[0009] In an exemplary embodiment of the present disclosure, the first image feature and the reference image feature are input into the reference attention module to obtain an enhanced sample feature after the first image and the reference image are fused, including: the reference attention module performs attention feature extraction on the first image feature to obtain a query image feature; the reference attention module performs attention feature extraction on the reference image feature to obtain a key image feature and a value image feature; and normalization is performed according to the query image feature, the key image feature and the value image feature to obtain the enhanced sample feature.
[0010] In an exemplary embodiment of the present disclosure, the enhanced sample features and the noise image features are input into the first attention module to obtain the output image features, including: the first attention module performs attention feature extraction on the noise image features to obtain query enhancement features; the first attention module performs attention feature extraction on the enhanced sample features to obtain key enhancement features and value enhancement features; and normalization is performed according to the query enhancement features, the key enhancement features, and the value enhancement features to obtain the output image features.
[0011] According to a second aspect of an embodiment of the present disclosure, there is provided an image super-resolution processing method, comprising: obtaining an image to be processed, and determining a reference image to be processed corresponding to the image to be processed; obtaining an image super-resolution model trained according to the above-mentioned image super-resolution model training method; inputting the image to be processed and the reference image to be processed into the image super-resolution model, and the image super-resolution model outputs a super-resolution reconstructed image corresponding to the image to be processed.
[0012] According to a third aspect of an embodiment of the present disclosure, a training device for an image super-resolution model is provided, comprising: a reference image determination module, used to obtain a sample image and a first image, and determine a reference image corresponding to the first image; an image denoising module, used to perform denoising processing on the sample image based on sample image data to obtain noisy image data; a noise prediction module, used to obtain a pre-constructed initial model, and the initial model performs fusion attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise; a loss function determination module, used to determine the actual image noise contained in the noisy image data, and construct a model loss function based on the predicted image noise and the actual image noise; and a model training module, used to train the initial model based on the model loss function to obtain an image super-resolution model.
[0013] In an exemplary embodiment of the present disclosure, the reference image determination module includes a reference image determination unit, which is used to: obtain a pre-constructed image activation model; and use the image activation model to perform denoising and optimization processing on the first image to obtain the reference image.
[0014] In an exemplary embodiment of the present disclosure, the image denoising module includes an image denoising unit, which is used to: encode the sample image to obtain encoded sample image data; obtain pre-configured random noise and diffusion steps, and add the random noise to the sample image data based on the diffusion steps to obtain the noisy image data.
[0015] In an exemplary embodiment of the present disclosure, the initial model includes a control network module and a reference fusion module, and the noise prediction module includes a noise prediction unit, which is used to: the control network module performs feature extraction processing on the first image and the reference image respectively to obtain first image features and reference image features; the reference fusion module performs fusion attention feature extraction on the first image features, the reference image features and the noisy image data to obtain output image features; and the output image features are decoded to obtain the predicted image noise.
[0016] In an exemplary embodiment of the present disclosure, the reference fusion module includes a reference attention module, a self-attention module and a first attention module; the noise prediction unit includes an output feature determination unit, which is used to: determine the noise image features corresponding to the noisy image data based on the self-attention module; input the first image features and the reference image features into the reference attention module to obtain enhanced sample features after the first image and the reference image are fused; input the enhanced sample features and the noise image features into the first attention module to obtain the output image features.
[0017] In an exemplary embodiment of the present disclosure, the output feature determination unit includes an enhanced feature determination subunit, which is used to: perform attention feature extraction on the first image feature by the reference attention module to obtain a query image feature; perform attention feature extraction on the reference image feature by the reference attention module to obtain a key image feature and a value image feature; perform normalization processing based on the query image feature, the key image feature and the value image feature to obtain the enhanced sample feature.
[0018] In an exemplary embodiment of the present disclosure, the output feature determination unit includes an output feature determination sub-unit, which is used to: perform attention feature extraction on the noise image feature by the first attention module to obtain a query enhancement feature; perform attention feature extraction on the enhanced sample feature by the first attention module to obtain a key enhancement feature and a value enhancement feature; perform normalization processing based on the query enhancement feature, the key enhancement feature and the value enhancement feature to obtain the output image feature.
[0019] According to a fourth aspect of an embodiment of the present disclosure, there is provided an image super-resolution processing device, comprising: an image determination module, used to acquire an image to be processed and determine a reference image to be processed corresponding to the image to be processed; a model acquisition module, used to acquire an image super-resolution model trained according to the above-mentioned image super-resolution model training method; and a super-resolution reconstruction module, used to input the image to be processed and the reference image to be processed into the image super-resolution model, and the image super-resolution model outputs a super-resolution reconstructed image corresponding to the image to be processed.
[0020] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to execute instructions to implement any one of the above-described image super-resolution model training methods, or to implement any one of the above-described image super-resolution processing methods.
[0021] According to the sixth aspect of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute any one of the above-mentioned image super-resolution model training methods and implement any one of the above-mentioned image super-resolution processing methods.
[0022] According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements any one of the above-mentioned image super-resolution model training methods and any one of the above-mentioned image super-resolution processing methods.
[0023] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0024] On the one hand, adding control over the input image during model training can make full use of the prior knowledge of the diffusion model, and to a certain extent, solve the problem that related solutions perform poorly when the low-definition image has a high degree of degradation. On the other hand, by fusing feature extraction of the model input data, the information of the input image can be mined and the quality of the model's generation of image super-resolution results can be improved.
[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0027] Figure 1 The figure is a flowchart of a method for training an image super-resolution model according to an exemplary embodiment.
[0028] Figure 2 It is a model framework diagram of an image super-resolution model according to an exemplary embodiment.
[0029] Figure 3 is a process diagram of feature aggregation based on a reference fusion module according to an exemplary embodiment.
[0030] Figure 4The figure is a flowchart of an image super-resolution processing method according to an exemplary embodiment.
[0031] Figure 5 It is a workflow diagram of an image super-resolution model according to an exemplary embodiment.
[0032] Figure 6 It is a comparison diagram of an input image to be processed and an output super-resolution reconstructed image according to an exemplary embodiment.
[0033] Figure 7 It is a block diagram of a training device for an image super-resolution model according to an exemplary embodiment.
[0034] Figure 8 The figure is a block diagram of an image super-resolution processing device according to an exemplary embodiment.
[0035] Fig. 9 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0036] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0038] GAN-based and diffusion model-based algorithms have become the mainstream choices for image super-resolution. However, they have different problems, such as: (1) GAN-based methods are difficult to generate realistic and vivid textures because their optimization goals focus too much on fidelity and there is a domain bias between training data and real test data; (2) Text to Image (T2I) diffusion model-based methods mostly rely on carefully designed text cues and semantic priors, but it is difficult to design corresponding text cues for images that have undergone complex degradation. Due to the powerful generation ability of the diffusion model, in this case, it usually generates content that deviates from the ground truth (GT) or introduces elements and artifacts that do not exist in the GT, resulting in a decrease in the quality of the super-resolution results. To address the above problems, a cascade diffusion model algorithm for image super-resolution is proposed in related solutions.
[0039] At present, most representative diffusion model-based image super-resolution algorithms improve the performance of T2I models in super-resolution tasks by designing more advanced language cues. For example, multimodal transformer models such as the BLIP model extract image annotation information from low-definition images and use the CLIP encoder to convert text information into image-level features to provide additional semantic information to guide the diffusion process. Alternatively, a pre-trained label model such as the Recognize Anything Model (RAM) is fine-tuned to obtain a language cue extractor with degradation perception capabilities, and the obtained language cues are combined with low-definition images to control the generation results of the pre-trained T2I model.
[0040] The language cue extractor designed by the above method is difficult to obtain high-quality generation results when facing complex scenes in practical applications. This is because the design of language cues has its inevitable limitations. On the one hand, its language cues cannot contain detailed information such as image distortion. On the other hand, the pre-trained label model it uses is difficult to provide information such as spatial position and global scene, and cannot fully describe the image after multiple degradations. In addition, its super-resolution process simply interpolates (such as bicubic interpolation) and upsamples the low-definition image, and does not fully utilize the image prior information.
[0041] Based on this, according to the embodiments of the present disclosure, a training method for an image super-resolution model, an image super-resolution processing method, a training device for an image super-resolution model, an image super-resolution processing device, an electronic device, a computer-readable storage medium and a computer program product are proposed.
[0042] Figure 1 is a flowchart of a method for training an image super-resolution model according to an exemplary embodiment. Figure 1As shown, the training method of the image super-resolution model can be used in a computer device, wherein the computer device described in the present disclosure may include mobile terminal devices such as mobile phones, tablet computers, laptops, PDAs, and fixed terminal devices such as desktop computers. This exemplary embodiment uses the method applied to a computer device as an example, and it can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.
[0043] In step S110, a sample image and a first image are acquired, and a reference image corresponding to the first image is determined.
[0044] In an exemplary embodiment of the present disclosure, the sample image and the first image pair are images that can be acquired in advance for model training and are included in the training data set of the model. The sample image can be a ground truth (GT) image, an original image, etc. The first image can be an image obtained after the sample image produces image degradation phenomenon. The sample image is generated after the image quality is reduced during the formation, recording, processing and transmission process. As the first image, the image degradation includes reduced resolution, blurred edges, increased noise, etc., that is, the first image is a low-definition image of the sample image. In this embodiment, it can be determined whether a certain image belongs to the first image according to the image parameter value. For example, the first image can be an image with an image resolution lower than a resolution threshold, the first image can also be an image with an edge sharpness value less than an edge sharpness threshold, and the first image can also be an image with a peak signal-to-noise ratio (PSNR) greater than a signal-to-noise ratio threshold. The reference image can be an image generated after denoising and optimizing the first image.
[0045] In order to adapt to the situation where the input image in the image super-resolution task is a low-definition image, the first image (i.e., the low-definition image of the sample image) is selected as one of the inputs of the model in this embodiment. The number of pixels of the low-definition image is small, resulting in poor image clarity and detail expression. For example, an image with an image resolution less than a resolution threshold is used as the first image. After the first image is obtained, the first image is preliminarily optimized, and the input first image is subjected to denoising optimization processing to reduce the degradation of the first image, and the result (i.e., the reference image) is used as one of the inputs of the model.
[0046] In step S120, the sample image is subjected to noise addition processing based on the sample image data to obtain noise added image data.
[0047] In an exemplary embodiment of the present disclosure, the sample image data may be data obtained after encoding the sample image, for example, the sample image data may be represented in a matrix form. The noise addition process may be an operation of adding random noise to the sample image. The noisy image data may be data obtained after adding noise to the sample image data.
[0048] In actual image super-resolution processing tasks, the input image may contain more or less noise, resulting in the introduction of non-existent distortion elements in the original image. To address the above problem, before model training, the sample image is encoded to obtain the corresponding sample image data, and then noise is added to the sample image data to obtain noisy image data, which is used as one of the inputs of the model.
[0049] In step S130, a pre-built initial model is obtained, and the initial model performs fusion attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise.
[0050] In an exemplary embodiment of the present disclosure, the initial model may be a pre-built network model, and the initial model includes multiple attention modules. The fused attention feature extraction may be a processing operation of extracting features and fusing features on multiple model input data using the network structure in the initial model. The predicted image noise may be the noise added to the sample image predicted after the initial model extracts features from the model input data.
[0051] In order to reduce the impact of the degradation of the input image on the super-resolution reconstruction task, the present invention uses the reference image and the original low-definition image (i.e., the first image) as input to more accurately control the model generation process. A pre-built initial model is obtained, and the low-definition image, the reference image, and the noisy image data are used as the model input data of the initial model. Feature extraction is performed based on the attention mechanism, and the predicted image noise corresponding to the output image is output. For example, the backbone network of the initial model extracts features based on the features of the low-definition image and the features of the reference image to predict the added noise.
[0052] In step S140, the actual image noise contained in the noisy image data is determined, and a model loss function is constructed based on the predicted image noise and the actual image noise.
[0053] In an exemplary embodiment of the present disclosure, the actual image noise may be the noise actually added to the sample image. The model loss function may be a loss function constructed based on the predicted image noise and the actual image noise, and is used to train the initial model.
[0054] For noisy image data, the actual image noise contained therein can be determined, and a model loss function can be constructed based on the predicted image noise and the actual image noise for model training.
[0055] In step S150, the initial model is trained based on the model loss function to obtain an image super-resolution model.
[0056] In an exemplary embodiment of the present disclosure, the image super-resolution model may be a model for performing an image super-resolution reconstruction task.
[0057] During the model training process, the parameters of the model part to be trained are adjusted, and the final model parameters are determined according to the model training objectives, and the image super-resolution model is determined.
[0058] According to the training method of the image super-resolution model in this example embodiment, on the one hand, by adding the control of the input image during the model training process, the prior knowledge of the diffusion model can be fully utilized, and to a certain extent, the problem that the related scheme performs poorly when the low-definition image degradation degree is high. On the other hand, by fusing the feature extraction of the model input data, the information of the input image can be mined, and the quality of the model's generation of the image super-resolution result can be improved.
[0059] Next, the training method of the image super-resolution model in this example embodiment will be further described.
[0060] In an exemplary embodiment of the present disclosure, for step S110, determining a reference image corresponding to the first image includes: acquiring a pre-constructed image activation model; and performing denoising and optimization processing on the first image using the image activation model to obtain a reference image.
[0061] The image activation model, also known as the image activation module, may be a network model that performs image denoising on an input image to optimize the input image, thereby reducing the degree of degradation of the input first image (i.e., the low-definition image). The denoising optimization process may be a process operation that performs image denoising on an image to optimize the image quality.
[0062] The image super-resolution processing disclosed in the present invention mainly includes two stages, the first stage is implemented by the image activation model, and the second stage is implemented by the reference fusion model. In the first stage, a pre-built image activation model is obtained, such as the image activation model can be a cascade diffusion model, and the image activation model is used to perform preliminary optimization on the input low-definition image.
[0063] Specifically, in the first stage, the image activation module aims to optimize the image through a relatively simple denoising method, so as to reduce the degradation of the input low-definition image. In this stage, only a lightweight, computationally efficient controllable image diffusion generation model is needed, such as a bidirectional super-resolution generative adversarial network model (BSRGAN) for image super-resolution reconstruction to implement the image activation model. After the first stage, the input low-definition image can generate a high-definition image with richer image details, less degradation, and closer to the target. As a reference image, the degradation of the image is reduced, which can provide more information for the subsequent pre-trained model to guide its diffusion process.
[0064] In an exemplary embodiment of the present disclosure, for step S120, the sample image is denoised based on the sample image data to obtain noisy image data, and the steps include: encoding the sample image to obtain encoded sample image data; obtaining pre-configured random noise and diffusion steps, and adding random noise to the sample image data based on the diffusion steps to obtain noisy image data.
[0065] The sample image may be a reference truth image, and the reference truth may be a true value or label of the training data, i.e., an optimization output target of the model. The sample image data may be image data generated after encoding the sample image. Random noise may be unnecessary or redundant interference information present in the image data. For example, image noise may be a random change in brightness or color information in the image. The number of diffusion steps may be the number of steps corresponding to the diffusion process of the initial model.
[0066] The model training of the present disclosure is divided into two stages, namely, the training of the image activation module and the training of the complete model. For the image activation module, any pre-trained image super-resolution model can be used. For the complete model, a sample image with noise added is required as one of the inputs for the model training process.
[0067] In the first stage, a sample image, i.e., a GT image, is obtained first. The GT image is encoded using a pre-trained variational auto-encoder (VAE) with fixed parameters to obtain the corresponding encoding result, i.e., the sample image data, which is used for z 0 Then, based on the encoded sample image data, the sample image is subjected to noise addition processing.
[0068] Since the initial model of the present disclosure is a cascade diffusion model, the denoising process of the input image is performed based on the number of diffusion steps. The diffusion model includes two steps: a fixed (or preset) forward diffusion process, which can gradually add noise to the image until pure noise is finally obtained. The trainable reverse denoising diffusion process refers to training a network structure to gradually denoise from pure noise until a real image is obtained. The specific denoising process is: the encoded sample image data z 0 Add random noise ∈ to get the noisy image data z t , where t represents the number of diffusion steps. Taking the noisy image data as an input of the model can further enrich the input information of the model.
[0069] In an exemplary embodiment of the present disclosure, for step S130, the initial model performs fused attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise, including: the control network module performs feature extraction processing on the first image and the reference image respectively to obtain the first image feature and the reference image feature; the reference fusion module performs fused attention feature extraction on the first image feature, the reference image feature and the noisy image data to obtain the output image feature; the output image feature is decoded to obtain the predicted image noise.
[0070] Among them, the initial model can be a controllable diffusion model, and the control network module can be a controller in the initial model as a pre-trained T2I model. For example, the control network module can be implemented using a skeleton model based on ControlNet. The first image feature can be a feature obtained by performing feature extraction processing on the first image using the control network module. The reference image feature can be a feature obtained by performing feature extraction processing on the reference image using the control network module. The reference fusion module can be a fusion attention module composed of multiple attention modules. The output image feature can be the image feature output after the reference fusion module performs feature extraction and fusion processing on the input feature data. The decoding process can be a process of restoring the encoded image data to the original image format, and the decoding process can be the inverse process of the image encoding process. The predicted image noise can be the noise contained in the sample image predicted by the initial model.
[0071] In the second stage, a skeleton model based on ControlNet is used as the controller of the pre-trained T2I model. However, since the ControlNet model is designed for high-level control such as edges, depth, and posture, it is not competent for fine super-resolution tasks. Although the reference image obtained by the image activation module has a high resolution, it may introduce content that does not exist in the original image, resulting in reduced fidelity. Therefore, a reference fusion module is introduced in the upsampling process of the initial model to improve the quality and fidelity of the output by integrating all conditional image information.
[0072] refer to Figure 2 , Figure 2 : is a model framework diagram of an image super-resolution model according to an exemplary embodiment. The first image and the reference image are input into the initial model, and the image encoder in the initial model encodes the first image and the reference image respectively to obtain the encoded first image and the reference image. The encoded first image and the reference image are input into the control network module, and the control network module extracts features from the encoded first image and the reference image respectively to obtain the first image feature and the reference image feature. For example, the feature of the input first image is recorded as L, and the feature of the reference image obtained after the first stage is recorded as R.
[0073] In addition, the noisy image data is input into the initial model, and the U-Net network structure in the initial model extracts its features, and the extracted feature vector X is used as an input of the reference fusion module, so that the reference fusion module extracts the input image features by fusion attention features to obtain the output image features; then, the output image features will be output to the next layer of decoder, and the decoder decodes the output image features to obtain the predicted image noise. The reference fusion module is introduced on the basis of the control network model, so that the reference fusion module can integrate all conditional image information to improve the quality and fidelity of the output.
[0074] In an exemplary embodiment of the present disclosure, a reference fusion module performs fused attention feature extraction on the first image features, the reference image features and the noisy image data to obtain the output image features, including: determining the noisy image features corresponding to the noisy image data based on the self-attention module; inputting the first image features and the reference image features into the reference attention module to obtain the enhanced sample features after the first image and the reference image are fused; inputting the enhanced sample features and the noise image features into the first attention module to obtain the output image features.
[0075] Among them, the reference attention module can be a module for extracting features from the first image features and the reference image features. The self-attention module can be a module for extracting features from the noisy image data. The first attention module can be a module that takes the output result of the self-attention module and the output result of the reference attention module as input and performs feature extraction. The noise image feature can be a feature obtained after feature extraction processing is performed on the noisy image data. The enhanced sample feature can be a feature obtained after the reference attention module extracts and enhances the first image feature based on the reference image feature.
[0076] refer to Figure 3 , Figure 3 is a process diagram of feature aggregation based on a reference fusion module according to an exemplary embodiment. Figure 3 The reference fusion module in the example may include three modules: reference attention module, self-attention module and first attention module. For the first image feature L output by the first stage and the reference image feature R, according to Figure 3 It can be seen that the first image feature L and the reference image feature R are first input into the reference attention module, the reference attention module performs feature extraction on the above input, and the reference attention module outputs the enhanced sample feature after the first image and the reference image are fused.
[0077] The feature vector X output by the U-Net network structure is used as the input of the self-attention module, and the self-attention module outputs the noise image features corresponding to the noised image data; then, the output of the self-attention module (noise image features) and the output of the reference attention module (enhanced sample features) are used as the input of the first attention module, and the first attention module extracts features from the input information to obtain output image features. The reference fusion module fuses the information of the reference image and the input low-definition image through the attention mechanism, which can be used to control the generation of the diffusion model.
[0078] In an exemplary embodiment of the present disclosure, a first image feature and a reference image feature are input into a reference attention module to obtain an enhanced sample feature after the first image and the reference image are fused, including: the reference attention module performs attention feature extraction on the first image feature to obtain a query image feature; the reference attention module performs attention feature extraction on the reference image feature to obtain a key image feature and a value image feature; and normalization is performed based on the query image feature, the key image feature and the value image feature to obtain an enhanced sample feature.
[0079] The query image feature may be a query vector outputted by the reference attention module after extracting features from the first image feature based on the attention mechanism. The key image feature may be a key vector obtained by the reference attention module after extracting features from the reference image feature based on the attention mechanism. The value image feature may be a value vector outputted by the reference attention module after extracting features from the reference image feature based on the attention mechanism.
[0080] For the reference fusion module, the query, key, and value in the attention mechanism are Q, K, and V respectively. Figure 3 First, the first image feature L and the reference image feature R are input into the reference attention module, and the reference attention module extracts the attention feature of the first image feature L to obtain the query image feature, that is, the query image feature Q r Calculated from the first image feature L. The reference attention module extracts the attention feature of the reference image feature R to obtain the key image feature K r and value image feature V r , that is, the key image feature K r and value image feature V r Calculated from the reference image feature R.
[0081] Furthermore, the reference attention module is used to analyze the query image feature Q r , key image feature K r and value image feature V r After normalization, the output of the reference attention module is obtained, that is, the enhanced sample feature L′. The calculation process of the reference attention module is shown in Formula 1.
[0082]
[0083] Among them, L′ can represent the enhanced sample features, that is, the output of the reference attention module; Q r Can represent the query image features; K r Can represent key image features; V r It can represent the value image feature; d can represent a constant; Softmax() can represent a normalized exponential function. The reference fusion module fuses the information of the reference image and the input low-definition image through the attention mechanism, which can make full use of the information in the input image to control the generation of the diffusion model.
[0084] In an exemplary embodiment of the present disclosure, enhanced sample features and noise image features are input into a first attention module to obtain output image features, including: the first attention module performs attention feature extraction on the noise image features to obtain query enhancement features; the first attention module performs attention feature extraction on the enhanced sample features to obtain key enhancement features and value enhancement features; and normalization is performed based on the query enhancement features, key enhancement features, and value enhancement features to obtain output image features.
[0085] The query enhancement feature may be a query vector obtained after the first attention module extracts the noise image features based on the attention mechanism. The key enhancement feature may be a key vector obtained after the first attention module extracts the enhanced sample features based on the attention mechanism. The value enhancement feature may be a value vector output by the first attention module after extracting the enhanced sample features based on the attention mechanism. The first attention module is also called the low-definition attention module.
[0086] Continue to refer Figure 3 , the output of the reference attention module, i.e., the key image feature K r and value image feature V r As one of the inputs of the first attention module; at the same time, the output of the self-attention module, that is, the noise image feature X', is used as another input of the first attention module.
[0087] The first attention module extracts the attention feature of the noise image feature X' to obtain the query enhancement feature Q lr , that is, query enhancement feature Q lr It is calculated from the noise image feature X'; the first attention module extracts the attention feature of the enhanced sample feature L' to obtain the key enhanced feature K lr Enhanced feature V with value lr , that is, the key enhancement feature K lr Enhanced feature V with value lr Calculated by the enhanced sample feature L′.
[0088] Furthermore, the first attention module enhances the query feature Q lr , key enhancement feature K lr Enhanced feature V with value lr Normalize it to get the output of the first attention module, that is, the output image feature X out The calculation process of the first attention module is shown in Formula 2. Then the image feature X is output out It is output to the next layer decoder to obtain the final output. The reference fusion module can improve the quality and fidelity of the output image by integrating all conditional image information.
[0089]
[0090] Among them, X out It can represent the output image features, that is, the output of the first attention module; Q lr Can represent query enhancement features; K lr Can represent bond enhancement features; V lr It can represent value enhancement features; d can represent a constant; Softmax() can represent a normalized exponential function.
[0091] Furthermore, during the model training process, the backbone network ∈ with θ as the parameter θ Using the first image feature z lr and the feature z of the reference image r , predict the added noise. In particular, during the training process, the present disclosure sets the text prompt of the T2I model to be empty, denoted as In the inference phase, only some generalized text prompts are used, such as "clean, high resolution", instead of using the label model to extract annotations or design special annotations for each image. The mean square error function can be used as the loss function. Based on the above content, the training target of the model is determined, and the specific model training target is shown in Formula 3.
[0092]
[0093] in, Can express expectations; ||x|| 2 It can represent the L2 norm of x; ∈ can represent the actual image noise; can represent the predicted image noise.
[0094] During the training process, Figure 3 The model structure, all parameters of the pre-trained T2I model (Stable Diffusionv1.5) are frozen, and the training targets are the ControlNet model and the newly added reference fusion module.
[0095] Table 1: Comparison results of the present disclosure and other image super-resolution methods on DIV2K-val (DIV2K test set), DrealSR dataset and RealSR dataset.
[0096]
[0097]
[0098] In the performance comparison, the commonly used DIV2K dataset was selected to quantitatively evaluate the performance of this method. As shown in Table 1, this method achieved the best (State-of-the-Art, SOTA) results in most indicators. In addition, experiments were conducted on DIV2K and specific test sets to qualitatively evaluate the performance of this method. Both the subjective performance results and the image / video quality evaluation (Kuaishou Visual Quality, KVQ) indicators were significantly improved.
[0099] In summary, the training method of the image super-resolution model disclosed in the present invention obtains a sample image and a first image, determines a reference image corresponding to the first image; performs noise processing on the sample image based on the sample image data to obtain the noisy image data; obtains a pre-built initial model, and the initial model performs fusion attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise; determines the actual image noise contained in the noisy image data, and constructs a model loss function based on the predicted image noise and the actual image noise; trains the initial model based on the model loss function to obtain the image super-resolution model. On the one hand, in the process of model training, adding control to the input image can make full use of the prior knowledge of the diffusion model, and to a certain extent solves the problem that the related scheme performs poorly when the low-definition image has a high degree of degradation. On the other hand, by performing fusion feature extraction on the model input data, the information of the input image can be mined, and the quality of the model's generation of image super-resolution results can be improved. On the other hand, the reference fusion module fuses the information of the reference image and the input low-definition image through the attention mechanism, fully mining the information of the input low-definition image and improving the quality of the super-resolution results.
[0100] Figure 4 is a flowchart of an image super-resolution processing method according to an exemplary embodiment. Figure 4 As shown, the image super-resolution processing method can be used in a computer device. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.
[0101] In step S410, an image to be processed is acquired, and a reference image to be processed corresponding to the image to be processed is determined.
[0102] In an exemplary embodiment of the present disclosure, the image to be processed may be an input image to be subjected to super-resolution reconstruction processing. The reference image to be processed may be an image obtained after denoising and optimization processing is performed on the image to be processed.
[0103] Get the image to be processed. The image to be processed can be a low-resolution image input by the user, or other image containing noise. Figure 5 , Figure 5 This is a workflow diagram of an image super-resolution model according to an exemplary embodiment. After obtaining the image to be processed, the image activation model is used to perform denoising on the image to be processed to optimize the image to be processed, thereby reducing the degradation degree of the input low-definition image to be processed, and the optimized image is used as the reference image to be processed.
[0104] In step S420, an image super-resolution model trained according to the image super-resolution model training method is obtained.
[0105] In an exemplary embodiment of the present disclosure, a pre-trained image super-resolution model is obtained, and the image super-resolution model is trained according to a training method of the image super-resolution model. The training method of the image super-resolution model has been explained in detail above, and the present disclosure will not repeat it again.
[0106] In step S430, the image to be processed and the reference image to be processed are input into the image super-resolution model, and the image super-resolution model outputs a super-resolution reconstructed image corresponding to the image to be processed.
[0107] In an exemplary embodiment of the present disclosure, the super-resolution reconstructed image may be a high-resolution image obtained by performing super-resolution reconstruction on the image to be processed.
[0108] After obtaining the image super-resolution model, the image to be processed and the reference image to be processed are used as the input of the image super-resolution model. The image super-resolution model performs super-resolution reconstruction on the input data, and finally obtains the super-resolution reconstructed image corresponding to the image to be processed. The image super-resolution model uses the reference image of the image to be processed as one of its inputs, which can control the generation process more accurately and reduce the impact of the quality degradation of the input image on the generation process, so as to obtain a higher quality image super-resolution result. Figure 6 , Figure 6 FIG. 1 is a comparison diagram of an input image to be processed and an output super-resolution reconstructed image according to an exemplary embodiment. Figure 6 It can be seen that the super-reconstructed image has higher resolution and clarity than the image to be processed.
[0109] According to the image super-resolution processing method in this example embodiment, on the one hand, the low-definition image to be processed and the corresponding reference image are used as the input of the model, which can make full use of the information in the input image and the prior knowledge of the diffusion model to more accurately control the generation process. On the other hand, using the reference image as input can reduce the impact of the quality degradation of the input image on the generation process, so as to obtain a higher quality image super-resolution result.
[0110] Figure 7FIG. 1 is a block diagram of a training device for an image super-resolution model according to an exemplary embodiment. Figure 7 The image super-resolution model training device 700 includes: a reference image determination module 710, an image noise addition module 720, a noise prediction module 730, a loss function determination module 740 and a model training module 750.
[0111] Specifically, the reference image determination module 710 is used to obtain the sample image and the first image, and determine the reference image corresponding to the first image; the image denoising module 720 is used to perform denoising on the sample image based on the sample image data to obtain the noisy image data; the noise prediction module 730 is used to obtain a pre-constructed initial model, and the initial model performs fusion attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise; the loss function determination module 740 is used to determine the actual image noise contained in the noisy image data, and construct a model loss function based on the predicted image noise and the actual image noise; the model training module 750 is used to train the initial model based on the model loss function to obtain an image super-resolution model.
[0112] In an exemplary embodiment of the present disclosure, the reference image determination module 710 includes a reference image determination unit, which is used to: obtain a pre-constructed image activation model; and use the image activation model to perform denoising and optimization processing on the first image to obtain a reference image.
[0113] In an exemplary embodiment of the present disclosure, the image denoising module 720 includes an image denoising unit, which is used to: encode the sample image to obtain encoded sample image data; obtain pre-configured random noise and diffusion steps, and add random noise to the sample image data based on the diffusion steps to obtain noisy image data.
[0114] In an exemplary embodiment of the present disclosure, the initial model includes a control network module and a reference fusion module, and the noise prediction module 730 includes a noise prediction unit, which is used to: the control network module performs feature extraction processing on the first image and the reference image respectively to obtain first image features and reference image features; the reference fusion module performs fusion attention feature extraction on the first image features, the reference image features and the noisy image data to obtain output image features; the output image features are decoded to obtain predicted image noise.
[0115] In an exemplary embodiment of the present disclosure, a reference fusion module includes a reference attention module, a self-attention module and a first attention module; the noise prediction unit includes an output feature determination unit, which is used to: determine the noise image features corresponding to the noisy image data based on the self-attention module; input the first image features and the reference image features into the reference attention module to obtain enhanced sample features after the first image and the reference image are fused; input the enhanced sample features and the noise image features into the first attention module to obtain output image features.
[0116] In an exemplary embodiment of the present disclosure, the output feature determination unit includes an enhanced feature determination subunit, which is used to: perform attention feature extraction on the first image feature by a reference attention module to obtain a query image feature; perform attention feature extraction on the reference image feature by the reference attention module to obtain a key image feature and a value image feature; perform normalization processing based on the query image feature, the key image feature and the value image feature to obtain an enhanced sample feature.
[0117] In an exemplary embodiment of the present disclosure, the output feature determination unit includes an output feature determination sub-unit, which is used to: perform attention feature extraction on noise image features by the first attention module to obtain query enhancement features; perform attention feature extraction on enhanced sample features by the first attention module to obtain key enhancement features and value enhancement features; perform normalization processing based on the query enhancement features, key enhancement features and value enhancement features to obtain output image features.
[0118] Figure 8 FIG. 1 is a block diagram of an image super-resolution processing device according to an exemplary embodiment. Figure 8 The image super-resolution processing device 800 includes: an image determination module 810, a model acquisition module 820 and a super-resolution reconstruction module 830.
[0119] Specifically, the image determination module 810 is used to obtain the image to be processed and determine the reference image to be processed corresponding to the image to be processed; the model acquisition module 820 is used to obtain the image super-resolution model trained according to the above-mentioned image super-resolution model training method; the super-resolution reconstruction module 830 is used to input the image to be processed and the reference image to be processed into the image super-resolution model, and the image super-resolution model outputs the super-resolution reconstructed image corresponding to the image to be processed.
[0120] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0121] Reference below Fig. 9 hereinafter describes an electronic device 900 according to such an embodiment of the present disclosure. Fig. 9The electronic device 900 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0122] like Fig. 9 As shown, the electronic device 900 is in the form of a general computing device. The components of the electronic device 900 may include, but are not limited to: the at least one processing unit 910, the at least one storage unit 920, a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910), and a display unit 940.
[0123] The storage unit stores program codes, which can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0124] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 921 and / or a cache memory unit 922 , and may further include a read-only memory unit (ROM) 923 .
[0125] The storage unit 920 may include a program / utility 924 having a set (at least one) of program modules 925, such program modules 925 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0126] Bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0127] The electronic device 900 may also communicate with one or more external devices 970 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 950. Furthermore, the electronic device 900 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0128] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the instructions can be executed by a processor of a device to complete the above-mentioned image super-resolution model training method and image super-resolution processing method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0129] In an exemplary embodiment, a computer program product is also provided, including a computer program, which, when executed by a processor, implements any one of the above-mentioned image super-resolution model training methods or any one of the above-mentioned image super-resolution processing methods.
[0130] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0131] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A training method for an image super-resolution model, characterized in that: include: Acquire a sample image and a first image, and determine a reference image corresponding to the first image; Performing noise processing on the sample image based on the sample image data to obtain noisy image data; Acquire a pre-built initial model, and use the initial model to extract fusion attention features on the first image, the reference image, and the noisy image data to determine predicted image noise; Determining actual image noise contained in the noisy image data, and constructing a model loss function based on the predicted image noise and the actual image noise; The initial model is trained based on the model loss function to obtain an image super-resolution model.
2. The method according to claim 1, characterized in that The determining a reference image corresponding to the first image includes: Get a pre-built image activation model; The image activation model is used to perform denoising and optimization processing on the first image to obtain the reference image.
3. The method according to claim 1, characterized in that The step of performing noise processing on the sample image based on the sample image data to obtain the noise-added image data comprises: Performing encoding processing on the sample image to obtain encoded sample image data; Preconfigured random noise and diffusion steps are obtained, and the random noise is added to the sample image data based on the diffusion steps to obtain the noisy image data.
4. The method according to claim 1, characterized in that: The initial model includes a control network module and a reference fusion module. The initial model extracts fusion attention features from the first image, the reference image and the noisy image data to determine predicted image noise, including: The control network module performs feature extraction processing on the first image and the reference image respectively to obtain first image features and reference image features; The reference fusion module performs fusion attention feature extraction on the first image feature, the reference image feature and the noisy image data to obtain an output image feature; The output image feature is decoded to obtain the predicted image noise.
5. The method according to claim 4, characterized in that The reference fusion module includes a reference attention module, a self-attention module and a first attention module; The reference fusion module extracts the first image feature, the reference image feature and the noisy image data by fusing attention features to obtain output image features, including: Determine, based on the self-attention module, a noise image feature corresponding to the noisy image data; Inputting the first image features and the reference image features into the reference attention module to obtain enhanced sample features after the first image and the reference image are fused; The enhanced sample features and the noise image features are input into the first attention module to obtain the output image features.
6. The method according to claim 5, characterized in that The step of inputting the first image feature and the reference image feature into the reference attention module to obtain an enhanced sample feature after the first image and the reference image are fused, comprises: The reference attention module performs attention feature extraction on the first image feature to obtain a query image feature; The reference attention module extracts attention features from the reference image features to obtain key image features and value image features; Normalization processing is performed according to the query image feature, the key image feature and the value image feature to obtain the enhanced sample feature.
7. The method according to claim 5, characterized in that The step of inputting the enhanced sample feature and the noise image feature into the first attention module to obtain the output image feature comprises: The first attention module performs attention feature extraction on the noise image feature to obtain a query enhancement feature; The first attention module performs attention feature extraction on the enhanced sample features to obtain key enhanced features and value enhanced features; Normalization processing is performed according to the query enhanced feature, the key enhanced feature and the value enhanced feature to obtain the output image feature.
8. An image super-resolution processing method, characterized in that: include: Acquire an image to be processed, and determine a reference image to be processed corresponding to the image to be processed; Obtain an image super-resolution model trained according to the image super-resolution model training method according to any one of claims 1 to 7; The image to be processed and the reference image to be processed are input into the image super-resolution model, and the image super-resolution model outputs a super-resolution reconstructed image corresponding to the image to be processed.
9. A training device for an image super-resolution model, characterized in that: include: A reference image determination module, used to obtain a sample image and a first image, and determine a reference image corresponding to the first image; An image noise adding module, used for performing noise adding processing on the sample image based on the sample image data to obtain the noise added image data; A noise prediction module is used to obtain a pre-built initial model, and the initial model performs fusion attention feature extraction on the first image, the reference image and the noisy image data to determine the predicted image noise; A loss function determination module, used to determine the actual image noise contained in the noisy image data, and construct a model loss function based on the predicted image noise and the actual image noise; The model training module is used to train the initial model based on the model loss function to obtain an image super-resolution model.
10. An image super-resolution processing device, characterized in that: include: An image determination module is used to obtain an image to be processed and determine a reference image to be processed corresponding to the image to be processed; A model acquisition module, used to acquire an image super-resolution model trained according to the image super-resolution model training method according to any one of claims 1 to 7; The super-resolution reconstruction module is used to input the image to be processed and the reference image to be processed into the image super-resolution model, and the image super-resolution model outputs the super-resolution reconstructed image corresponding to the image to be processed.
11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image super-resolution model training method as described in any one of claims 1 to 7, or to implement the image super-resolution processing method as described in claim 8.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the image super-resolution model as described in any one of claims 1 to 7, or implements the image super-resolution processing method as described in claim 8.