Image super-resolution reconstruction method and system based on hybrid experts and stable diffusion
Through the image super-resolution reconstruction method based on hybrid expert and stable diffusion, combined with the degradation perception decoder and multi-scale control conditions, the problem of inaccurate image degradation simulation and smooth reconstruction results in the prior art is solved, and high-quality super-resolution reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510024781.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Existing image super-resolution reconstruction methods are difficult to accurately simulate the real-world image degradation process, and the reconstruction results are too smooth to effectively capture complex texture information.
Using a hybrid expert and stable diffusion-based image super-resolution reconstruction method, the reconstruction hidden layer characterization is decoded into the pixel space through an image quality-driven degradation perception decoder, combining degradation embedding, multi-scale control conditions and spatial control conditions to improve the predicted noise accuracy of the denoising backbone network.
Realize the real and clear super-resolution reconstruction effect, restore clear and natural local details, and improve the visual experience of the reconstructed image.
Smart Images

Figure CN119444578B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image super-resolution reconstruction, and in particular relates to an image super-resolution reconstruction method and system based on mixed experts and stable diffusion. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The goal of image super-resolution (SR) is to recover a high-resolution (HR) image with more details from a low-resolution (LR) image. This task is widely used in medical image analysis, satellite remote sensing image processing, video enhancement, and security monitoring. Early super-resolution methods were mainly based on signal processing and statistical modeling, and usually assumed that the image had certain prior characteristics. Representative methods include interpolation methods, reconstruction-based methods, and dictionary learning-based methods. Among them, the interpolation method is the simplest to calculate, but it is easy to introduce artifacts and produce jagged edges, and it is impossible to recover high-frequency details; the reconstruction-based method can restore details to a certain extent, but it is sensitive to parameter settings and it is difficult to effectively capture complex texture information; and the dictionary learning-based method has limited feature expression capabilities and is difficult to process large-scale complex data.
[0004] In order to solve the limitations of traditional image super-resolution reconstruction methods, super-resolution networks based on deep learning have made significant progress in recent years. Early super-resolution network models (such as SRCNN, VDSR, RCAN, etc.) use LR and high-resolution real images (Ground-Truth, GT) to achieve super-resolution through end-to-end training. However, since such models usually assume a known degradation operator and only use pixel-level loss for optimization, it is difficult to accurately simulate the image degradation process in the real world and the reconstruction results are too smooth. With the widespread application of Generative Adversarial Network (GAN) in the field of image generation, it has performed well in improving the clarity and realism of super-resolution images. GAN-based super-resolution models (such as SRGAN, ESRGAN, Real-ESRGAN, etc.) usually use end-to-end networks as generators and design discriminators for adversarial training. However, due to the characteristics of the GAN network, the reconstruction results are often accompanied by unrealistic textures, which affects the visual experience. Recently, with the development of large-scale text-to-image diffusion models, such models have accumulated rich prior knowledge through pre-training on massive datasets, providing a potential efficient solution for super-resolution tasks. Many works (such as DiffBIR, PASD, SeeSR, SUPIR, etc.) use low-resolution images as conditional controls by training control networks (Controlnet) to guide the diffusion model to gradually generate high-resolution images (HR) from noise. Although the diffusion model has made good progress in visual quality, it has not fully considered the impact of different degradation types on image reconstruction. At the same time, the prompt condition mainly relies on text input, ignoring the rich prompt information contained in the image itself. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides an image super-resolution reconstruction method and system based on hybrid experts and stable diffusion, which decodes the reconstructed latent layer representation into pixel space through an image quality-driven degradation-aware decoder, and can obtain a true and clear super-resolution reconstruction effect.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] A first aspect of the present invention provides an image super-resolution reconstruction method based on mixed experts and stable diffusion.
[0008] In one or more embodiments, a method for image super-resolution reconstruction based on mixed experts and stable diffusion is provided, comprising:
[0009] Get a low-resolution image and extract its degradation embedding and feature embedding;
[0010] The low-resolution image is encoded using degradation embedding and a trained degradation-aware conditional image encoder, and then passed through a trained hybrid expert conditional control network to obtain a multi-scale control condition; the feature embedding is mapped from the image prompt space to the text prompt space using a trained hybrid expert image inversion adapter to obtain a spatial control condition;
[0011] The low-resolution image is encoded by a trained variational autoencoder, and then fused with the noise sampled by the noise scheduler to generate a noisy feature hidden layer space representation; or pure Gaussian noise is directly sampled to generate a noisy feature hidden layer space representation;
[0012] In each sampling time step, the latent spatial representation of the noisy feature is used as the input of the trained denoising backbone network, and the noise predicted by the denoising backbone network at each sampling time step is obtained by combining the multi-scale control condition and the spatial control condition;
[0013] After multiple sampling time steps, the noise predicted by the denoising backbone network at the current time step is continuously subtracted from the input of the denoising backbone network at the current sampling time step, and then the denoising result is used as the input of the denoising backbone network at the next time step. After multiple sampling time steps are completed, the latent spatial expression of the reconstructed image is obtained;
[0014] By using the trained image quality-driven degradation-aware decoder and combining the intermediate features and degradation embedding of low-resolution image encoding with the variational autoencoder, the latent spatial expression of the reconstructed image is decoded into the image pixel space to obtain the image super-resolution reconstruction result.
[0015] As an implementation method, in each sampling time step, the hidden spatial representation of the noisy features is used as the input of the denoising backbone network, the multi-scale control conditions are added to the down-sampled part of the denoising backbone network, and are spliced and mapped with the features of the up-sampled part in the channel dimension; the spatial control conditions are cross-attention calculated with the intermediate features of the denoising backbone network.
[0016] As an implementation mode, the training process of the degradation-perceptual conditional image encoder is:
[0017] The low-resolution images in the training samples are used as the input of the degradation-aware conditional image encoder;
[0018] Each basic module of the degradation-aware conditional image encoder embeds the degradation of the low-resolution image into the intermediate features through adaptive modulation, and then passes the intermediate features to the next basic module;
[0019] The hidden layer representation output by the degradation-perceptual conditional image encoder is decoded by a decoder composed of basic modules, and the calculated loss of the real image corresponding to the decoded image and the low-resolution image is used as the training objective function to optimize the parameters of the degradation-perceptual conditional image encoder.
[0020] As an implementation method, the training objective function of the degradation-perceptual conditional image encoder is:
[0021]
[0022] in, is the training objective function of the degradation-aware conditional image encoder; are the height, width and number of channels of the image respectively; For the real image The position pixel Channel values; To decode the image No. The position pixel channel value.
[0023] As an implementation method, the expression of the latent space representation of the noise-added feature is:
[0024] ;
[0025] or ;
[0026] in, is the hidden space representation of the noise feature, represents the weight factor; represents the pure Gaussian noise sampled from the Gaussian distribution by the noise scheduler; is the feature hidden space representation.
[0027] As an implementation method, the training process of the hybrid expert conditional control network and the hybrid expert image inversion adapter is:
[0028] Extract degradation embedding and feature embedding for low-resolution images in training samples;
[0029] Combining the low-resolution images and their degradation embeddings in the training samples, the trained degradation-aware conditional image encoder is used to encode the low-resolution images while preliminarily removing the degradation.
[0030] The encoded features are passed through a hybrid expert conditional control network to generate multi-scale control conditions;
[0031] The feature embedding is mapped from the image cue space to the text cue space through a hybrid expert image inversion adapter to obtain the spatial control condition;
[0032] The high-resolution real image corresponding to the low-resolution image in the training sample is encoded by a variational autoencoder, and then fused with the noise sampled by the noise scheduler to generate a noisy feature hidden layer space representation;
[0033] The hidden spatial representation of the noise-added features is used as the input of the denoising backbone network. The multi-scale control conditions are added to the downsampled part of the denoising backbone network, and are concatenated and mapped with the features of the upsampled part in the channel dimension. The spatial control conditions are cross-attentioned with the intermediate features of the denoising backbone network to obtain the noise predicted by the denoising backbone network.
[0034] The parameters of the hybrid expert conditional control network and the hybrid expert image inversion adapter are optimized based on the computational loss of the sampled original noise and the predicted noise.
[0035] As an implementation, the training process of the image quality driven degradation-aware decoder includes:
[0036] The variational autoencoder is used to encode the real image corresponding to the low-resolution image in the training sample to generate the decoder input hidden layer features;
[0037] Use variational autoencoders to encode low-resolution images in training samples and obtain intermediate features of each layer;
[0038] The hidden features of the real image are used as the input of the image quality driven degradation-aware decoder. Each layer fuses the low-resolution image intermediate features and degradation embedding of the corresponding layer of the encoder to decode the hidden features into pixel space.
[0039] Freeze the discriminative model parameters and optimize the parameters of the decoder corresponding to the variational autoencoder based on the calculated pixel loss, perceptual loss, and GAN loss of the decoded image and the real image;
[0040] Freeze the parameters of the variational autoencoder and its corresponding decoder, and optimize the parameters of the discriminative model based on the GAN loss calculated for the decoded image and the real image respectively;
[0041] Repeat the above steps of freezing the discriminant model parameters and freezing the parameters of the variational autoencoder and its corresponding decoder until the decoder and the discriminant model converge.
[0042] A second aspect of the present invention provides an image super-resolution reconstruction system based on hybrid experts and stable diffusion.
[0043] In one or more embodiments, an image super-resolution reconstruction system based on hybrid experts and stable diffusion includes:
[0044] An embedding extraction module, which is used to acquire a low-resolution image and extract its degradation embedding and feature embedding;
[0045] A control condition generation module is used to encode the low-resolution image using degradation embedding and a trained degradation-aware conditional image encoder, and then obtain multi-scale control conditions through a trained hybrid expert conditional control network; and map feature embedding from an image prompt space to a text prompt space using a trained hybrid expert image inversion adapter to obtain a spatial control condition;
[0046] A feature noisy module is used to encode the low-resolution image through a trained variational autoencoder, and then fuse it with the noise sampled by the noise scheduler to generate a noisy feature latent space representation; or directly sample pure Gaussian noise to generate a noisy feature latent space representation;
[0047] The noise prediction module is used to use the hidden spatial representation of the noise-added feature as the input of the trained denoising backbone network in each sampling time step, and combine the multi-scale control condition and the spatial control condition to obtain the noise predicted by the denoising backbone network at each sampling time step;
[0048] The hidden space expression module is used to continuously subtract the noise predicted by the denoising backbone network at the current time step from the input of the denoising backbone network at the current sampling time step through multiple sampling time steps, and then use the denoising result as the input of the denoising backbone network at the next time step. After multiple sampling time steps are completed, the hidden space expression of the reconstructed image is obtained;
[0049] The super-resolution image reconstruction module is used to utilize the trained image quality-driven degradation-aware decoder, combined with the intermediate features and degradation embedding of the variational autoencoder for low-resolution image encoding, to decode the latent spatial expression of the reconstructed image into the image pixel space to obtain the image super-resolution reconstruction result.
[0050] A third aspect of the present invention provides a computer-readable storage medium.
[0051] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the image super-resolution reconstruction method based on mixed experts and stable diffusion as described above.
[0052] A fourth aspect of the present invention provides an electronic device.
[0053] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the image super-resolution reconstruction method based on mixed experts and stable diffusion as described above are implemented.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] The present invention is applicable to super-resolution reconstruction of common images. Starting from the perspective of hybrid expert model and stable diffusion, the fine-grained image processing capability of the hybrid expert model and the reconstruction capability of stable diffusion for image detail texture are utilized. First, degradation embedding and feature embedding extracted from low-resolution images are used, and then degradation embedding, trained degradation-aware conditional image encoder and hybrid expert conditional control network are used to form multi-scale control conditions, thereby realizing fine-grained processing of degradation embedding. The hybrid expert image inversion adapter is used to combine feature embedding to obtain spatial control conditions, and richer semantic representations are mined from low-resolution images, thereby improving the accuracy of noise prediction by a denoising backbone network. Finally, the reconstructed hidden layer representation is decoded into pixel space through an image quality-driven degradation-aware decoder, thereby further improving the texture of the reconstruction result, so that the entire super-resolution reconstruction method is conducive to restoring clear and natural local details, thereby improving the visual experience of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0057] Figure 1 is a schematic diagram of an image super-resolution reconstruction process based on hybrid experts and stable diffusion according to an embodiment of the present invention;
[0058] Figure 2 is a structural diagram of a degraded perceptual condition image encoder according to an embodiment of the present invention;
[0059] Figure 3 2D degradation-aware residual block and a 2D residual block structure diagram according to an embodiment of the present invention;
[0060] Figure 4 is a structural diagram of a Top-P hybrid expert module according to an embodiment of the present invention;
[0061] Figure 5 is a structural diagram of a hybrid expert image inversion adapter according to an embodiment of the present invention;
[0062] Figure 6 is a structural diagram of a degradation-aware decoder driven by image quality according to an embodiment of the present invention;
[0063] Figure 7 is a flow chart of an image super-resolution reconstruction method based on hybrid experts and stable diffusion according to an embodiment of the present invention;
[0064] Figure 8It is a schematic diagram of the structure of an image super-resolution reconstruction system based on hybrid experts and stable diffusion according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0066] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0067] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0068] Embodiment 1
[0069] Combination Figure 1 and Figure 7 , the image super-resolution reconstruction method based on mixed experts and stable diffusion in this embodiment may include:
[0070] S701: Acquire low-resolution image And extract its degenerate embedding and feature embedding.
[0071] Extraction model via pre-trained degradation and feature embedding right Extracting Degenerate Embeddings and feature embedding ; The process is as follows: .
[0072] S702: Using Degenerate Embedding and trained degradation-aware conditional image encoder For the low-resolution image Encoding, and then trained hybrid expert conditional control network , and obtain the multi-scale control condition ; Using the trained mixture of experts image inversion adapter Embedding features The spatial control condition is obtained by mapping from the image prompt space to the text prompt space.
[0073] In this embodiment, the training process of the degradation-perceptual conditional image encoder is:
[0074] Step A: The low-resolution images in the training samples are used as the input of the degradation-aware conditional image encoder; the training samples are composed of high-resolution-low-resolution image pairs. The construction process of the high-resolution-low-resolution image data pairs in the training samples is:
[0075] For existing high-definition real-world image data (i.e. high-resolution images) , a second-order degradation pipeline simulating the real-world degradation process is used to obtain the corresponding low-resolution image , that is, we get a high-resolution-low-resolution image pair .
[0076] Among them, each order of degradation includes downsampling , upsampling , low pass filter ,noise and JPEG compression etc., and then randomly crop the high-resolution image to size, the corresponding low-resolution image size is , and interpolate to Resolution .
[0077] Among them, low-resolution images The expression is:
[0078] .
[0079] Step B: Each basic module of the degradation-aware conditional image encoder embeds the degradation of the low-resolution image into the intermediate features through adaptive modulation, and then passes the intermediate features to the next basic module;
[0080] like Figure 2 As shown in Figure 1, the degradation-aware conditional image encoder includes an encoder side and a decoder side. The encoder side first uses a pre-trained degradation and feature embedding extraction model extract The corresponding degenerate embedding ;Then After four consecutive residual blocks for feature extraction and downsampling, the intermediate representation is finally obtained. Each residual block contains two 2D degradation-aware residual blocks. Figure 3 (a) in the figure is a 2D degradation-aware residual block, which mainly includes group normalization , linear layer , activation function And convolutional layers The 2D degradation-aware residual block will Or intermediate features as input, combined with degenerate embedding To preliminarily remove the degradation in the image and provide more accurate feature input for the subsequent generation of Controlnet control conditions. This process is shown in the following formula:
[0081]
[0082]
[0083]
[0084]
[0085]
[0086] in, Represents the intermediate features, represents the modulation parameters for modulating the degradation features into the 2D degradation-aware residual block.
[0087] Degradation-aware conditional image encoder The degradation embedding of the fused image is achieved by adaptive modulation , this modulation method maps the degradation embedding into the modulation parameter through a linear layer , and then scale the normalized input features so that the encoder can preliminarily remove the degradation in the image and be robust.
[0088] Step C: The hidden layer representation output by the degraded perceived conditional image encoder is decoded by a decoder composed of basic modules, and the loss is calculated based on the decoded image and the real image corresponding to the low-resolution image as the training objective function to optimize the parameters of the degraded perceived conditional image encoder.
[0089] The decoder mainly uses 2D residual blocks with degraded embedding removed for feature mapping and upsampling. Figure 3 (b) in the figure gives the structure of the 2D residual block, where each 2D residual block includes group normalization , activation function And convolutional layers The decoder converts the intermediate representation output by the encoder into Mapping back to the feature space, we get the first stage reconstructed image Finally, the reconstructed image and real images Calculate the loss to optimize the entire network model. This process can be expressed as the following formula.
[0090]
[0091]
[0092] in, Represents intermediate features.
[0093] The training objective function of the degradation-perceptual-condition image encoder of this embodiment is:
[0094]
[0095] in, is the training objective function of the degradation-aware conditional image encoder; are the height, width and number of channels of the image respectively; For the real image The position pixel Channel values; To decode the image No. The position pixel channel value.
[0096] S703: The low-resolution image Through the trained variational autoencoder Encode and then sample the noise with the noise scheduler Fusion, or direct sampling of pure Gaussian noise, to generate a hidden spatial representation of the noise-added feature .
[0097] Specifically, the expression of the hidden space representation of the noise-added feature is:
[0098] or
[0099] in, is the hidden space representation of the noise feature, represents the weight factor, represents the noise sampled from Gaussian distribution by the noise scheduler; is the feature hidden space representation.
[0100] S704: In each sampling time step, the noise feature hidden space representation As the input of the trained denoising backbone network, combined with multi-scale control conditions And spatial control conditions, the noise predicted by each time step of the denoising backbone network is obtained .
[0101] Specifically, in each sampling time step, the noise feature hidden space representation is As the input of the denoising backbone network, the multi-scale control condition It is added with the downsampled part of the denoising backbone network, and concatenated and mapped with the features of the upsampled part in the channel dimension; the spatial control condition is cross-attended with the intermediate features of the denoising backbone network.
[0102] Degenerate embedding For Degradation-Aware Conditional Image Encoder right Encode and get the intermediate representation , and through a hybrid expert conditional control network Output multi-scale control conditions ,here yes Each layer outputs features The specific process is:
[0103]
[0104]
[0105] Among them, the forward transmission process of the Top-P mixed expert module is:
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116] In the formula, represents the weight factor of each expert, represents the intermediate input features, represents the expert mapping matrix, represents the noisy expert mapping matrix, Representative The intermediate features of experts are output as , It is the output of the Top-P mixed expert module. represents the Top-P selection strategy, Respectively represent function, function, represents the standard normally distributed noise, represents the layer normalization operation, represents the convolution operation, represents the activation function, represent Convolution, whose output is .
[0117] Hybrid Expert Conditional Control Network of Embodiment A Top-P hybrid expert module consisting of multiple expert models is added after each basic module. Figure 4 As shown in the figure, the Top-P hybrid expert module includes: Top-P degradation-aware distribution routing, a set of experts for handling different types of degradation, and a regular expert. The Top-P hybrid expert module is designed to achieve fine-grained targeted degradation repair and avoid interference between different types of degradation. At the same time, the Top-P dynamically activated expert enables the network to flexibly select the most suitable expert model according to the characteristics of the input sample, reducing the computational overhead while ensuring performance.
[0118] This embodiment uses the Top-P hybrid expert module to perform fine-grained feature processing on the input control features. At the same time, the Top-P selection strategy ensures that different numbers of expert models are dynamically selected according to the degradation types actually contained in the image, unlike the Top-K selection strategy, which selects a fixed number of experts each time and limits the degradation processing capability. At the same time, compared with direct text input, mining semantic information from images has higher information density and cross-language expression capabilities, which is conducive to fine-grained conditional control of the stable diffusion model and ensures the semantic consistency of the reconstruction result and the real image.
[0119] Feature Embedding Hybrid Expert Image Inversion Adapter To perform feature space mapping, map feature embedding from image space to text space, and obtain control conditions , to motivate the text-to-image prior of the stable diffusion model. Figure 5 As shown, the hybrid expert image inversion adapter It is composed of Top-K distribution routing and several expert models. The several expert models are composed of multi-layer perceptrons to form an expert set. This embodiment dynamically distributes input to a diverse expert set through Top-K routing, thereby improving computing efficiency and task adaptability.
[0120] The Hybrid-of-Experts Image Inversion Adapter aims to achieve a fine-grained mapping from image cue space to text cue space, motivating the stable diffusion model of text-to-image priors. The process is shown in the following formula.
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] in, represents the weight factor of each expert, represents the expert mapping matrix, represents the noisy expert mapping matrix, Representative The intermediate features of experts are output as , is the output of the Top-P hybrid expert module; represents the Top-K selection strategy, Respectively represent function, function; represents the standard normally distributed noise, represents the layer normalization operation, represents a linear layer, is the activation function, Represents a dimension expansion operation.
[0127] In the specific implementation process, the training process of the hybrid expert conditional control network and the hybrid expert image inversion adapter is:
[0128] Step a: extract degradation embedding and feature embedding from low-resolution images in training samples;
[0129] Step b: Combining the low-resolution image and its degradation embedding in the training sample, the trained degradation-aware conditional image encoder is used to encode the low-resolution image while preliminarily removing the degradation;
[0130] Step c: Pass the encoded features through a hybrid expert conditional control network to generate multi-scale control conditions;
[0131] Step d: Map the feature embedding from the image cue space to the text cue space through a hybrid expert image inversion adapter to obtain the spatial control condition;
[0132] Step e: Encode the real image corresponding to the high-resolution real image in the training sample through a variational autoencoder, and then fuse it with the noise sampled by the noise scheduler to generate a noisy feature hidden layer space representation;
[0133] Step f: The hidden spatial representation of the noise-added feature is used as the input of the denoising backbone network, the multi-scale control condition is added to the down-sampled part of the denoising backbone network, and the multi-scale control condition is concatenated and mapped with the features of the up-sampled part in the channel dimension; the spatial control condition is cross-attended with the intermediate features of the denoising backbone network to obtain the noise predicted by the denoising backbone network;
[0134] Step g: Optimize the parameters of the hybrid expert conditional control network and the hybrid expert image inversion adapter based on the computational loss of the sampled original noise and predicted noise.
[0135] Combined with the current sampling time step , and finally obtain the noise output predicted by the denoising backbone network . Predicted noise output and the original sampling noise Calculate the loss to optimize the trainable parameters of the entire model. The process is:
[0136]
[0137] Training objective function of the hybrid expert conditional control network and the hybrid expert image inversion adapter for:
[0138] ;
[0139] in, is the sampling noise; To denoise the prediction noise of the backbone network; It is the hidden space representation of the noise-added feature; It is the spatial control condition; It is a multi-scale control condition; is the sampling time step; It is the denoising backbone network.
[0140] In this embodiment, the degradation-aware conditional image encoder, the variational autoencoder, the hybrid expert conditional control network, the hybrid expert image inversion adapter and the denoising backbone network constitute a stable diffusion model.
[0141] S705: After multiple sampling time steps, the noise predicted by the denoising backbone network at the current time step is continuously subtracted from the input of the denoising backbone network at the current sampling time step, and then the result is used as the input of the denoising backbone network at the next time step. After multiple sampling time steps are completed, the latent spatial expression of the reconstructed image is obtained. .
[0142] S706: Using the trained image quality driven degradation-aware decoder, combined with the intermediate features and degradation embedding of the variational autoencoder for low-resolution image encoding, the latent space representation of the reconstructed image is Decode to image pixel space to obtain image super-resolution reconstruction results.
[0143] Specifically, combined with the variational autoencoder for Encoded intermediate features and degenerate embedding , decode the hidden layer representation into pixel space to obtain the final super-resolution reconstructed image .
[0144] in, Figure 6 The image quality driven degradation-aware decoder structure of the embodiment of the present invention is given. Figure 6 , the training process of the image quality driven degradation-aware decoder includes:
[0145] Step 1: Use the variational autoencoder to encode the real image corresponding to the high-resolution real image in the training sample to generate the decoder input hidden layer features;
[0146] Step 2: Use the variational autoencoder to encode the low-resolution images in the training samples to obtain the intermediate features of each layer;
[0147] Step 3: The image quality driven degradation-aware decoder takes the real image hidden layer features as input, fuses the low-resolution image intermediate features and degradation embedding of the corresponding layer of the encoder at each layer, and decodes the hidden layer features into pixel space;
[0148] Step 4: Freeze the discriminative model parameters and optimize the parameters of the decoder corresponding to the variational autoencoder based on the calculated pixel loss, perceptual loss, and GAN loss of the decoded image and the real image;
[0149] Step 5: Freeze the parameters of the variational autoencoder and its corresponding decoder, and optimize the parameters of the discriminant model (such as the U-Net discriminant model) based on the GAN loss calculated for the decoded image and the real image respectively;
[0150] Step 6: Repeat the above steps of freezing the discriminant model parameters and freezing the parameters of the variational autoencoder and its corresponding decoder (i.e., steps 4 and 5) until the decoder and the discriminant model converge.
[0151] Specifically, combined Figure 6 In the process of training the degradation-aware decoder combined with the encoder, the variational autoencoder is first Encoding real images , generated as a decoder Input hidden features ;Then Continue coding , get the intermediate features of each layer ; The decoder converts the hidden layer features As input, fuse the intermediate features of the corresponding layer of each layer of the fusion encoder and degenerate embedding , decoding hidden layer features into pixel space .Decoder Degenerate embedding is added based on the original 2D residual block The fusion module is similar to the encoder, which embeds the degradation into the intermediate features through adaptive modulation. The process is shown in the following formula:
[0152]
[0153]
[0154]
[0155]
[0156]
[0157] in, is the intermediate feature, is the output feature, For channel splicing operations, represents the convolution operation, represents a linear layer, represents the activation function, Represents a 2D residual block.
[0158] Freeze the discriminant model Parameters, decoded image and real images Calculating pixel loss , Perceptual Loss and GAN loss , the weighted loss above is Optimize decoder parameters; freeze variational autoencoder and decoder parameters, the decoded image and the real image respectively calculate the GAN loss and optimize the discriminant model parameters; this process can be expressed as the following formula.
[0159]
[0160]
[0161]
[0162]
[0163] in, is a pre-trained feature extraction model, is the loss function, is the weight factor; is true.
[0164] Optimizing the discriminant model :
[0165]
[0166]
[0167]
[0168] in, is false; To optimize the discriminant model The loss function of and They are the loss functions for distinguishing true and false respectively.
[0169] It should be noted here that the discriminator is only used in the training process of the image quality driven degradation-aware decoder. After the image quality driven degradation-aware decoder training is completed, the discriminator is discarded.
[0170] The input of the image quality driven degradation-aware decoder of this embodiment is the intermediate features of the decoder and the residual connection from the encoder, and the degradation embedding is fused through an adaptive modulation strategy similar to the degradation-aware conditional image encoder. .
[0171] Embodiment 2
[0172] In one or more embodiments, Figure 8 As shown, an image super-resolution reconstruction system based on hybrid experts and stable diffusion may include:
[0173] An embedding extraction module 801 for acquiring a low-resolution image and extracting its degradation embedding and feature embedding;
[0174] A control condition generation module 802 is used to encode the low-resolution image using degradation embedding and a trained degradation-aware conditional image encoder, and then obtain multi-scale control conditions through a trained hybrid expert conditional control network; and map feature embedding from an image prompt space to a text prompt space using a trained hybrid expert image inversion adapter to obtain a spatial control condition;
[0175] The feature noise adding module 803 is used to encode the low-resolution image through the trained variational autoencoder, and then fuse it with the noise sampled by the noise scheduler to generate a noisy feature hidden layer space representation; or directly sample pure Gaussian noise to generate a noisy feature hidden layer space representation;
[0176] The noise prediction module 804 is used to use the hidden spatial representation of the noise-added feature as the input of the trained denoising backbone network in each sampling time step, and combine the multi-scale control condition and the spatial control condition to obtain the noise predicted by the denoising backbone network at each sampling time step;
[0177] The hidden space expression module 805 is used to continuously subtract the noise predicted by the denoising backbone network at the current time step from the input of the denoising backbone network at the current sampling time step through multiple sampling time steps, and then use the denoising result as the input of the denoising backbone network at the next time step. After multiple sampling time steps are completed, the hidden space expression of the reconstructed image is obtained;
[0178] The super-resolution image reconstruction module 806 is used to utilize the trained image quality driven degradation-aware decoder, combined with the intermediate features and degradation embedding of the variational autoencoder for low-resolution image encoding, to decode the latent spatial expression of the reconstructed image into the image pixel space, and obtain the image super-resolution reconstruction result.
[0179] It should be noted here that the various modules in the image super-resolution reconstruction system based on hybrid experts and stable diffusion correspond one to one with the various steps in the image super-resolution reconstruction method based on hybrid experts and stable diffusion, and their specific implementation processes are the same and will not be repeated here.
[0180] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains program code for executing the above-mentioned image super-resolution reconstruction method based on hybrid experts and stable diffusion.
[0181] The computer program instructions corresponding to the above-mentioned image super-resolution reconstruction method based on hybrid experts and stable diffusion can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which is implemented in the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0182] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0183] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An image super-resolution reconstruction method based on mixed experts and stable diffusion, characterized in that: include: Get a low-resolution image and extract its degradation embedding and feature embedding; The low-resolution image is encoded using degradation embedding and a trained degradation-aware conditional image encoder, and then passed through a trained hybrid expert conditional control network to obtain a multi-scale control condition; the feature embedding is mapped from the image prompt space to the text prompt space using a trained hybrid expert image inversion adapter to obtain a spatial control condition; The low-resolution image is encoded by a trained variational autoencoder, and then fused with the noise sampled by the noise scheduler to generate a noisy feature hidden layer space representation; or pure Gaussian noise is directly sampled to generate a noisy feature hidden layer space representation; In each sampling time step, the latent spatial representation of the noisy feature is used as the input of the trained denoising backbone network, and the noise predicted by the denoising backbone network at each sampling time step is obtained by combining the multi-scale control condition and the spatial control condition; After multiple sampling time steps, the noise predicted by the denoising backbone network at the current time step is continuously subtracted from the input of the denoising backbone network at the current sampling time step, and then the denoising result is used as the input of the denoising backbone network at the next time step. After multiple sampling time steps are completed, the latent spatial expression of the reconstructed image is obtained; By using the trained image quality-driven degradation-aware decoder and combining the intermediate features and degradation embedding of low-resolution image encoding with the variational autoencoder, the latent spatial expression of the reconstructed image is decoded into the image pixel space to obtain the image super-resolution reconstruction result.
2. The image super-resolution reconstruction method based on hybrid experts and stable diffusion as claimed in claim 1, characterized in that: In each sampling time step, the hidden spatial representation of the noisy features is used as the input of the denoising backbone network, the multi-scale control conditions are added to the down-sampled part of the denoising backbone network, and are concatenated and mapped with the features of the up-sampled part in the channel dimension; the spatial control conditions are cross-attention calculated with the intermediate features of the denoising backbone network.
3. The image super-resolution reconstruction method based on hybrid experts and stable diffusion as claimed in claim 1, characterized in that: The training process of the degradation-aware conditional image encoder is as follows: The low-resolution images in the training samples are used as the input of the degradation-aware conditional image encoder; Each basic module of the degradation-aware conditional image encoder embeds the degradation of the low-resolution image into the intermediate features through adaptive modulation, and then passes the intermediate features to the next basic module; The hidden layer representation output by the degradation-perceptual conditional image encoder is decoded by a decoder composed of basic modules, and the calculated loss of the real image corresponding to the decoded image and the low-resolution image is used as the training objective function to optimize the parameters of the degradation-perceptual conditional image encoder.
4. The image super-resolution reconstruction method based on hybrid experts and stable diffusion as claimed in claim 3, characterized in that: The training objective function of the degradation-aware conditional image encoder is: in, is the training objective function of the degradation-aware conditional image encoder; are the height, width and number of channels of the image respectively; For the real image The position pixel Channel values; To decode the image No. The position pixel channel value.
5. The image super-resolution reconstruction method based on hybrid experts and stable diffusion as claimed in claim 1, characterized in that: The expression of the hidden space representation of the noise feature is: ; or ; in, is the hidden space representation of the noise feature, represents the weight factor; represents the pure Gaussian noise sampled from the Gaussian distribution by the noise scheduler; is the feature hidden space representation.
6. The image super-resolution reconstruction method based on hybrid experts and stable diffusion as claimed in claim 1, characterized in that: The training process of the hybrid expert conditional control network and the hybrid expert image inversion adapter is: Extract degradation embedding and feature embedding for low-resolution images in training samples; Combining the low-resolution images and their degradation embeddings in the training samples, the trained degradation-aware conditional image encoder is used to encode the low-resolution images while preliminarily removing the degradation. The encoded features are passed through a hybrid expert conditional control network to generate multi-scale control conditions; The feature embedding is mapped from the image cue space to the text cue space through a hybrid expert image inversion adapter to obtain the spatial control condition; The high-resolution real image corresponding to the low-resolution image in the training sample is encoded by a variational autoencoder, and then fused with the noise sampled by the noise scheduler to generate a noisy feature hidden layer space representation; The hidden spatial representation of the noise-added features is used as the input of the denoising backbone network. The multi-scale control conditions are added to the downsampled part of the denoising backbone network, and are concatenated and mapped with the features of the upsampled part in the channel dimension. The spatial control conditions are cross-attentioned with the intermediate features of the denoising backbone network to obtain the noise predicted by the denoising backbone network. The parameters of the hybrid expert conditional control network and the hybrid expert image inversion adapter are optimized based on the computational loss of the sampled original noise and the predicted noise.
7. The image super-resolution reconstruction method based on hybrid experts and stable diffusion as claimed in claim 1, characterized in that: The training process of the image quality driven degradation-aware decoder includes: The variational autoencoder is used to encode the real image corresponding to the low-resolution image in the training sample to generate the decoder input hidden layer features; Use variational autoencoders to encode low-resolution images in training samples and obtain intermediate features of each layer; The hidden features of the real image are used as the input of the image quality driven degradation-aware decoder. Each layer fuses the low-resolution image intermediate features and degradation embedding of the corresponding layer of the encoder to decode the hidden features into pixel space. Freeze the discriminative model parameters and optimize the parameters of the decoder corresponding to the variational autoencoder based on the calculated pixel loss, perceptual loss, and GAN loss of the decoded image and the real image; Freeze the parameters of the variational autoencoder and its corresponding decoder, and optimize the parameters of the discriminative model based on the GAN loss calculated for the decoded image and the real image respectively; Repeat the above steps of freezing the discriminant model parameters and freezing the parameters of the variational autoencoder and its corresponding decoder until the decoder and the discriminant model converge.
8. An image super-resolution reconstruction system based on hybrid experts and stable diffusion, characterized in that: include: An embedding extraction module, which is used to acquire a low-resolution image and extract its degradation embedding and feature embedding; A control condition generation module is used to encode the low-resolution image using degradation embedding and a trained degradation-aware conditional image encoder, and then obtain multi-scale control conditions through a trained hybrid expert conditional control network; and map feature embedding from an image prompt space to a text prompt space using a trained hybrid expert image inversion adapter to obtain a spatial control condition; A feature noisy module is used to encode the low-resolution image through a trained variational autoencoder, and then fuse it with the noise sampled by the noise scheduler to generate a noisy feature latent space representation; or directly sample pure Gaussian noise to generate a noisy feature latent space representation; The noise prediction module is used to use the hidden spatial representation of the noise-added feature as the input of the trained denoising backbone network in each sampling time step, and combine the multi-scale control condition and the spatial control condition to obtain the noise predicted by the denoising backbone network at each sampling time step; The hidden space expression module is used to continuously subtract the noise predicted by the denoising backbone network at the current time step from the input of the denoising backbone network at the current sampling time step through multiple sampling time steps, and then use the denoising result as the input of the denoising backbone network at the next time step. After multiple sampling time steps are completed, the hidden space expression of the reconstructed image is obtained; The super-resolution image reconstruction module is used to utilize the trained image quality-driven degradation-aware decoder, combined with the intermediate features and degradation embedding of the variational autoencoder for low-resolution image encoding, to decode the latent spatial expression of the reconstructed image into the image pixel space to obtain the image super-resolution reconstruction result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the image super-resolution reconstruction method based on mixed experts and stable diffusion as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the image super-resolution reconstruction method based on mixed experts and stable diffusion as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Diffusion-based multi-modal image fusion
CN118401958A
Image super-resolution reconstruction method based on FSRCNN and OPE, electronic equipment and storage medium
CN118469818A