A lightweight single-step diffusion image super-resolution method based on content-adaptive time steps

CN121767186BActive Publication Date: 2026-09-01WUHAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511974857.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-09-01
Estimated Expiration
2045-12-25

AI Technical Summary

Technical Problem

在推理时,传统扩散采样需经过 50-200 步迭代去噪,在主流移动端芯片上处理512×512分辨率 LR 图像的单帧推理时间仍远远无法满足时效性需求

Benefits of technology

[0017] This invention provides a lightweight single-step diffusion image super-resolution method, system, storage medium, and electronic device based on content-adaptive time steps. First, a teacher network based on a latent space diffusion model is constructed. A pre-restoration module estimates the uncertainty of latent features in the image, and a content-adaptive optimal time step is derived by aligning the signal-to-noise ratio. Training with a multi-objective loss function yields a baseline model with accurate super-resolution capabilities. Then, a lightweight student network is constructed based on the teacher network's supervisory information. A convolutional fusion module overcomes the information bottleneck of VAEs and UNets. An adaptive strategy employing directional-frequency co-convolution, query-driven global aggregation, and hierarchical pruning compresses the network structure. The network is then trained using confidence-weighted distillation to ensure the student model approximates the teacher model's performance. After training, the input low-resolution image undergoes preprocessing, latent space mapping, single-step denoising, and decoding to directly output a high-resolution image. This invention significantly reduces the number of parameters and inference time while ensuring the richness of detail and semantic consistency of the reconstructed results, significantly improving the deployment efficiency of super-resolution reconstruction models in resource-constrained scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767186B_ABST
    Figure CN121767186B_ABST
Patent Text Reader

Abstract

This invention discloses a lightweight single-step diffusion image super-resolution method based on content-adaptive time steps. The method includes: acquiring paired high-resolution and low-resolution images to form a training set; fine-tuning a teacher network using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework; constructing a lightweight student network based on the teacher network and performing distillation training; wherein the lightweight student network is obtained by lightweighting and improving the teacher network; and inputting the low-resolution image to be reconstructed into the trained student network to obtain a high-resolution image. This invention can significantly reduce the number of model parameters and inference time while ensuring the detail richness and semantic consistency of the super-resolution results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a lightweight single-step diffusion image super-resolution method, system, storage medium, and electronic device based on content-adaptive time step. Background Technology

[0002] Image super-resolution reconstruction (SR), a core task in computer vision, aims to recover detailed and semantically consistent high-resolution (HR) images from low-resolution (LR) images. It is widely used in multimedia content distribution, remote sensing image interpretation, medical image diagnosis, and security monitoring image enhancement. With the rapid development of deep learning technology, super-resolution reconstruction methods have evolved from traditional optimization algorithms to deep generative methods based on convolutional neural networks (CNNs), Transformers, and diffusion models. Among these, diffusion models, represented by latent space diffusion models, have become one of the mainstream techniques in the current super-resolution field due to their powerful probabilistic generation capabilities and ability to model complex image texture distributions.

[0003] However, existing super-resolution reconstruction methods based on latent space diffusion models face numerous core challenges in practical applications, severely limiting their deployment in computationally limited scenarios (such as mobile devices, edge devices, and real-time super-resolution systems). Existing diffusion models suffer from large parameter counts, numerous inference steps, and inability to perform real-time inference in practical deployments. The core components of commonly used latent space diffusion models (VAE codec, denoising UNet, text encoder) often have a cumulative parameter count exceeding 1 billion, and the computational load typically reaches hundreds of billions of floating-point operations, placing extremely high demands on storage and computing power. During inference, traditional diffusion sampling requires 50-200 iterative denoising steps, and the single-frame inference time for processing 512×512 resolution LR images on mainstream mobile chips is still far from meeting timeliness requirements. Although some technical solutions have simplified the inference steps of diffusion models to a single step, and there are attempts to reduce the number of parameters by implementing structured pruning for single-step diffusion models, these improvements have not fundamentally solved the problem of balancing lightweight design with reconstruction quality. Summary of the Invention

[0004] This invention provides a lightweight single-step diffusion image super-resolution method, system, storage medium, and electronic device based on content-adaptive time step, which can significantly reduce the number of model parameters and inference time while ensuring the detail richness and semantic consistency of the super-resolution results.

[0005] This invention provides a lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, comprising: Acquire paired high-resolution and low-resolution images to form a training set; The teacher network is fine-tuned using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework. A lightweight student network is constructed based on the teacher network and distilled for training; wherein the lightweight student network is obtained by lightweighting and improving the teacher network. The low-resolution image to be reconstructed is input into the trained student network to obtain a high-resolution image.

[0006] Furthermore, according to the above-described lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, the teacher network includes a VAE encoder, a pre-restoration unit, a UNet denoising network, and a VAE decoder; the training process of the teacher network includes: The low-resolution image is input into the frozen VAE encoder, and the corresponding first low-quality latent feature is obtained through latent space mapping. The first low-quality latent feature is input into the pre-repair unit to obtain the mean and variance of the latent feature; The optimal time step is calculated based on the mean and variance of the latent features; The optimal time step, the first low-quality latent feature, and the embedding features of the optional text prompt words are input into the LoRA-optimized UNet denoising network to obtain the perceptually optimized first latent feature. The first latent feature, optimized by perception, is input into the LoRA-optimized decoder to obtain the initial super-resolution output; A multi-objective loss function is constructed to jointly train the pre-repairer, the LoRA-optimized UNet network, and the VAE decoder.

[0007] Furthermore, according to the above-described lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, the lightweight student network includes a lightweight VAE encoder, a lightweight UNet denoising network, and a lightweight VAE decoder. Two convolutional fusion units are embedded in the lightweight VAE encoder, the lightweight UNet denoising network, and the lightweight VAE decoder. The last layer of the lightweight VAE encoder and the first convolutional layer of the lightweight UNet denoising network constitute one convolutional fusion unit, and the first layer of the lightweight VAE decoder and the last convolutional layer of the lightweight UNet denoising network constitute another convolutional fusion unit. The convolutional fusion units are used to perform feature compression operations.

[0008] Furthermore, according to the above-described lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, the step of inputting the low-resolution image into the trained student network to obtain the high-resolution image includes: The low-resolution image to be reconstructed is preprocessed and then input into a lightweight VAE encoder to obtain the second low-quality latent feature. The second low-quality latent feature is input into the lightweight UNet denoising network to obtain the perceptually optimized second latent feature; The second latent feature, optimized by perception, is input into a lightweight VAE decoder to obtain a high-resolution image.

[0009] Furthermore, according to the above-mentioned lightweight single-step diffusion image super-resolution method based on content adaptive time step, the lightweight UNet denoising network is obtained by replacing the standard convolutional blocks in the traditional UNet denoising network with orientation-frequency co-modules. The orientation-frequency coordination module includes an orientation convolution module and a frequency separation convolution module; The directional convolution module includes horizontal convolution kernels and vertical convolution kernels, and the processing procedure of the directional convolution module includes: The first input features are passed through horizontal and vertical convolution kernels respectively to obtain horizontal and vertical features; The vertical features are input into the horizontal convolution kernel to obtain the diagonal features; The first input feature is subjected to global average pooling and then passed through a linear layer to obtain the first weighting coefficient; The horizontal feature, the vertical feature, and the diagonal feature are fused based on the first weighting coefficient to obtain the first fused feature; The first fused feature is input into a dimension-reducing linear layer and then into a dimension-increasing linear layer to obtain the first output feature.

[0010] Furthermore, according to the above-described lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, the processing procedure of the frequency separation convolution module includes: The second input feature is input into a large kernel convolutional layer to obtain a low-frequency smoothed feature, and the low-frequency smoothed feature and the second input feature are convolved to obtain a high-frequency smoothed feature. The second input feature is subjected to global average pooling and then passed through a linear layer to obtain the second weighting coefficient. The low-frequency smoothing feature and the high-frequency smoothing feature are fused based on the second weighting coefficient to obtain the second fused feature; The second fused feature is input into a dimension-reducing linear layer and then into a dimension-increasing linear layer to obtain the second output feature.

[0011] Furthermore, according to the above-described lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, the lightweight UNet denoising network further includes a query-driven global attention aggregation module, the processing procedure of which includes: The system performs a cross-attention operation on multiple pre-defined learnable query vectors and a third input feature, and then aggregates the global semantic features. The aggregated global semantic features are then subjected to a cross-attention operation with the third input features to obtain the third output features.

[0012] Furthermore, according to the above-described lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, the method further includes: During the joint training of the teacher network, a discriminator is used to identify the initial super-resolution output of the teacher network. Using the discriminator as a confidence estimator, the teacher-student distillation loss is weighted and averaged, as expressed by the following formula:

[0013] in, For teacher-student distillation loss, These refer to the intermediate layer features generated during the denoising process of UNet in the student network and teacher network, respectively. For discriminator, This is the initial super-resolution output for the teacher network.

[0014] The present invention also provides a lightweight single-step diffusion image super-resolution system based on content-adaptive time steps, comprising: The acquisition module is used to acquire paired high-resolution and low-resolution images to form a training set. The teacher network training fine-tuning module is used to fine-tune the teacher network using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework. The student network construction and training module is used to construct a lightweight student network based on the teacher network and train it; wherein the lightweight student network is obtained by lightweighting and improving the teacher network. The high-resolution reconstruction module is used to input the low-resolution image to be reconstructed into the trained student network to obtain a high-resolution image.

[0015] The present invention also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute any of the above-described lightweight single-step diffusion image super-resolution methods based on content-adaptive time steps.

[0016] The present invention also provides an electronic device including a processor and a memory, the processor being electrically connected to the memory, the memory being used to store instructions and data, and the processor being used in the steps of the lightweight single-step diffusion image super-resolution method based on content adaptive time step as described in any of the preceding claims.

[0017] This invention provides a lightweight single-step diffusion image super-resolution method, system, storage medium, and electronic device based on content-adaptive time steps. First, a teacher network based on a latent space diffusion model is constructed. A pre-restoration module estimates the uncertainty of latent features in the image, and a content-adaptive optimal time step is derived by aligning the signal-to-noise ratio. Training with a multi-objective loss function yields a baseline model with accurate super-resolution capabilities. Then, a lightweight student network is constructed based on the teacher network's supervisory information. A convolutional fusion module overcomes the information bottleneck of VAEs and UNets. An adaptive strategy employing directional-frequency co-convolution, query-driven global aggregation, and hierarchical pruning compresses the network structure. The network is then trained using confidence-weighted distillation to ensure the student model approximates the teacher model's performance. After training, the input low-resolution image undergoes preprocessing, latent space mapping, single-step denoising, and decoding to directly output a high-resolution image. This invention significantly reduces the number of parameters and inference time while ensuring the richness of detail and semantic consistency of the reconstructed results, significantly improving the deployment efficiency of super-resolution reconstruction models in resource-constrained scenarios. Attached Figure Description

[0018] The technical solution and other beneficial effects of the present invention will become apparent from the following detailed description of specific embodiments of the invention, in conjunction with the accompanying drawings.

[0019] Figure 1 The flowchart illustrates a lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, as provided in an embodiment of the present invention.

[0020] Figure 2 This is a schematic diagram of the structure of a teacher network provided in an embodiment of the present invention.

[0021] Figure 3 This is a schematic diagram of the structure of the convolutional fusion unit provided in an embodiment of the present invention.

[0022] Figure 4 This is a schematic diagram of the structure of the directional convolution module provided in an embodiment of the present invention.

[0023] Figure 5 This is a schematic diagram of the frequency separation convolution module provided in an embodiment of the present invention.

[0024] Figure 6 This is a schematic diagram of the structure of the query-driven global aggregation module provided in an embodiment of the present invention.

[0025] Figure 7This is a schematic diagram of the structure of a lightweight single-step diffusion image super-resolution system based on content-adaptive time step provided in an embodiment of the present invention.

[0026] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Current single-step super-resolution reconstruction schemes based on diffusion generation models have not fundamentally solved the problem of balancing lightweight design and reconstruction quality, and still have the following shortcomings in practical applications: (1) Using a fixed time step: Current single-step diffusion SR methods all use a fixed time step (e.g., using a fixed time step of 0 or 999), without considering the sensitivity of different image content and degradation degree to the degree of noise addition.

[0029] (2) The feature interaction logic between VAE (Variational Autoencoder) and UNet in the latent space diffusion model is not considered: There are usually feature dimensionality reduction and feature dimensionality increase operations at the connection between UNet and VAE in the latent space diffusion model. Such operations will lose the high-frequency detail information (such as texture edges and fine structures) carried by the features.

[0030] 3) Lightweight network structures lack design for the characteristics of natural images: The structures of objects in natural images have strong directionality, and their frequency distribution has obvious hierarchical clustering characteristics. Specifically, the edges, textures, and other structures of natural images exhibit significant directional correlations, and image information can be decomposed into low-frequency structural information and high-frequency detail information. However, existing lightweight SR networks have failed to fully utilize these features in their structural design.

[0031] 4) Lack of content awareness in teacher-student training: Conventional uniform distillation strategies often use uniform weighting, which cannot adaptively distill based on the quality and confidence differences of the teacher's output. This can easily introduce low-quality supervision information and affect the training effect of the student model.

[0032] To address the aforementioned problems, embodiments of the present invention provide a lightweight single-step diffusion image super-resolution method, system, storage medium, and electronic device based on content-adaptive time-stepping. The lightweight single-step diffusion image super-resolution system based on content-adaptive time-stepping provided by embodiments of the present invention can be integrated into an electronic device, which can be a terminal, server, or other device. The terminal can include tablet computers, laptops, personal computers (PCs), miniature processing boxes, or other devices.

[0033] Please see Figure 1 , Figure 1 The flowchart illustrates a lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, provided in an embodiment of the present invention. This method, applied in electronic devices, includes the following steps: S1: Acquire paired high-resolution and low-resolution images to form a training set.

[0034] S2, fine-tuning the teacher network using the training set; the teacher network is constructed based on a pre-trained latent space diffusion model framework.

[0035] The teacher network includes a VAE encoder, a pre-restor, a UNet denoising network, and a VAE decoder. Figure 2 This is a schematic diagram of the structure of a teacher network provided in an embodiment of the present invention; the training process of the teacher network includes: S21, input the low-resolution image into the frozen VAE encoder, and obtain the corresponding first low-quality latent feature through latent space mapping.

[0036] First, for low-resolution images Bicubic interpolation is used to upsample it to the target super-resolution. Then, input the frozen VAE encoder and obtain the corresponding first low-quality latent features through latent space mapping. .

[0037] S22, input the first low-quality latent feature into the pre-repairer to obtain the mean and variance of the latent feature.

[0038] The pre-repair unit is obtained by cascading multiple residual blocks. Low-quality hidden features are then integrated. The input is fed into a pre-repairing unit, which interacts with activation functions through multiple convolutions to output the mean of the latent features. With variance The variance of the latent features Used to characterize the uncertainty of the reconstructed features at each location in the latent space.

[0039] S23, calculate the optimal time step based on the mean and variance of the latent features.

[0040] Signal-to-noise ratio (SNR) definition based on diffusion process ( The cumulative product coefficient of the diffusion process, (For time steps), the mean of the latent features output by the pre-repairer As a measure of uncertainty, an alternative expression for SNR under the definition of uncertainty can be obtained. Combining the two equations yields the optimal time step. Expression: This method enables adaptive matching of time steps with the uncertainty in the latent space of the image.

[0041] S24. The optimal time step, the first low-quality latent feature, and the embedding features of the optional text prompt words are input into the LoRA-optimized UNet denoising network to obtain the perceptually optimized first latent feature.

[0042] Optimal time step First low-quality hidden features and the embedding features of optional text prompts The common input is the LoRA-optimized UNet denoising network, which then processes the LoRA-optimized denoising network. Perform noise prediction and removal, and output the first latent feature after perceptual optimization. .

[0043] LoRA (Low-Rank Adaptation Optimization) is an efficient fine-tuning technique for large models. Its core idea is to inject a set of additional, very small trainable matrices (low-rank adapters) instead of directly modifying the weights of the original large neural network (such as the UNet denoising network). During training, only these small adapters are trained, while the original model weights remain frozen. During inference, the effects of the adapters are incorporated into the original weights, with almost no increase in computational cost.

[0044] S25, the first latent feature optimized by perception is input into the LoRA-optimized decoder to obtain the initial super-resolution output. .

[0045] Furthermore, the initial super-resolution output is then discriminated using a discriminator.

[0046] S26. Construct a multi-objective loss function to jointly train the pre-repairer, the LoRA-optimized UNet network, and the VAE decoder.

[0047] The multi-objective loss function is:

[0048] in, For the true value of the target high-resolution image, For pixel-level loss, In order to perceive loss, For the losses of conflict and confrontation, The loss used for pre-repair uncertainty estimation is specifically expressed as:

[0049] In a specific implementation, the weight coefficients of each loss function are taken as follows: , , , The AdamW optimizer was used, with an initial learning rate of 3e-4 and a weight decay coefficient of 1e-5. After training for 300K epochs, a stable teacher model was obtained. .

[0050] S3 constructs a lightweight student network based on the teacher network and trains it; the lightweight student network is obtained by making lightweight improvements based on the teacher network.

[0051] The lightweight student network consists of a lightweight VAE encoder, a lightweight UNet denoising network, and a lightweight VAE decoder; Two convolutional fusion units are embedded in the lightweight VAE encoder, the lightweight UNet denoising network, and the lightweight VAE decoder. The last layer of the lightweight VAE encoder and the first convolutional layer of the lightweight UNet denoising network constitute one convolutional fusion unit, and the first layer of the lightweight VAE decoder and the last convolutional layer of the lightweight UNet denoising network constitute another convolutional fusion unit. The convolutional fusion units are used to perform feature compression operations.

[0052] Figure 3 This is a schematic diagram of the structure of the convolutional fusion unit provided in an embodiment of the present invention, as shown below. Figure 3 As shown, convolutional fusion units are designed between the lightweight VAE encoder and the UNet connection layer in the student model, and between the VAE decoder and the UNet connection layer, to replace the traditional feature compression operation. The last layer of the lightweight VAE encoder and the first layer of the UNet convolutional layer together constitute the convolutional fusion unit. In the traditional VAE encoder and UNet connection layer, the former significantly reduces the number of feature channels ( , (This represents the number of input channels in the last convolutional layer of the VAE encoder), which in turn significantly increases the number of feature channels. , This represents the number of output channels of the first convolutional layer in UNet. To alleviate the information bottleneck problem caused by intermediate layer features, this embodiment of the invention merges these two convolutional layers to form a convolutional fusion unit that does not require a significant reduction in the number of channels. ).

[0053] Furthermore, a pruning strategy is adopted to improve the UNet denoising network to achieve lightweight design.

[0054] In one embodiment, the lightweight UNet denoising network is obtained by replacing the standard convolutional blocks in the traditional UNet denoising network with orientation-frequency co-operation modules.

[0055] The orientation-frequency co-convolution module includes an orientation convolution module and a frequency separation convolution module.

[0056] The directional convolution module includes horizontal convolution kernels and vertical convolution kernels. Figure 4 This is a schematic diagram of the structure of the directional convolution module provided in an embodiment of the present invention. The processing procedure of the directional convolution module includes: S311, the first input features are passed through horizontal and vertical convolution kernels respectively to obtain horizontal and vertical features.

[0057] It should be noted that the first input feature refers to the feature input to the directional convolution module. The size of the horizontal convolution kernel is 1×3, and the size of the vertical convolution kernel is 3×1.

[0058] S312 inputs the vertical features into the horizontal convolution kernel to obtain the diagonal features.

[0059] S313, after global average pooling of the first input features, pass them through a linear layer to obtain the first weighting coefficient; S314, based on the first weighting coefficient, the horizontal features, vertical features and diagonal features are fused to obtain the first fused feature.

[0060] Specifically, the fusion process can be represented by the following formula:

[0061] in, The first fusion feature, It is a horizontal feature. As a vertical feature, It is a slanted feature. , The first weighting coefficient corresponds to the importance weights of the horizontal, vertical, and diagonal features, respectively.

[0062] S315, the first fused feature is input into the dimension reduction linear layer and then into the dimension increase linear layer to obtain the first output feature.

[0063] Figure 5 This is a schematic diagram of the frequency separation convolution module provided in an embodiment of the present invention. The processing procedure of the frequency separation convolution module includes: S321, the second input feature is input into the large kernel convolutional layer to obtain the low-frequency smooth feature, and the low-frequency smooth feature and the second input feature are convolved to obtain the high-frequency smooth feature.

[0064] It should be noted that the second input feature value is input to the frequency separation convolution module, which is the first output feature of the directional convolution module. The convolution size of the large kernel convolution layer is 7×7.

[0065] S322, after global average pooling of the second input features, the second weighting coefficients are obtained by passing them through a linear layer.

[0066] S323, based on the second weighting coefficient, the low-frequency smoothing feature and the high-frequency smoothing feature are fused to obtain the second fused feature.

[0067] Specifically, the fusion process can be represented by the following formula:

[0068] in, This is the second fusion feature. It has low-frequency smoothing characteristics. It is a high-frequency smoothing feature. These are the corresponding importance weights.

[0069] S324, the second fused feature is input into the dimension reduction linear layer and then into the dimension increase linear layer to obtain the second output feature.

[0070] Furthermore, the lightweight UNet denoising network also includes a query-driven global attention aggregation module. Figure 6 This is a schematic diagram of the structure of the query-driven global aggregation module provided in an embodiment of the present invention. The processing procedure of the query-driven global attention aggregation module includes: S331 performs a cross-attention operation on multiple preset learnable query vectors and a third input feature to aggregate global semantic features.

[0071] It should be noted that the third input feature refers to the feature input into the query-driven global aggregation module. The process of this cross-attention operation can be represented by the following formula:

[0072] in, To aggregate global semantic features, For cross-attention operations, For learnable vectors, This is the third input feature.

[0073] S332, the aggregated global semantic features are then subjected to a cross-attention operation with the third input features to obtain the third output features.

[0074] The process of this cross-attention operation can be represented by the following formula:

[0075] in, This is the third output feature.

[0076] In this way, the complexity of attention operations will be reduced from that of global attention. Reduce to (in This refers to the number of tokens in the input features. (This) while retaining the ability to extract global features.

[0077] It should be noted that the deep modules (the latter half of the decoder) in the UNet network are responsible for semantic reconstruction, while the shallow modules (the first few layers of the encoder and decoder) are responsible for detail restoration. However, traditional methods always use uniform channel pruning when pruning, without taking into account the fact that most of the semantic information already exists in the low-resolution image. Therefore, this embodiment of the invention adopts a hierarchical pruning strategy, pruning all deep modules in UNet, while performing the lightweight strategy described in steps S311-S315, S321-S324, and S331-S332 for shallow modules.

[0078] Furthermore, a multi-objective loss function is constructed for joint training of the UNet and VAE encoders and decoders of the student network. This multi-objective loss function integrates pixel-level loss, perceptual loss, adversarial loss, and teacher-student distillation loss. The resulting multi-objective loss function is:

[0079] in, These refer to the intermediate layer features generated during the denoising process of UNet in the student network and teacher network, respectively.

[0080] Considering that the supervision provided by the teacher network has different quality for different image content, and that the distillation loss should be assigned different weight factors for supervision of different quality, we use the discriminator from the teacher network training as a confidence estimator to perform weighted averaging of the distillation loss, thereby improving the supervision quality through confidence-weighted distillation.

[0081] in, For teacher-student distillation loss, For discriminator, These refer to the intermediate layer features generated during the denoising process of UNet in the student network and teacher network, respectively. This is the initial super-resolution output for the teacher network.

[0082] In a specific implementation, the weight coefficients of each loss function are taken as follows: , , , The AdamW optimizer was used, with an initial learning rate of 3e-4 and a weight decay coefficient of 1e-5. After training for 100K epochs, a stable student model was obtained. .

[0083] S4. Input the low-resolution image to be reconstructed into the trained student network to obtain a high-resolution image.

[0084] In one embodiment, step S4 specifically includes: S41, after preprocessing the low-resolution image to be reconstructed, it is input into the lightweight VAE encoder to obtain the second low-quality latent feature. S42, the second low-quality hidden feature is input into the lightweight UNet denoising network to obtain the perceptually optimized second hidden feature; S43, the perceptually optimized second latent feature is input into the decoder to obtain a high-resolution image.

[0085] It should be noted that the first and second books mentioned above are for distinguishing between the inputs and outputs of the teacher network and the inputs and outputs of the student network.

[0086] To quantitatively verify the effectiveness of the method of this invention, multi-dimensional image quality evaluation metrics are used for performance evaluation, including pixel fidelity metrics, perceptual quality metrics, and no-reference quality evaluation metrics, comprehensively covering evaluation dimensions of objective accuracy and subjective visual effect. The training process uses standard high-resolution image datasets containing various scene types to ensure the generalization performance of the model; the testing phase uses benchmark datasets for mainstream super-resolution reconstruction tasks to ensure the objectivity and comparability of the evaluation results.

[0087] Based on the method described in the above embodiments, this embodiment will further describe it from the perspective of a lightweight single-step diffusion image super-resolution system based on content adaptive time step. Specifically, the lightweight single-step diffusion image super-resolution system based on content adaptive time step can be implemented as an independent entity or integrated into an electronic device. The electronic device can be a terminal, server, or other device. The terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessor box, or other devices.

[0088] Please see Figure 7 , Figure 7 This invention specifically describes a lightweight single-step diffusion image super-resolution system based on content-adaptive time steps, applicable to electronic devices. The system may include: The acquisition module is used to acquire paired high-resolution and low-resolution images to form a training set. The teacher network fine-tuning training module is used to train the teacher network using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework. The student network construction and training module is used to construct a lightweight student network based on the teacher network and train it; wherein the lightweight student network is obtained by lightweighting and improving the teacher network. The high-resolution reconstruction module is used to input the low-resolution image to be reconstructed into the trained student network to obtain a high-resolution image.

[0089] In specific implementation, the above modules and / or units can be implemented as independent entities, or they can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of the above modules and / or units, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments, which will not be repeated here.

[0090] In addition, embodiments of the present invention also provide an electronic device, which may be a computer, tablet computer, or other similar device. This electronic device can implement the steps of any embodiment of the lightweight single-step diffusion image super-resolution method based on content-adaptive time step provided in the embodiments of the present invention. Therefore, it can achieve the beneficial effects achievable by any lightweight single-step diffusion image super-resolution method based on content-adaptive time step provided in the embodiments of the present invention, as detailed in the preceding embodiments, and will not be repeated here.

[0091] Figure 8 A specific structural block diagram of an electronic device provided in an embodiment of the present invention is shown. This electronic device can be used to implement the lightweight single-step diffusion image super-resolution method based on content adaptive time step provided in the above embodiments. The electronic device 500 can be a terminal, server, or other device. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.

[0092] The memory 520 can be used to store software programs and modules, such as the program instructions / modules corresponding to those in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520. The memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 520 may further include memory remotely located relative to the processor 580, and these remote memories can be connected to the electronic device 500 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0093] The input unit 530 can be used to receive input numeric or character information, and to generate a keyboard and mouse related to user settings and function control. Display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which can be composed of graphics, text, icons, video, and any combination thereof. Display unit 540 may include display panel 541, which may optionally be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms.

[0094] Electronic device 500, through transmission module 570 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 570 is shown in the figure, it is understood that it is not an essential component of electronic device 500 and can be omitted as needed without changing the essence of the invention.

[0095] The processor 580 is the control center of the electronic device 500. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 500 by running or executing software programs and / or modules stored in the memory 520, and by calling data stored in the memory 520, thereby providing overall monitoring of the electronic device. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 580.

[0096] Electronic device 500 also includes a power supply 590 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to processor 580 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 590 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0097] Although not shown, the electronic device 500 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the following operations: Acquire paired high-resolution and low-resolution images to form a training set; The teacher network is fine-tuned using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework. A lightweight student network is constructed based on the teacher network and trained; wherein the lightweight student network is obtained by lightweighting and improving the teacher network. The low-resolution image to be reconstructed is input into the trained student network to obtain a high-resolution image.

[0098] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.

[0099] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any embodiment of the content-adaptive time-step-based lightweight single-step diffusion image super-resolution method provided by the present invention.

[0100] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0101] Since the instructions stored in the storage medium can execute the steps in any embodiment of the content-adaptive time-step-based lightweight single-step diffusion image super-resolution method provided in the embodiments of the present invention, the beneficial effects that any content-adaptive time-step-based lightweight single-step diffusion image super-resolution method provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0102] The foregoing has provided a detailed description of a lightweight single-step diffusion image super-resolution method, system, storage medium, and electronic device based on content-adaptive time step provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A lightweight single-step diffusion image super-resolution method based on content-adaptive time steps, characterized in that, The method includes: Acquire paired high-resolution and low-resolution images to form a training set; The teacher network is fine-tuned using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework. A lightweight student network is constructed based on the teacher network and distilled for training; wherein the lightweight student network is obtained by lightweighting and improving the teacher network. The teacher network includes a VAE encoder, a pre-restoration unit, a UNet denoising network, and a VAE decoder. The training process of the teacher network includes: inputting a low-resolution image into a frozen VAE encoder to obtain the corresponding first low-quality latent features through latent space mapping; inputting the first low-quality latent features into the pre-restoration unit to obtain the mean and variance of the latent features; calculating the optimal time step based on the mean and variance of the latent features; inputting the optimal time step, the first low-quality latent features, and the embedding features of optional text prompts into a LoRA-fine-tuned and optimized UNet denoising network to obtain the perceptually optimized first latent features; inputting the perceptually optimized first latent features into a LoRA-fine-tuned and optimized VAE decoder to obtain the initial super-resolution output; and constructing a multi-objective loss function to jointly train the pre-restoration unit, the LoRA-optimized UNet network, and the VAE decoder. The lightweight student network includes a lightweight VAE encoder, a lightweight UNet denoising network, and a lightweight VAE decoder. Two convolutional fusion units are embedded in the lightweight VAE encoder, the lightweight UNet denoising network, and the lightweight VAE decoder. The last layer of the lightweight VAE encoder and the first convolutional layer of the lightweight UNet denoising network constitute one convolutional fusion unit, and the first layer of the lightweight VAE decoder and the last convolutional layer of the lightweight UNet denoising network constitute another convolutional fusion unit. These convolutional fusion units are used for feature compression operations. During the joint training of the teacher network, a discriminator is used to identify the initial super-resolution output of the teacher network; the discriminator is used as a confidence estimator to perform a weighted average of the teacher-student distillation loss, expressed by the following formula: in, For teacher-student distillation loss, These refer to the intermediate layer features generated during the denoising process of UNet in the student network and teacher network, respectively. For discriminator, The initial super-resolution output for the teacher network output; The low-resolution image to be reconstructed is input into the trained student network to obtain a high-resolution image.

2. The lightweight single-step diffusion image super-resolution method based on content-adaptive time step as described in claim 1, characterized in that, The step of inputting the low-resolution image to be reconstructed into the trained student network to obtain a high-resolution image includes: The low-resolution image to be reconstructed is preprocessed and then input into a lightweight VAE encoder to obtain the second low-quality latent feature. The second low-quality latent feature is input into the lightweight UNet denoising network to obtain the perceptually optimized second latent feature; The second latent feature, optimized by perception, is input into a lightweight VAE decoder to obtain a high-resolution image.

3. The lightweight single-step diffusion image super-resolution method based on content-adaptive time step as described in claim 1, characterized in that, The lightweight UNet denoising network is obtained by replacing the standard convolutional blocks in the traditional UNet denoising network with orientation-frequency co-modules. The orientation-frequency coordination module includes an orientation convolution module and a frequency separation convolution module; The directional convolution module includes horizontal convolution kernels and vertical convolution kernels, and the processing procedure of the directional convolution module includes: The first input features are passed through horizontal and vertical convolution kernels respectively to obtain horizontal and vertical features; The vertical features are input into the horizontal convolution kernel to obtain the diagonal features; The first input feature is subjected to global average pooling and then passed through a linear layer to obtain the first weighting coefficient; The horizontal feature, the vertical feature, and the diagonal feature are fused based on the first weighting coefficient to obtain the first fused feature; The first fused feature is input into a dimension-reducing linear layer and then into a dimension-increasing linear layer to obtain the first output feature.

4. The lightweight single-step diffusion image super-resolution method based on content-adaptive time step as described in claim 3, characterized in that, The processing steps of the frequency separation convolution module include: The second input feature is input into a large kernel convolutional layer to obtain a low-frequency smoothed feature, and the low-frequency smoothed feature and the second input feature are convolved to obtain a high-frequency smoothed feature. The second input feature is subjected to global average pooling and then passed through a linear layer to obtain the second weighting coefficient. The low-frequency smoothing feature and the high-frequency smoothing feature are fused based on the second weighting coefficient to obtain the second fused feature; The second fused feature is input into a dimension-reducing linear layer and then into a dimension-increasing linear layer to obtain the second output feature.

5. The lightweight single-step diffusion image super-resolution method based on content-adaptive time step as described in claim 1, characterized in that, The lightweight UNet denoising network also includes a query-driven global attention aggregation module, the processing of which includes: The system performs a cross-attention operation on multiple pre-defined learnable query vectors and a third input feature, and then aggregates the global semantic features. The aggregated global semantic features are then subjected to a cross-attention operation with the third input features to obtain the third output features.

6. A lightweight single-step diffusion image super-resolution system based on content-adaptive time step, wherein the lightweight single-step diffusion image super-resolution system based on content-adaptive time step is used to implement the lightweight single-step diffusion image super-resolution method based on content-adaptive time step as described in claim 1, characterized in that, include: The acquisition module is used to acquire paired high-resolution and low-resolution images to form a training set. The teacher network fine-tuning training module is used to train the teacher network using the training set; wherein the teacher network is constructed based on a pre-trained latent space diffusion model framework. The student network construction and training module is used to construct a lightweight student network based on the teacher network and train it; wherein the lightweight student network is obtained by lightweighting and improving the teacher network. The high-resolution reconstruction module is used to input the low-resolution image to be reconstructed into the trained student network to obtain a high-resolution image.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to execute the lightweight single-step diffusion image super-resolution method based on content-adaptive time steps as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Lightweight image super-resolution reconstruction method based on multi-dimensional knowledge distillation

    CN113240580A