Sketch guide image generation method and system based on stable diffusion model
Through multimodal feature fusion and orthogonal butterfly fine-tuning technology, the stable diffusion model is optimized, which solves the problems of detail loss and texture blur in sketch image generation, and achieves the improvement of high-quality image generation and training efficiency.
Patent Information
- Application Number
- CN202510692800.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-22
AI Technical Summary
When generating sketch images, existing stable diffusion models have insufficient heterologous feature matching leads to loss of details, texture blur and structural distortion, and full-parameter fine-tuning calculation cost is high and easy to overfit.
The multimodal feature fusion method is used to extract the sketch edge structure features and the CLIP text encoder through ControlNet to generate semantic feature vectors, combine orthogonal butterfly fine-tuning technology (BOFT) to optimize attention layer parameters, avoid full parameter fine-tuning, and use composite loss function to optimize image quality.
It improves the detail accuracy and structural consistency of image generation, reduces the calculation cost, avoids the risk of gradient explosion, and improves training stability and efficiency.
Smart Images

Figure CN120525985A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a sketch-guided image generation method and system based on a stable diffusion model. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Sketch-to-image generation is a classic cross-modal research topic that aims to learn a mapping between real images and freehand sketches to bridge the domain gap between the two. Sketches are an intuitive and flexible form of expression, but due to differences in drawing skills, they may exhibit varying degrees of abstraction.
[0004] Existing technologies use a stable diffusion model to generate clear images from sketches. These clear images are characterized by high quality and richer details, creating realistic images similar to natural landscapes and human figures. However, the existing pre-trained Stable Diffusion v1.5 uses hand-drawn sketches and text as input to generate images. Because sketches are geometrically distorted and lack color, texture, and other visual details, the generative model is unable to accurately match heterogeneous features, resulting in loss of detailed features and affecting the generation of texture and structural features. This leads to blurred textures and distorted structures, ultimately resulting in low-quality generated images with distorted textures, colors, and shapes. Furthermore, existing full-parameter fine-tuning of pre-trained diffusion models incurs high computational costs due to updating all network layers and is prone to overfitting when the target dataset is small. Summary of the Invention
[0005] To overcome the above-mentioned shortcomings of the existing technology, the present invention proposes a sketch-guided image generation method and system based on a stable diffusion model, which solves the problem of detail loss caused by mismatch of heterogeneous feature spaces in traditional methods, significantly improves the texture blurring and structural distortion problems in traditional methods, and avoids the overfitting problem in parameter fine-tuning.
[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, a sketch-guided image generation method based on a stable diffusion model is disclosed, comprising: Obtain sketch data and text data; Inputting the sketch data and text data into a constructed multimodal conditional input module, wherein the multimodal conditional input module includes an edge feature extraction module and a text encoder, wherein the edge feature extraction module extracts edge structural features of the sketch data, and the text encoder extracts text features of the text data to generate a semantic feature vector, thereby obtaining a multimodal feature; Based on a pre-trained stable diffusion model, the multimodal features are diffused to generate a clear image.
[0007] In a second aspect, a sketch-guided image generation system based on a stable diffusion model is disclosed, comprising: A data acquisition module is configured to: acquire sketch data and text data; A feature extraction module is configured to: input the sketch data and text data into a constructed multimodal conditional input module, wherein the multimodal conditional input module includes an edge feature extraction module and a text encoder, wherein the edge feature extraction module extracts edge structural features of the sketch data, and the text encoder extracts text features of the text data to generate a semantic feature vector to obtain a multimodal feature; The image generation module is configured to generate a clear image by using the multimodal features through a diffusion process based on a pre-trained stable diffusion model.
[0008] In a third aspect, an electronic device is disclosed, comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps of the sketch-guided image generation method based on the stable diffusion model are completed.
[0009] In a fourth aspect, a computer-readable storage medium is disclosed for storing computer instructions. When the computer instructions are executed by a processor, the steps of the sketch-guided image generation method based on the stable diffusion model are completed.
[0010] Compared with the prior art, the present invention has the following beneficial effects: 1. Improved precision of multimodal feature fusion: This paper uses ControlNet and CLIP text encoders to achieve deep fusion of sketch edge structure and text semantic features, effectively solving the problem of detail loss caused by mismatch of heterogeneous feature spaces in traditional methods.
[0011] 2. Optimized generated image quality: A composite loss function is used to jointly optimize pixel-level reconstruction and perceptual similarity, significantly improving the texture blur and structural distortion problems in traditional methods.
[0012] 3. Enhanced Training Efficiency and Stability: This paper employs the Orthogonal Butterfly Tuning technique (BOFT) to constrain parameter updates through low-rank decomposition. This technique optimizes only the parameters in the attention layer and freezes the remaining parameters. This effectively preserves pre-trained knowledge and prevents catastrophic forgetting. It also reduces the optimization difficulty and improves training stability by reducing the number of trainable parameters. Model optimization requires only updating 1% of the attention layer parameters, effectively avoiding the risk of gradient explosion associated with traditional full-parameter tuning while maintaining the expressive power of the underlying model.
[0013] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0015] Figure 1 This is a flow chart of the sketch-guided image generation method based on the stable diffusion model described in Example 1 of the present invention.
[0016] Figure 2 Schematic diagram of the training process of the sketch-text-to-image generation model described in Example 1 of the present invention.
[0017] Figure 3 This is a framework diagram of the image-to-sketch conversion module described in Example 1 of the present invention.
[0018] Figure 4 This is a flowchart of the BOFT fine-tuning process of the stable diffusion model training process described in Example 1 of the present invention. DETAILED DESCRIPTION
[0019] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0020] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0021] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0022] Example 1 In one or more embodiments, a sketch-guided image generation method based on a stable diffusion model is disclosed, such as Figure 1As shown, the following steps are included: Step S1: Obtain sketch data and corresponding text prompt data, and use a sketch-text-to-image generation model to generate a clear image. The sketch-text-to-image generation model includes a multimodal conditional input module and a stable diffusion model.
[0023] Step S2: Input the sketch data and text data into the constructed multimodal conditional input module. The multimodal conditional input module includes an edge feature extraction module and a text encoder. The edge feature extraction module extracts the edge structure features of the sketch data, and the text encoder extracts the text features of the text data to generate a semantic feature vector to obtain multimodal features.
[0024] Specifically, in this embodiment, the edge feature extraction module adopts the pre-trained ControlNet network structure, and the text encoder adopts the CLIP encoder. The sketch data and the corresponding text prompts in the dataset are input into the pre-trained ControlNet network structure and the CLIP encoder respectively. ControlNet extracts the edge structure features of the sketch, and the CLIP text encoder extracts the text features to generate a semantic feature vector. The sketch features are directly superimposed on the UNet feature map through the residual connection of ControlNet, and the text features are aligned with the image block relationship through the cross-attention mechanism of UNet.
[0025] Among them, the pre-trained ControlNet network introduces sketch conditional control signals and uses a deep conditional encoder to perform nonlinear mapping on the input sketch to generate a feature tensor that is isomorphic to the Stable Diffusion latent space dimension. It also uses a residual connection strategy initialized with zero convolution to achieve multi-scale fusion of the control signal and the diffusion process, realize fine-grained control of image generation, and thus achieve precise spatial constraints on the generation process.
[0026] The CLIP text encoder module converts text prompt words into semantic embedding vectors, and through the cross-attention layer in UNet, interacts the text features with the image latent space features to guide the direction of image generation.
[0027] Step S3: Based on the pre-trained stable diffusion model, the multimodal features are diffused to generate a clear image.
[0028] In this embodiment, the pre-trained Stable Diffusion v1.5 model (also known as the diffusion model) is used as the pre-trained Stable Diffusion v1.5 model. Multimodal features are input into the pre-trained Stable Diffusion v1.5 model's UNet architecture, which generates high-resolution images through a diffusion process. The diffusion model consists of a UNet module and a variational autoencoder (VAE) connected in sequence.
[0029] The Stable Diffusion model performs forward and backward denoising operations in the latent space. The StableDiffusion model adopts a two-stage approach. First, the sketch edge structure features extracted by the ControlNet network are fed into the Unet module. The Unet denoiser performs conditional denoising directly in the latent space based on the textual prompt p. The UNet module iteratively denoises the latent space, generating a latent representation that conforms to the text description and sketch structure. The decoder of the VAE then upsamples the latent representation to produce a clear image.
[0030] In this embodiment, the pre-training process of the sketch-text to image generation model is as follows: Figure 2 As shown in the figure, it includes: first, taking images and text as input, extracting sketch edge structure features through the ControlNet architecture, and using the CLIP text encoder to encode the text to generate a semantic feature vector. Secondly, the VAE encoder is used to add noise to the real image. The pre-trained image-sketch conversion module is used to form a closed-loop feedback with the LPIPS evaluation system, and the mean square error (MSE) loss is used to finally perform a targeted parameter update on the UNet attention layer through the BOFT technology. The Adam optimizer is used to update the parameters of the model to minimize the total loss function. The hyperparameter settings of the Adam optimizer are as follows: the learning rate is 5e-6, the momentum parameter beta1 is 0.9, the momentum parameter beta2 is 0.999, and the epsilon is 1e-08.
[0031] It should be noted that, in this embodiment, the training data set during the model pre-training process includes sketches, edge maps, texts, and real images, which are one-to-one corresponding data.
[0032] The model training process is as follows: Step S3-1: Construct a multimodal conditional input module. The sketch data and corresponding text prompts in the dataset are fed into the pre-trained ControlNet network structure and CLIP encoder, respectively. ControlNet extracts edge structure features from the sketch, while the CLIP encoder extracts text features to generate semantic feature vectors.
[0033] Step S3-2: Input the edge structure features and semantic feature vectors and the noisy data into the UNet module to generate a first generated image, and use the image-to-sketch conversion module to inversely convert the first generated image into a first generated sketch.
[0034] Step S3-2-1. Input the real image in the training data set into the VAE encoder for encoding, convert it into a low-dimensional latent space representation, and create a random Gaussian noise matrix with exactly the same dimension as the latent space. The noise value obeys the standard normal distribution. The diffusion time step is randomly selected independently for each training sample to control the current noise intensity. The noise function of the DDIM scheduler is called, and Gaussian noise is gradually added to the real image through the forward diffusion process to generate noisy latent space data as the training target.
[0035] Among them, generating noisy data , the noisy data and time step t are input into Unet, and the output is the predicted noise , compared with the predicted noise With real noise , calculate the mean square error, and the mean square error is used as the first optimization objective function for optimizing the Unet network.
[0036] The UNet module iteratively denoises in the latent space to generate a latent representation that conforms to the text description and sketch structure. The VAE decoder upsamples the latent representation and converts the latent representation generated by the diffusion model back into a clear image in the pixel space, that is, generating the first generated image.
[0037] Step S3-2-2: construct a pre-trained image-to-sketch conversion module based on a generative adversarial network (GAN), and use a conditional generator to inversely convert the output image into the corresponding sketch; Specifically, a pre-trained generative adversarial network (GAN)-based image-to-sketch conversion module is constructed, consisting of a generator and a discriminator. During training, the discriminator uses an adversarial style loss to encourage the generated line drawings to match the style of the training set. The generator implements an end-to-end image generation architecture using a downsampling-residual-upsampling structure. Initial convolutional blocks perform preliminary feature extraction on the input first generated image, maintaining the spatial size and using reflective padding to reduce boundary artifacts. The downsampling module uses two downsampling operations to gradually compress the image size and extract high-level abstract features. The residual block module uses deep residual learning to enhance feature representation, preventing gradient vanishing and preserving detailed information. The upsampling module uses two upsampling operations to gradually restore the image size and reconstruct details. The output layer maps the features to the target image space, resulting in the first generated sketch.
[0038] like Figure 3As shown in the figure, the pre-trained bidirectional conversion mechanism based on the adversarial generative network is demonstrated. After receiving real image input, the generator calculates the CLIP semantic loss, appearance similarity loss, and geometric structure loss to form an adversarial training system with the discriminator. The hand-drawn airplane sketch at the output intuitively verifies the effectiveness of this module in image style transfer while preserving key features. The superposition of the three loss functions provides a multi-dimensional evaluation benchmark for conversion quality.
[0039] Step S3-3: introducing a learning-based perceptual image patch similarity (LPIPS) evaluation system, establishing an image conversion quality evaluation module, and evaluating the conversion quality of the first generated sketch; Specifically, by calculating the structural similarity between the first generated sketch and the corresponding edge map in the multi-layer feature space, an optimization objective function based on perceptual loss is constructed.
[0040] The Learned Perceptual Patch Similarity (LPIPS) is an image similarity assessment metric based on deep neural network features. It aims to more closely align image quality or similarity metrics with human visual perception. It uses multiple convolutional layers of a pretrained deep convolutional network (AlexNet) to extract features from the generated sketch and its corresponding edge map. The feature maps of each layer of the AlexNet network are unit-normalized in the channel dimension to eliminate feature amplitude differences while preserving directional information. The spatial dimensions (height and width) of each feature map are then averaged to produce a global feature vector for each layer. A weighted Euclidean distance is calculated between the features of each layer of the generated sketch and its corresponding edge map, ultimately outputting the distance between the two image patches.
[0041] Step S3-4: construct a second optimization objective function based on perceptual loss according to the conversion quality (distance value), and perform weighted fusion of the second optimization objective function and the first optimization objective function to obtain a total loss function.
[0042] The first optimization objective function is the mean square error loss, which is expressed as: (1) Where, is the mean square error loss; E is the mathematical expectation; t is the time step number of the diffusion process; is the original input data; is random noise that follows a standard normal distribution, and ; is the noisy data state at time step t; is a parameterized noise prediction neural network. The loss function optimizes the network parameters through the back propagation algorithm θ , so that the model gradually learns from noisy data The precise mapping relationship between the noise separation.
[0043] The second optimization objective function is the perceptual similarity index, which is expressed as: (2) Where, is the perceptual similarity index; and Represents the feature vectors at the (h, w) position of the input image x and y in the feature map of the lth layer of the deep network respectively; x represents the sketch generated by the generative adversarial network; y represents the corresponding edge map loaded by the dataset; is the trainable channel weight vector of layer l; For Hadamard; and Represents the height and width of the l-th layer feature map.
[0044] The total loss function is: (3) Where α and β are adjustable weight coefficients.
[0045] Step S3-5: Use the Orthogonal Butterfly Tuning Technique (BOFT) to update the parameters of the Stable Diffusion v1.5 model, specifically including: A trainable butterfly matrix is inserted into the cross-attention layer of the U-Net module. The butterfly matrix consists of orthogonal basis vectors and is used to adjust the mapping relationship between text-sketch and image features. The basic model parameters are frozen, and only the butterfly matrix parameters of the attention layer are gradient updated. Low-rank decomposition technology is used to constrain the parameter update direction to maintain the stability of model generation.
[0046] like Figure 4 Figure 2 shows the model parameter optimization process. Starting with noisy data input, the model uses the UNet attention layer to extract features. A composite loss function is constructed by combining the mean squared error (Lmse) and the learned perceptual patch similarity (LPIPS). Finally, the orthogonal butterfly fine-tuning technique (BOFT) is used to achieve a targeted update of the attention layer parameters. The BOFT hyperparameters are set as follows: boft_block_num=8, boft_block_size=0.
[0047] The sketch-text-to-image generation model is optimized based on the total loss function until the loss function meets the preset requirements to obtain a trained sketch-text-to-image generation model.
[0048] Through the synergistic effect of the aforementioned technical features, sketch encoding and text embedding vectors are generated. ControlNet injects sketch features into the UNet residual block via a zero-convolution layer, while CLIP text embeddings are integrated into the decoding process via a cross-attention mechanism. The UNet network uses a DDIM sampler to perform progressive denoising under dual constraints, ultimately generating high-resolution images. This achieves simultaneous optimization of image generation quality and structural consistency under multimodal conditions, ensuring that the generated images meet structural requirements and effectively alleviate local distortion issues in the generated images.
[0049] Example 2 In one or more embodiments, a sketch-guided image generation system based on a stable diffusion model is disclosed, specifically comprising: A data acquisition module is configured to: acquire sketch data and text data; A feature extraction module is configured to: input the sketch data and text data into a constructed multimodal conditional input module, wherein the multimodal conditional input module includes an edge feature extraction module and a text encoder, wherein the edge feature extraction module extracts edge structural features of the sketch data, and the text encoder extracts text features of the text data to generate a semantic feature vector to obtain a multimodal feature; The image generation module is configured to generate a clear image by using the multimodal features through a diffusion process based on a pre-trained stable diffusion model.
[0050] Example 3 This embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the sketch-guided image generation method based on the stable diffusion model are completed.
[0051] Example 4 This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the sketch-guided image generation method based on the stable diffusion model are completed.
[0052] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0053] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0054] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide the functions for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0055] The descriptions of the various embodiments in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0056] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A sketch-guided image generation method based on a stable diffusion model, characterized in that: include: Obtain sketch data and text data; Inputting the sketch data and text data into a constructed multimodal conditional input module, the multimodal conditional input module comprising an edge feature extraction module and a text encoder, the edge feature extraction module extracting edge structural features of the sketch data, the text encoder extracting text features of the text data to generate a semantic feature vector, and obtaining a multimodal feature; Based on a pre-trained stable diffusion model, the multimodal features are diffused to generate a clear image.
2. The sketch-guided image generation method based on a stable diffusion model according to claim 1, characterized in that: The edge feature extraction module adopts a pre-trained ControlNet network structure, and the text encoder adopts a CLIP encoder to obtain multimodal features.
3. The sketch-guided image generation method based on a stable diffusion model according to claim 1, characterized in that: The pre-trained stable diffusion model adopts the UNet architecture of the pre-trained Stable Diffusion v1.5 model, including a variational autoencoder and a Unet module, wherein the variational autoencoder includes an encoder and a decoder. The encoder converts a real image into a first latent representation and adds Gaussian noise to the latent representation in the forward process to generate a noisy latent representation; the Unet module denoises the noisy latent representation to obtain a second latent representation, and the decoder upsamples the second latent representation to obtain a clear image.
4. The sketch-guided image generation method based on a stable diffusion model according to claim 1, wherein: The pre-training process of the stable diffusion model includes: Inputting the multimodal features into a UNet module to generate a first generated image, and using an image-to-sketch conversion module to inversely convert the first generated image into a first generated sketch; evaluating a conversion quality of the first generated sketch based on an image conversion quality evaluation module; Constructing a second optimization objective function based on perceptual loss according to the conversion quality, and weightedly fusing the first optimization objective function with the second optimization objective function to obtain a total loss function; Based on the total loss function, the orthogonal butterfly fine-tuning technology is used to update the parameters of the stable diffusion model to obtain a pre-trained stable diffusion model.
5. The sketch-guided image generation method based on a stable diffusion model according to claim 4, characterized in that: The step of inputting the multimodal features into the UNet module to generate the first generated image is specifically as follows: The real images in the training dataset are input into the VAE encoder for encoding and converted into a low-dimensional latent space representation. A random Gaussian noise matrix with the same dimension as the latent space is created, and the noise value follows a standard normal distribution. Gaussian noise is gradually added to the real images through the forward diffusion process to generate noisy latent space data as the training target; The UNet module iteratively denoises in the latent space to generate a latent representation that conforms to the text description and sketch structure, and the VAE decoder upsamples the latent representation to generate the first generated image.
6. The sketch-guided image generation method based on a stable diffusion model according to claim 4, characterized in that: The image-to-sketch conversion module is a pre-trained generative adversarial network, comprising a generator and a discriminator. The generator comprises a convolution block, a downsampling module, a residual module, an upsampling module, and an output layer connected in sequence. The convolution block performs feature extraction on the first generated image. The downsampling module uses two downsampling operations to compress the size and extract high-level features. The residual module enhances the feature representation. The upsampling module uses two upsampling operations to restore the image size. The output layer maps the features to the target image space to obtain the first generated sketch.
7. The sketch-guided image generation method based on a stable diffusion model according to claim 4, characterized in that: By calculating the structural similarity between the first generated sketch and the corresponding edge map in the multi-layer feature space, an optimization objective function based on perceptual loss is constructed: Where, is the perceptual similarity index; and Represents the feature vectors at the (h, w) position of the input image x and y in the feature map of the lth layer of the deep network respectively; x represents the sketch generated by the generative adversarial network; y represents the corresponding edge map loaded by the dataset; is the trainable channel weight vector of layer l; For Hadamard; and Represents the height and width of the l-th layer feature map.
8. A sketch-guided image generation system based on a stable diffusion model, characterized in that: include: A data acquisition module is configured to: acquire sketch data and text data; A feature extraction module is configured to: input the sketch data and text data into a constructed multimodal conditional input module, wherein the multimodal conditional input module includes an edge feature extraction module and a text encoder, wherein the edge feature extraction module extracts edge structural features of the sketch data, and the text encoder extracts text features of the text data to generate a semantic feature vector to obtain a multimodal feature; The image generation module is configured to generate a clear image by using the multimodal features through a diffusion process based on a pre-trained stable diffusion model.
9. An electronic device, characterized in that: The invention comprises a memory and a processor and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the sketch-guided image generation method based on the stable diffusion model according to any one of claims 1 to 7 is completed.
10. A computer-readable storage medium, characterized in that Used to store computer instructions, which, when executed by a processor, complete the sketch-guided image generation method based on a stable diffusion model as described in any one of claims 1 to 7.
Citation Information
Cited By
Diffusion model image restoration method based on regional mask and dynamic exit
CN121860872A
Image Restoration Method Based on Region Masking and Dynamic Exit Diffusion Model
CN121860872B