Model generation method and device

By improving the network structure and training strategy, and combining it with a conditionally controlled generative network, the problem of insufficient efficiency and quality of existing diffusion models in high-resolution image generation is solved, and a highly efficient high-definition image enhancement effect is achieved.

CN121883256APending Publication Date: 2026-04-17HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing image generation methods based on pre-trained diffusion models suffer from significantly reduced enhancement effects and increased inference time when processing high-resolution images above a certain resolution, making it difficult to effectively improve image quality and efficiency.

Method used

By improving the network structure and training strategy, an initial model is constructed and the sampling module is adjusted to adapt to high-resolution images. Combined with a conditional control generative network, the target generative model is trained to improve image enhancement capabilities.

Benefits of technology

It achieves effective enhancement of high-resolution images, improves generation efficiency and image quality, and can adapt to high-resolution input with a small amount of training data to generate high-quality, highly realistic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883256A_ABST
    Figure CN121883256A_ABST
Patent Text Reader

Abstract

The invention provides a model generation method and device, and relates to the technical field of image processing. The electronic device obtains the target generation model by improving the model network structure and the training strategy, and the target generation model can effectively enhance the high-resolution image. The method comprises the steps that an electronic device obtains an initial model, the initial model comprises an initial base model and an initial condition control generation network, and the initial condition control generation network is used for introducing a control signal for the initial base model; an electronic device constructs a plurality of image pairs including a first image and a second image. And the electronic equipment adjusts the output size of the sampling module of the initial model to obtain a to-be-trained model. The to-be-trained model comprises a to-be-trained model and a to-be-trained condition control generation network. And the electronic equipment uses the first image to train the to-be-trained base model into a fixed base model. And the electronic equipment uses the plurality of image pairs to train the to-be-trained model into a target generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a model generation method and apparatus. Background Technology

[0002] With the rapid development of computer vision and image processing technologies, high-definition image generation and enhancement technologies have been widely applied in various fields, such as film and television production, virtual reality, security monitoring, and social media. Among these applications, portrait enhancement is an important and challenging task. Portrait enhancement requires the generated images to be rich in detail and close to real-life quality, while also demanding that the algorithm have high computational efficiency and flexibility.

[0003] Currently, image generation methods based on pre-trained diffusion models, such as those based on Stable Diffusion, have made significant progress in the field of image generation. These methods generate high-quality images step-by-step through noise addition and denoising processes, demonstrating powerful generation capabilities and broad application potential. However, when applied to high-resolution image enhancement, these pre-trained diffusion model-based image generation methods are limited by the model's performance at specific resolutions, such as 512×512 or 1024×1024. When the input image resolution exceeds this limitation, the enhancement effect of the model is significantly reduced, and the inference time is greatly increased.

[0004] Therefore, how to effectively enhance high-resolution images has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a model generation method and apparatus, applicable to the field of terminal technology. In this application, the electronic device, through an improved network structure and training strategy, obtains a target generation model capable of effectively enhancing high-resolution images.

[0006] To achieve the above-mentioned technical objectives, the embodiments of this application provide the following technical solutions:

[0007] Firstly, a model generation method is provided, comprising: an electronic device acquiring an initial model, the initial model including an initial base model and an initial conditional control generative network, the initial base model being a pre-trained diffusion model of the U-Net architecture, and the initial conditional control generative network being a branch network of ControlNet, the initial conditional control generative network being used to introduce control signals to the initial base model. The electronic device constructs multiple image pairs, each image pair including a first image and a second image, the first image being a training image with a first resolution and a first sharpness, the second image being a training image with a first resolution and a second sharpness, the first image and the second image being training images based on the same scene but with different sharpnesses, and the first sharpness being higher than the second sharpness. The electronic device adjusts the output size of the sampling module of the initial model to match the first resolution, obtaining a model to be trained, the model to be trained including a base model to be trained and a conditional control generative network to be trained. The electronic device uses the first image to train the base model to be trained until the output value of the loss function of the base model to be trained is less than or equal to a preset threshold, obtaining a fixed base model. With the fixed base model fixed, the electronic device uses multiple image pairs to train the model to be trained until the output value of the loss function of the model to be trained is less than or equal to the preset threshold, obtaining a target generative model.

[0008] In some examples, the electronic device integrates an initial conditional control generative network into the initial base model to introduce control signals to the initial base model, thereby enhancing its generative capabilities. The electronic device constructs multiple image pairs based on a first resolution as training data for the model. Subsequently, the electronic device can use these multiple image pairs to train the model to be trained, enabling the model to learn how to augment a second image into a target image (the first image). By adjusting the output size of the sampling module of the model to be trained, the electronic device can enable the model to effectively process image pairs at the first resolution. By training the base model first, the electronic device can then fix the base model and train other parts of the model, thus improving training efficiency. With the fixed base model fixed, the electronic device trains the model to be trained using multiple image pairs. Training stops when the output value of the loss function of the model to be trained is less than or equal to a preset threshold, and the trained model is used as the target generative model.

[0009] Thus, the target generation model obtained by the electronic device based on the improved model network structure and training strategy described above can effectively enhance high-resolution images.

[0010] According to the first aspect, the initial base model includes multiple first initial encoder blocks, a first initial intermediate block, and multiple first initial decoder blocks. The first initial encoder blocks are used to downsample and extract features from the first initial image data to obtain a first initial feature map, which is then passed to the first initial intermediate block. The first initial image data is either a first image with added noise or data after an electronic device encodes the first image into a latent space representation and adds noise to the latent space representation corresponding to the first image. The first initial encoder block includes a first initial downsampling module, which is used to reduce the resolution of the first initial image data and extract features from it. The first initial intermediate block is located between the first initial encoder block and the first initial decoder block, and is used to extract features from the received first initial feature map to obtain a first initial intermediate feature map. The first initial decoder block is used to upsample the first initial intermediate feature map to generate a first output image. The first initial decoder block also includes a first upsampling module, which is used to upsample the first initial intermediate feature map.

[0011] According to the first aspect, or any implementation of the first aspect above, the initial condition control generator network includes multiple second initial encoder blocks, a second initial intermediate block, and multiple zero-convolutional layers. The second initial encoder blocks are used to downsample and extract features from the second initial image data input to the second initial encoder block to obtain a second initial feature map. The second initial encoder block is also used to pass the second initial feature map to the second initial intermediate block, and to pass the second initial feature map to the first initial decoder block of the initial base model through zero-convolutional layers. The second initial feature map is used to introduce control signals to the initial base model. The second initial image data is the data corresponding to an image pair after adding noise to the first image in an image pair, or the data corresponding to an image pair after an electronic device encodes an image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image in an image pair. The second initial intermediate block is used to extract features from the second initial feature map to obtain a second intermediate feature map, and to pass the second intermediate feature map to the first intermediate block of the initial base model through zero-convolutional layers. The second intermediate feature map is used to introduce control signals to the initial base model. The second initial encoder block also includes a second initial downsampling module, which is used to reduce the resolution of the second image data and extract the features of the second image data.

[0012] According to the first aspect, or any implementation of the first aspect above, the electronic device adjusts the sampling rate of the first initial downsampling module and the first initial upsampling module corresponding to the first initial downsampling module in the initial base model to match the first resolution. The electronic device adjusts the sampling rate of the second initial downsampling module in the initial condition control generation network to match the first resolution.

[0013] Thus, once the sampling rate of the sampling module of the model to be trained is matched with the first resolution by the electronic device, the model to be trained is able to accurately process image pairs at the first resolution.

[0014] According to the first aspect, or any implementation of the first aspect above, the output resolution of the first initial downsampling module, the first initial upsampling module and the second initial downsampling module are both the second resolution, which is smaller than the first resolution.

[0015] According to the first aspect, or any implementation of the first aspect above, the electronic device increases the downsampling ratio of the first initial downsampling module, the upsampling ratio of the first initial upsampling module, and the downsampling ratio of the second initial downsampling module according to the ratio of the first resolution and the second resolution, to obtain the first training downsampling module, the first training upsampling module, and the second training downsampling module.

[0016] In this way, by increasing the sampling rate of the initial model's sampling module according to the ratio of the first resolution to the second resolution, the electronic device can improve the resolution that the sampling module of the adjusted model to be trained can handle from the second resolution to the first resolution.

[0017] According to the first aspect, or any implementation of the first aspect above, the base model to be trained includes multiple first encoder blocks to be trained, multiple intermediate blocks to be trained, and multiple first decoder blocks to be trained. The first encoder block to be trained includes a first downsampling module to be trained, and the first decoder block to be trained includes a first upsampling module to be trained. The fixed base model includes multiple first fixed encoder blocks, multiple fixed intermediate blocks, and multiple first fixed decoder blocks. The conditionally controlled generative network to be trained includes multiple second encoder blocks to be trained, multiple intermediate blocks to be trained, and multiple zero-convolutional layers. The second encoder blocks to be trained include a second downsampling module to be trained.

[0018] According to the first aspect, or any implementation of the first aspect above, the electronic device inputs the first training image data into the first training encoder block of the training base model after N training iterations. The first training image data is either a first image with added noise or data after the electronic device encodes the first image into a latent space representation and adds noise to the latent space representation corresponding to the first image, where N is a positive integer. The electronic device uses the first training encoder block to downsample and extract features from the first training image data to obtain a first training feature map. The electronic device uses the first training intermediate block to further extract features from the first training feature map to obtain a first training intermediate feature map. The electronic device uses the first training decoder block to upsample the first training intermediate feature map and generate a second output image. Based on the second output image, the electronic device calculates the output value of the base model loss function of the training base model.

[0019] In this way, electronic devices can measure the difference between the second output image generated by the training model and the target image (first image) by calculating the output value of the base model loss function of the base model to be trained, thereby optimizing the performance of the training model.

[0020] According to the first aspect, or any implementation of the first aspect above, if the output value of the base model loss function is not greater than the preset threshold of the base model loss function, the electronic device will use the base model to be trained after N training iterations as the fixed base model.

[0021] According to the first aspect, or any implementation of the first aspect above, if the output value of the base model loss function is greater than the preset threshold of the base model loss function, the electronic device adjusts the parameters of the base model to be trained. The electronic device uses another first image to train the base model to be trained until the output value of the base model loss function is less than or equal to the preset threshold of the base model loss function, thus obtaining a fixed base model.

[0022] In some examples, when the output value of the base model loss function decreases to less than or equal to a preset threshold, the electronic device determines that the performance of the base model to be trained has reached the target level, and further training may not significantly improve its performance. At this point, the electronic device determines that the base model to be trained has completed sufficient training and stops training it. The electronic device then designates the trained base model as a fixed base model.

[0023] In this way, after the electronic device completes the training of the base model in the model to be trained, it can then focus on training the other parts of the model to be trained.

[0024] According to the first aspect, or any implementation of the first aspect above, the electronic device inputs the second training image data into the first fixed encoder block of the fixed base model. The second training image data is data after the electronic device adds noise to the first image in an image pair, or data after the electronic device encodes the first image in an image pair into a latent space representation and then adds noise to the latent space representation corresponding to the first image. The electronic device uses the first fixed encoder block in the fixed base model to downsample and extract features from the second training image data to obtain the second training feature map. The electronic device uses the second training encoder block in the training conditional control generator network, which has been trained M times, to downsample and extract features from the third training image data input to the second training encoder block to obtain the third training feature map. The third training image data includes the second training image data. The third initial image data is data corresponding to an image pair after adding noise to the first image in an image pair, or data corresponding to an image pair after the electronic device encodes an image pair into a latent space representation and then adds noise to the latent space representation corresponding to the first image in an image pair. M is a positive integer. The electronic device uses a second intermediate block to extract features from a third intermediate feature map, obtaining a second intermediate feature map. This second intermediate feature map is used to introduce control signals to the fixed-base model. The electronic device then passes the third intermediate feature map to the first fixed decoder block of the fixed-base model through a zero-convolutional layer. The electronic device uses the first fixed intermediate block to extract features from the second intermediate feature map based on the control signals provided by the second intermediate feature map, obtaining a target intermediate feature map. The electronic device uses the first fixed decoder block to upsample and decode the target intermediate feature map based on the control signals provided by the third intermediate feature map, generating a third output image. Based on the third output image, the electronic device calculates the output value of the loss function of the training model.

[0025] In this way, the electronic device trains the training model using a pair of images corresponding to a third training image, enabling the training model to learn how to reconstruct the target image (first image) from the second image. By calculating the output value of the training model's loss function, the electronic device can measure the difference between the third output image generated by the training model and the target image (first image), thereby optimizing the performance of the training model.

[0026] According to the first aspect, or any implementation of the first aspect above, if the output value of the loss function of the model to be trained is not greater than the preset threshold of the loss function to be trained, the electronic device will use the base model to be trained after M training cycles as the target generation model.

[0027] According to the first aspect, or any implementation of the first aspect above, if the output value of the loss function of the model to be trained is greater than the preset threshold of the loss function, the electronic device adjusts the parameters of the generative network under the control of the training conditions. The electronic device inputs another pair of images from multiple image pairs into the model to be trained and continues to train the model. The electronic device trains the model to be trained until the output value of the loss function of the model to be trained is less than or equal to the preset threshold of the loss function, thus obtaining the target generative model.

[0028] In this way, electronic devices can train a higher-performing model by minimizing the output value of the loss function corresponding to the model to be trained, thereby improving its performance in high-resolution image enhancement.

[0029] According to the first aspect, or any implementation of the first aspect above, the electronic device calculates the output value of the perceptual loss function based on the third output image. If the output value of the perceptual loss function is greater than the preset threshold of the perceptual loss function, the parameters of the network to be trained are adjusted.

[0030] In this way, the electronic device minimizes the feature differences between the visual features corresponding to the third output image and the target image through a perceptual loss function. During this process, the model to be trained can learn how to adjust the generated third output image so that its visual features are as close as possible to the visual features of the target image.

[0031] According to the first aspect, or any implementation of the first aspect above, the model to be trained further includes a discriminator, which is used to adversarially compete with a generator composed of a fixed-base model and a conditionally controlled generative network to be trained. Both the generator and the discriminator include adversarial loss functions. The electronic device fixes the parameters of the generator, inputs one image from a plurality of image pairs into the model to be trained, and obtains a fourth output image from a first fixed decoder block. Based on the fourth output image, the electronic device calculates the output value of the first adversarial loss function of the discriminator. If the output value of the first adversarial loss function is greater than a preset threshold of the first adversarial loss function, the electronic device adjusts the parameters of the discriminator. The electronic device inputs another image from a plurality of image pairs into the model to be trained, and trains the model until the output value of the first adversarial loss function is less than or equal to the preset threshold of the first adversarial loss function. The electronic device fixes the parameters of the discriminator, inputs another image from a plurality of image pairs into the model to be trained, and obtains a fifth output image from a first fixed decoder block. Based on the fifth output image, the electronic device calculates the output value of the second adversarial loss function of the generator. If the output value of the second adversarial loss function is greater than a preset threshold of the second adversarial loss function, the electronic device adjusts the parameters of the generator. The electronic device inputs another pair of images from multiple image pairs into the model to be trained, and trains the model until the output value of the second adversarial loss function is less than or equal to a preset threshold of the second adversarial loss function.

[0032] In some examples, the discriminator is used against a generator consisting of a fixed-base model and a conditionally controlled generative network to be trained. Electronic devices allow the generator and discriminator to compete against each other, minimizing the output values ​​of the first adversarial loss function corresponding to the discriminator and the first adversarial loss function corresponding to the generator, respectively, to improve the quality and realism of the images generated by the model under training.

[0033] In this way, the electronic device forces the generator and discriminator to compete with each other. The generator learns to produce images that are difficult for the discriminator to distinguish, while the discriminator continuously improves its ability to differentiate between generated and target images. Ultimately, this enables the model under training to generate high-quality, highly realistic images.

[0034] In a second aspect, an electronic device is provided. The electronic device includes one or more processors and a memory. The memory is coupled to the one or more processors and is used to store computer program code, which includes computer instructions. The one or more processors invoke the computer instructions to cause the electronic device to perform the method of the first aspect or any embodiment of the first aspect.

[0035] Thirdly, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program (also referred to as instructions or code) that, when executed by an electronic device, causes the electronic device to perform the method of the first aspect or any embodiment of the first aspect.

[0036] Fourthly, a computer program product is provided that, when run on an electronic device, causes the electronic device to perform the method of the first aspect or any one of the embodiments of the first aspect.

[0037] Fifthly, a circuit system is provided, the circuit system including processing circuitry configured to perform the method of the first aspect or any embodiment of the first aspect.

[0038] In a sixth aspect, a chip system is provided, including at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor, and when the at least one processor executes the instructions, the at least one processor performs the method of the first aspect or any embodiment of the first aspect.

[0039] The technical effects of the aforementioned aspects can be referenced from each other, and will not be elaborated further here. Attached Figure Description

[0040] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic diagram of the forward and reverse processes of the diffusion model provided in the embodiments of this application;

[0042] Figure 2 This is a schematic diagram of an initial model provided in an embodiment of this application;

[0043] Figure 3 This is a schematic flowchart of a model generation method provided in an embodiment of this application;

[0044] Figure 4 This is a schematic diagram of the structure of an initial base model provided in an embodiment of this application;

[0045] Figure 5 This is a schematic diagram of the structure of an initial condition control generation network provided in an embodiment of this application;

[0046] Figure 6 This is a schematic diagram of the model to be trained obtained after adjusting the sampling rate of the sampling module;

[0047] Figure 7 This is a flowchart illustrating a method for increasing the upsampling factor provided in an embodiment of this application;

[0048] Figure 8 This is a schematic diagram of the structure of a training model including a base model to be trained, provided in an embodiment of this application;

[0049] Figure 9 This is a schematic diagram of the structure of a training model including a fixed-base model provided in an embodiment of this application;

[0050] Figure 10 This is a schematic diagram of a method for training a base model to be trained in an electronic device, provided in an embodiment of this application;

[0051] Figure 11 This is a schematic diagram of a method for training a model to be trained using an electronic device, provided in an embodiment of this application.

[0052] Figure 12 This is a schematic diagram of a data processing flow for a training model including a fixed-base model, provided in an embodiment of this application.

[0053] Figure 13 This is a schematic diagram illustrating a method for training a model using a loss function in an electronic device, as provided in an embodiment of this application.

[0054] Figure 14 This is a schematic diagram of a method for training a model of an electronic device based on an adversarial loss function, provided in an embodiment of this application.

[0055] Figure 15 This is a schematic diagram of the hardware structure of the electronic device 100 provided in the embodiments of this application;

[0056] Figure 16 This is a software structure block diagram of the electronic device provided in the embodiments of this application;

[0057] Figure 17 This is a block diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0058] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the protection scope of this application.

[0059] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0060] Furthermore, in this application, directional terms such as "upper," "lower," "inner," and "outer" are defined relative to the indicated placement of the components in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the placement of the components in the accompanying drawings.

[0061] Current image generation methods based on diffusion models are limited by the specific resolutions adapted when training the diffusion model (such as 512×512 or 1024×1024). When the resolution of the input image to be processed is higher than this specific resolution (such as 2048×2048), the trained diffusion model cannot directly and effectively process the image to be processed. If high-resolution images need to be processed, a new diffusion model needs to be trained separately.

[0062] In this embodiment, the electronic device introduces a conditional control generation network to adapt the pre-trained diffusion model to the specific resolution, thereby forming an initial model. Then, the structure of the initial model is appropriately adjusted (e.g., changing the downsampling and upsampling ratios of the sampling module in the initial model) to obtain the model to be trained. The electronic device first trains the pre-trained diffusion model (hereinafter referred to as the initial base model) in the model to be trained to obtain a fixed base model; then, it fixes this fixed base model and trains the entire model to be trained to obtain the target generation model. This target generation model can adapt to high-resolution input and achieve effective image enhancement for high-resolution images using only a small amount of training data, exhibiting high generation efficiency.

[0063] To facilitate a clear description of the technical solutions in the embodiments of this application, the following is a brief introduction to some of the terms and technologies involved in the embodiments of this application:

[0064] Diffusion model:

[0065] Diffusion models are a type of generative model, primarily consisting of two processes: a forward process and a backward process. For example... Figure 1 This is a schematic diagram illustrating the forward and reverse processes of the diffusion model provided in the embodiments of this application. For example... Figure 1 As shown, the arrows from right to left illustrate a forward process. The forward process involves adding Gaussian noise to the real images in the dataset. The arrows from left to right illustrate a backward process. The backward process involves denoising the noisy images to reconstruct the real images.

[0066] The following describes the forward pass process. The forward pass involves adding Gaussian noise to the original image until it becomes a completely noisy image. In each step of the forward pass, the amount of noise added gradually increases, causing the original image to become increasingly blurry until it completely loses its original information. This process is fixed, and the noise addition follows predefined rules. In the forward pass, the original image contains data point x0. After adding Gaussian noise, data point x0 is transformed into data point X. t-1 Then, noise is added to the original image, data point X. t-1 Transform into data point X t Then, noise is added to the original image, data point X. t Transform into data point X T The probability distribution of data point x0 is q(x0). The forward process gradually destroys the data structure of the data points in the original image by repeatedly applying the following Markov diffusion kernel:

[0067]

[0068] Where t∈{1,2,…,T}, It is a predefined or learned variation of noise variance. Through... With proper design, q(x) is theoretically guaranteed. t It converges to a unit spherical Gaussian distribution. It is noteworthy that the marginal distribution at any time step t has the following analytical form:

[0069]

[0070] in

[0071] The following describes the reverse process. The reverse process is the opposite of the forward process. It starts with a purely noisy image and gradually reconstructs the original image by removing noise from it. The electronic device converts the data points X from the purely noisy image... T Gradually recover to data point X t X t-1、 This continues until X0 is recovered. However, directly solving for the conditional probability distribution of the reverse process is very difficult; therefore, a neural network is needed to approximate this distribution. The goal of the reverse process is to recover X0 from the data point x. t to data point x t-1 We study a transfer kernel, which is defined by the following Gaussian distribution:

[0072] Where θ is a learnable parameter. Using such a learned transition kernel, we can approximate the data distribution q(x0) using the following marginal distribution:

[0073] in

[0074] In this way, the diffusion model simulates gradually adding noise to the data distribution (forward process), and then learns the inverse process (backward process), that is, how to restore the original image from a purely noisy image. In this process, the diffusion model can learn image features and data distribution, thus possessing the ability to generate high-resolution images based on low-resolution images.

[0075] Unet architecture (U-shapedNetwork, U-Net):

[0076] The UNet architecture is a convolutional neural network structure widely used in image segmentation tasks. Its role is to learn how to remove noise from noisy images, thereby recovering the original image. During training, the UNet architecture receives a series of images with varying degrees of added noise as input and learns to predict the added noise in each image. Through iterative training, the UNet architecture learns how to remove noise from a purely noisy image, ultimately generating a clear original image. The UNet architecture is characterized by its unique encoder-decoder structure and skip-link design. Specifically:

[0077] Encoder: Used to extract features from an image. An encoder consists of multiple convolutional and pooling layers. Convolutional layers extract local features from the image by sliding convolutional kernels, while pooling layers reduce the spatial resolution of the feature maps to decrease computation. As the encoder progresses, the number of channels in the feature maps gradually increases, while the spatial resolution gradually decreases.

[0078] Decoder: Used to remap the features extracted by the encoder onto the original image. The decoder consists of multiple deconvolutional and convolutional layers. The deconvolutional layers are used to increase the spatial resolution of the feature map, while the convolutional layers are used to further extract and refine the features.

[0079] Skip links: These connect feature maps from the encoder to the decoder, combining low-level and high-level feature information. This design helps preserve image detail and improves image segmentation accuracy.

[0080] Pre-trained diffusion model:

[0081] A pre-trained diffusion model is a diffusion model that has been pre-trained on a large amount of image data, and has learned rich image features and data distribution.

[0082] Stable Diffusion:

[0083] Stable Diffusion is an efficient pre-trained diffusion model that has already learned how to generate high-quality images based on text prompts or other input conditions. Users do not need to train the entire model from scratch; instead, they can directly use the pre-trained weights for image generation tasks.

[0084] Pre-trained diffusion model based on Unet architecture:

[0085] In pre-trained diffusion models based on the Unet architecture, the Unet architecture is used as part of the generative model. The model first extracts features from the input image using the encoder within the Unet architecture, and then removes added noise using the decoder within the Unet architecture to reconstruct the original image or generate a new image. The skip-link design of the Unet architecture helps preserve image detail during decoding, resulting in more realistic and clearer generated images.

[0086] In this model, the pre-trained diffusion model based on the U-Net architecture uses noise estimates and latent space feature maps from the U-Net output to iteratively generate or predict images. Latent space feature maps refer to feature representations extracted from the intermediate layers (hidden layers) of the model; these feature maps contain high-level semantic information about the image. For example, the pre-trained diffusion model based on the U-Net architecture can use the U-Net architecture to extract noise estimates and latent space feature maps from noisy data points (state z). t Recover the intermediate data point state z t-1 Until the original data point state is recovered in t iterations.

[0087] ControlNet (Control Network):

[0088] ControlNet is a conditional control generative network used to control and regulate images or other data. For example, in diffusion models, ControlNet is used as a conditional control generative network to guide the image generation process of the diffusion model by learning features of specific conditions, thereby improving the quality and consistency of images generated by the diffusion model.

[0089] Variational Auto-Encoders (VAE-Encoders):

[0090] Variational autoencoders (VAEs) are generative models that introduce probabilistic methods into the autoencoder (AE) model, describing the latent space probabilistically. This results in unique advantages in data generation, compression, and feature extraction. A VAE primarily consists of two parts: an encoder and a decoder. The encoder maps the input data to a latent variable distribution, and the decoder then maps this latent variable back to the data space. A key feature of VAEs is that they constrain the distribution of the latent variable by minimizing the KL divergence, making it approximate a prior distribution (usually a standard normal distribution). This approach allows for more diverse and readily sampled samples generated by the model.

[0091] Perceptual Loss:

[0092] The perceptual loss function is a loss function used in image translation tasks, designed to capture high-level feature differences between images, rather than just pixel-level differences. It evaluates the quality of the generated image by comparing the differences between the generated and target images in a high-level feature space.

[0093] Adversarial Loss:

[0094] Adversarial loss is a core concept in Generative Adversarial Networks (GANs) used to train an adversarial process between the generator and the discriminator. By making the generator and discriminator compete with each other, adversarial loss improves the quality and realism of the generated images.

[0095] The following describes the initial model provided in the embodiments of this application.

[0096] Figure 2 This is a schematic diagram of an initial model provided in an embodiment of this application. For example... Figure 2 As shown, the left half of the overall model architecture is Stable Diffusion, whose parameters are frozen and not trained. The right half of the overall model architecture is ControlNet, whose parameters are trainable. It should be understood that Stable Diffusion is selected as the pre-trained diffusion model in the initial model as an example in this embodiment. Optionally, the pre-trained diffusion model in the initial model in this embodiment can also be other pre-trained diffusion models, and this embodiment does not limit this.

[0097] While Stable Diffusion boasts powerful image generation capabilities, it still exhibits limitations in certain complex scenes and for fine-grained control. To further enhance the generation capabilities of Stable Diffusion, ControlNet is introduced to provide control signals and improve the quality and consistency of the generated images. Introducing ControlNet into Stable Diffusion provides pre-trained diffusion models with greater control and flexibility, enabling them to achieve functions that are otherwise difficult to accomplish.

[0098] The following section introduces the structure and functions of Stable Diffusion.

[0099] Stable Diffusion includes encoder blocks SD Encoder Block_1, SD Encoder Block_2, SD Encoder Block_3, and SD Encoder Block_4. Stable Diffusion includes intermediate blocks SD Middle Block. Stable Diffusion includes decoder blocks SD Decoder Block_4, SD Decoder Block_3, SD Decoder Block_2, and SD Decoder Block_1.

[0100] In some examples, the inputs or outputs of Stable Diffusion and ControlNet can be connected to autoencoders, which can be variational autoencoders. A variational autoencoder is connected to the input of Stable Diffusion. The electronic device can use this variational autoencoder to encode the input image, converting it into a representation in a latent space. A variational autodecoder is connected to the output of Stable Diffusion. The electronic device can use this variational autodecoder to map the representation in the latent space back into the original data space, generating new samples that are similar to but different from the original data.

[0101] For example, after an image is input to the Stable Diffusion input, the encoder in the variational autoencoder encodes the image to obtain a latent space representation. The encoder also downsamples the image by a factor, for example, 8. Thus, by connecting the variational autoencoder to the Stable Diffusion input, the variational autoencoder can capture the key features of the input image and provide a compact, numerical image representation for subsequent processing steps. This accelerates the processing of the subsequent initial model.

[0102] In some examples, a variational autoencoder is connected to the input of ControlNet. The main function of this variational autoencoder is to refine the features of the image input to ControlNet and then use ControlNet to generate more specific conditional features or control signals. In this way, by connecting a variational autoencoder to the input of ControlNet, efficient image encoding and feature extraction can be performed, so that these conditional features can be used during the generation process to achieve higher quality image reconstruction and enhancement.

[0103] In this way, connecting an electronic device to an autoencoder in Stable Diffusion and ControlNet can reduce the computational load when the initial model processes the input image, thereby improving the efficiency of generating high-quality output images.

[0104] The following section describes the functions of each module in Stable Diffusion by examining the image processing process.

[0105] In some examples, the input image is a high-resolution (HR) image. This HR image refers to the image to be processed with a resolution higher than the specific resolution required for Stable Diffusion adaptation. Optionally, the electronics encode the input image (HR image) into a latent spatial representation using a variational autoencoder before feeding it into Stable Diffusion. This latent spatial representation contains key information about the image but removes unnecessary details. "Add Gaussian Noise" refers to the electronics adding noise to this latent spatial representation to enhance the robustness and generalization ability of the initial model. The electronics can extract local features, such as edges and textures, from this latent spatial representation using convolutional layers before feeding it into the encoder block; these features form the basis for subsequent processing.

[0106] In some examples, the electronic device progressively downsamples and extracts features from the noisy latent spatial representation using multiple encoder blocks in Stable Diffusion, resulting in a feature map. Each encoder block includes a downsampling module, which extracts features from the latent spatial representation and reduces the image size within it. The image size in the latent spatial representation is gradually reduced through the progressive processing of multiple encoder blocks. For example... Figure 2 As shown, the latent space represents the image sizes after processing by encoder blocks SD Encoder Blcok_1, SD Encoder Blcok_2, SD Encoder Blcok_3, and SD Encoder Blcok_4, which are 64×64, 32×32, 16×16, and 8×8, respectively.

[0107] In some examples, after processing the latent space representation using the encoder block, the electronic device passes the resulting feature map to an intermediate block. The electronic device then uses the intermediate block to perform feature extraction and fusion processing on the received feature map. Finally, the electronic device inputs the feature map processed by the intermediate block into the decoder block. The decoder block can then combine the features extracted by the encoder to denoise and reconstruct the received feature map.

[0108] In some examples, each decoder block includes an upsampling module, and multiple decoder blocks can progressively recover the size of the feature map received by the decoder. For example... Figure 2 As shown, the decoder blocks SD Decoder Blcok_4, SD Decoder Blcok_3, SD Decoder Blcok_2, and SD Decoder Blcok_1 progressively restore the feature map size from 8×8 to 16×16, 16×16, and 64×64.

[0109] Finally, the electronic device uses a variational autodecoder to decode the feature map obtained from the decoder block back into the image space.

[0110] The following describes the structure of each module in the Conditional Control Generative Network (ControlNet) in the initial model.

[0111] In some examples, after Stable Diffusion is trained, the electronic device fixes the parameters of Stable Diffusion and trains an initial model. The initial model includes the trained Stable Diffusion and ControlNet. The ControlNet structure includes encoder blocks SD Encoder Block_1 (trainable copy), SD Encoder Block_2 (trainable copy), SD Encoder Block_3 (trainable copy), SD Encoder Block_4 (trainable copy), and intermediate blocks SD Middle Block (trainable copy). Here, "trainable copy" means that the encoder blocks and intermediate blocks in ControlNet are copied from the corresponding encoder blocks and intermediate blocks trained in Stable Diffusion. The encoder blocks and intermediate blocks in ControlNet are trainable within the ControlNet model and can be further trained and optimized.

[0112] The following section describes the functions of each module in ControlNet based on the data processing procedure for low-resolution images using the initial model.

[0113] In some examples, the electronic device uses a variational autoencoder to encode a low-resolution image into a latent spatial representation. Then, the electronic device uses a zero-convolution layer to integrate the latent spatial representation encoded from the low-resolution image and the latent spatial representation encoded from the noisy high-resolution image. In some examples, the electronic device can achieve seamless integration of these two latent spatial representations through the zero-convolution layer.

[0114] In some examples, the electronic device inputs the integrated latent spatial representation into an encoder block in ControlNet for downsampling and feature extraction to obtain a conditional feature map. The encoder block in ControlNet includes a downsampling module that extracts features from the latent spatial representation and reduces the image size within the latent spatial representation. The image size in the latent spatial representation is gradually reduced through progressive processing across multiple encoder blocks. For example... Figure 2 As shown, the latent space represents images with dimensions of 64×64, 32×32, 16×16, and 8×8 after processing by the encoder blocks SD Encoder Block_1 (trainable copy), SD Encoder Block_2 (trainable copy), SD Encoder Block_3 (trainable copy), and SD Encoder Block_4 (trainable copy) in ControlNet. The electronic device uses intermediate blocks in ControlNet to fuse the conditional feature maps. The electronic device then passes the conditional features extracted by the encoder blocks and intermediate blocks in ControlNet to the corresponding intermediate and decoder blocks in Stable Diffusion for fusion with the feature maps in the corresponding intermediate and decoder blocks in Stable Diffusion, guiding the image generation process of Stable Diffusion.

[0115] In some examples, both the initial model's Stable Diffusion and ControlNet incorporate text input (Prompt) and time input. The initial model uses a text encoder to encode the user's text input into a numerical form that the model can understand and process, typically called an embedding. This text encoder can be the CLIP text encoder. The initial model also uses a time encoder to encode the user's time input into a text embedding. This allows the electronic device to provide additional conditional information to the initial model by introducing text and time inputs, thereby enhancing the model's control capabilities and improving the accuracy of the images it generates.

[0116] The model generation method provided in the embodiments of this application is described below with reference to the accompanying drawings and the foregoing content.

[0117] For a detailed implementation of the model generation method provided in this application, please refer to the descriptions in the various embodiments below.

[0118] For example, Figure 3 This is a flowchart illustrating a model generation method provided in an embodiment of this application, which is applied to electronic devices. It should be noted that this method does not rely on... Figure 3 The specific order described below is a limitation. It should be understood that in other embodiments, the order of some steps in the method can be interchanged according to actual needs, or some steps can be omitted or deleted. The method includes the following steps S301-S305.

[0119] S301. Electronic equipment acquires initial model.

[0120] In some examples, the initial model includes an initial base model and an initial conditional control generator network. The initial base model is a pre-trained diffusion model of the U-Net architecture, and the initial conditional control generator network is a branch of ControlNet. The initial conditional control generator network is used to introduce control signals to the initial base model.

[0121] For example, such as Figure 2 As shown, the left half of the overall model architecture is a pre-trained diffusion model based on the U-Net architecture. Optionally, this pre-trained diffusion model is Stable Diffusion. The right half of the overall model architecture is the initial conditional control generator network, which can optionally be a branch of ControlNet. (The above text...) Figure 2The corresponding embodiments describe the specific structure of the initial model provided in the embodiments of this application, which will not be repeated here.

[0122] S302, Electronic device constructs multiple image pairs.

[0123] In some examples, since the goal of model training in this embodiment is to obtain a target generation model that can effectively enhance high-resolution images, the electronic device determines a first resolution based on this goal and constructs multiple image pairs as training data for the model based on the first resolution. For example, the initial base model is limited to enhancing images with a resolution of 512×512 or less. The goal of model generation in this embodiment is to obtain a target generation model that can enhance images with a resolution of 1024×1024. Therefore, the electronic device determines 1024×1024 resolution as the first resolution and constructs multiple image pairs based on the first resolution of 1024×1024.

[0124] In some examples, since the goal of model generation is to enable the model to enhance images of a first resolution and a second sharpness to an image of the same first resolution and a first sharpness, each image pair constructed by the electronic device includes a first image and a second image. The first image is a training image with the first resolution and the first sharpness. The second image is a training image with the first resolution and the second sharpness. The first image and the second image are training images of the same scene but with different sharpnesses, and the first sharpness is higher than the second sharpness. In some examples, the electronic device can obtain the first image from an open-source high-definition database or from a high-definition camera; this application embodiment does not limit this. The format of the first image can be: portable network graphics (PNG), bitmap (BMP), etc.; this application embodiment does not limit this. The electronic device can obtain a second image based on the first image; the difference between the second image and the first image is that the sharpness of the second image is lower than that of the first image.

[0125] In this way, subsequent electronic devices can use the first and second images to train the model to be trained, so that the model can learn how to enhance the clarity of the second image to the clarity of the first image.

[0126] S303. The electronic device adjusts the output size of the sampling module of the initial model to match the first resolution, thus obtaining the model to be trained.

[0127] In some examples, the model to be trained consists of two components: a base model to be trained and a conditionally controlled generative network to be trained.

[0128] In some examples, the initial model is built from a pre-trained diffusion model based on the U-Net architecture and a ControlNet branch network. The U-Net-based pre-trained diffusion model is limited by a specific resolution and cannot effectively augment a first image at a resolution greater than that specific resolution. For example, downsampling aims to reduce the image resolution to extract higher-level features or reduce computation. However, if the output size of the sampling module does not match the resolution of the first image to be processed, it may result in the loss or distortion of information in the first image. Similarly, upsampling aims to restore the image resolution for display on higher-resolution devices or for more refined analysis. If the output size of the upsampling module does not match the desired output resolution, it may result in blurred, distorted, or pixelated images.

[0129] In other examples, the model's performance largely depends on its ability to accurately process the input image and generate the desired output image. If the output size of the sampling module does not match the first resolution of the first image to be processed, the model may fail to effectively learn the features in the input first image, leading to a performance degradation. During the model inference phase, if the sampling magnification and the first resolution of the first image to be processed do not match, the electronic device may need to perform additional preprocessing or postprocessing steps on the input first image to match the initial model's sampling magnification. This increases the complexity and computational cost of inference. The encoder blocks and intermediate blocks of the ControlNet branch network are copied from the encoder blocks and intermediate blocks of the pre-trained diffusion model of the U-Net architecture, and they also suffer from the problem of being unable to effectively process image pairs at the first resolution.

[0130] In this way, the electronic device adjusts the output size of the sampling module of the initial model to match the first resolution, thus obtaining the model to be trained. After the sampling ratio of the sampling module of the model to be trained matches the first resolution, the model to be trained is able to accurately process image pairs at the first resolution.

[0131] In some embodiments, the electronic device adjusts the output size of the sampling module of the initial model to match the first resolution, including: adjusting the sampling rate of the first initial downsampling module and the first initial upsampling module corresponding to the first initial downsampling module in the initial base model to match the first resolution; and adjusting the sampling rate of the second initial downsampling module in the initial conditional control generation network to match the first resolution.

[0132] In some examples, to facilitate understanding of the position and function of the first initial downsampling module, the first initial upsampling module, and the second initial downsampling module in the initial model, the following is combined with... Figure 4 and Figure 6The structure and function of the initial base model in the embodiments of this application are described.

[0133] Figure 4 This is a schematic diagram of the structure of an initial base model provided in an embodiment of this application. For example, as shown... Figure 4 As shown, the initial base model includes multiple first initial encoder blocks, multiple first initial intermediate blocks, and multiple first initial decoder blocks.

[0134] In some examples, the first initial encoder block is used to downsample and extract features from the first initial image data to obtain a first initial feature map, which is then passed to the first initial intermediate block. The first initial image data is either a first image with added noise or data where noise is added to the latent space representation corresponding to the first image after the electronic device encodes the first image into a latent space representation.

[0135] For example, with Figure 2 Taking the initial model shown as an example, multiple first initial encoder blocks correspond to... Figure 2 The left half of the model shows SD Encoder Blocks 1, 2, 3, and 4. The first initial image data can be data after the variational autoencoder (AutoEncoder-Encoder) encodes the first image (HR Image) into a latent space representation, and then Gaussian noise is added to the latent space representation corresponding to the first image. Optionally, the initial model can pass the first initial image data to the first initial encoder block through a convolutional layer. The convolutional layer can perform processing on the first initial image data, such as feature extraction, enhancing robustness to noise, filtering noise, enhancing features, and reducing data dimensionality, thereby promoting the effective processing and accurate recognition of the first initial image data by the initial model.

[0136] In some examples, the first initial encoder block includes a first initial downsampling module, which is used to reduce the resolution of the first initial image data and extract features from the first initial image data. In some examples, a first initial intermediate block is located between the first initial encoder block and the first initial decoder block, and the first initial intermediate block is used to extract features from the received first initial feature map to obtain a first initial intermediate feature map. For example, using... Figure 2 Taking the initial model shown as an example, the first initial intermediate block corresponds to... Figure 2 The left half of the model shows the SD Middle Block.

[0137] In some examples, the first initial decoder block is used to upsample the first initial intermediate feature map to generate the first output image. For example, using... Figure 2 Taking the initial model shown as an example, multiple first initial decoder blocks correspond to... Figure 2 The left half of the model shows SD Decoder Blocks 4, 3, 2, and 1. In some examples, the first initial decoder block also includes a first upsampling module, which is used to upsample the first initial intermediate feature map.

[0138] Figure 5 This is a schematic diagram of the structure of an initial condition control generation network provided in an embodiment of this application. For example, as shown... Figure 5 As shown, the initial condition control generator network includes multiple second initial encoder blocks, a second initial intermediate block, and multiple zero convolutional layers.

[0139] In some embodiments, the second initial encoder block is used to downsample and extract features from the second initial image data input to the second initial encoder block to obtain a second initial feature map. The second initial encoder block is also used to pass the second initial feature map to a second initial intermediate block, and to pass the second initial feature map to a first initial decoder block of the initial base model through a zero convolutional layer. The second initial feature map is used to introduce control signals to the initial base model.

[0140] In some embodiments, the second initial image data is the data corresponding to an image pair after adding noise to the first image in an image pair, or the data corresponding to an image pair after the electronic device encodes an image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image in the image pair.

[0141] For example, with Figure 2 Taking the initial model shown as an example, the second initial encoder block corresponds to... Figure 2 The right half of the model shows SD Encoder Blcok_1 (trainable copy), SD Encoder Blcok_2 (trainable copy), SD Encoder Blcok_3 (trainable copy), and SD Encoder Blcok_4 (trainable copy).

[0142] In some embodiments, the second initial intermediate block is used to extract features from the second initial feature map to obtain a second intermediate feature map. The second initial intermediate block is also used to pass the second intermediate feature map to a first intermediate block of the initial base model through a zero-convolutional layer. The second intermediate feature map is used to introduce control signals into the initial base model.

[0143] For example, with Figure 2 Taking the initial model shown as an example, the second initial encoder block corresponds to... Figure 2 The right half of the initial model shows the SD Middle Block (trainable copy).

[0144] In some embodiments, the second initial encoder block further includes a second initial downsampling module, which is used to reduce the resolution of the second image data and extract features from the second image data.

[0145] In some embodiments, the output resolutions of the first initial downsampling module, the first initial upsampling module, and the second initial downsampling module are all second resolutions, which are smaller than the first resolution. To enable the initial base model to process the first image at the first resolution, the electronic device adjusts the sampling rate of the first initial downsampling module to match the first resolution. Simultaneously, since the sampling rate of the first initial upsampling module corresponding to the first initial downsampling module also needs to match the first resolution for the initial base model to output an image at the first resolution, the electronic device also needs to adjust the sampling rate of the first initial upsampling module in the initial base model to match the first resolution.

[0146] In some examples, the second initial downsampling module in the initial condition-controlled generative network is copied from the first initial downsampling module in the initial base model; therefore, the sampling rate of the second initial downsampling module is the same as that of the first initial downsampling module. To ensure that the initial condition-controlled generative network can process images at the first resolution, and also to ensure the consistency and accuracy of the model being trained when processing data, the electronic device also needs to adjust the sampling rate of the second initial downsampling module to match the first resolution.

[0147] In this way, the electronic device adjusts the sampling rate of the first initial downsampling module and the sampling rate of the first initial upsampling module corresponding to the first initial downsampling module, as well as the sampling rate of the second initial downsampling module, to match the first resolution. This ensures that the model to be trained can effectively process images at the first resolution, while also ensuring the consistency and accuracy of the model when processing data.

[0148] In some embodiments, the electronic device increases the downsampling ratio of the first initial downsampling module, the upsampling ratio of the first initial upsampling module, and the downsampling ratio of the second initial downsampling module according to the ratio of the first resolution and the second resolution, to obtain the first training downsampling module, the first training upsampling module, and the second training downsampling module.

[0149] In some examples, the second resolution is a specific resolution adapted to be processed by the sampling module in the initial model. In this embodiment, the electronic device needs to adjust the resolution of the sampling module of the initial model so that the resolution of the image that the sampling module of the adjusted model to be trained can process is increased to the first resolution. Therefore, the electronic device increases the downsampling ratio of the first initial downsampling module, the upsampling ratio of the first initial upsampling module, and the downsampling ratio of the second initial downsampling module according to the ratio of the first resolution to the second resolution.

[0150] In this way, by increasing the sampling rate of the initial model's sampling module according to the ratio of the first resolution to the second resolution, the electronic device can improve the resolution of the image processed by the sampling module of the adjusted model to be trained from the second resolution to the first resolution.

[0151] For example, the second resolution of the image processed by the initial model adaptation is 512×512. The first resolution of the image processed by the model to be trained is 1024×1024. The ratio of the first resolution to the second resolution is 4. The electronic device increases the sampling rate of the sampling module of the initial model according to this ratio of 4 to obtain the model to be trained.

[0152] For example, Figure 2 The sampling rate of the sampling module in the initial model shown is the sampling rate before adjustment. Figure 6 This illustrates a training model obtained after adjusting the sampling rate of the sampling module. For example... Figure 2 As shown, the output resolution of the first initial encoder block SD Encoder Block_1 and the first initial decoder block SD Decoder Block_1 in the initial model is 64×64, which means that the sampling rate of the sampling modules included in the first initial encoder block SD Encoder Block_1 and the first initial decoder block SD Decoder Block_1 is 64×64. Figure 6The output resolution of the second initial encoder block (SD Encoder Block_1) and the second initial decoder block (SD Decoder Block_1) in the model to be trained shown is 128×128. This means that the sampling ratio of the sampling modules included in the second initial encoder block (SD Encoder Block_1) and the second initial encoder block (SD Decoder Block_1) in the model to be trained is 128×128. That is, after increasing the downsampling ratio of the first initial downsampling module (64×64), the upsampling ratio of the first initial upsampling module (64×64), and the downsampling ratio of the second initial downsampling module (64×64) according to the ratio of the first resolution to the second resolution (4), the resulting sampling ratios of the first downsampling module to be trained are 128×128, the first upsampling module to be trained is 128×128, and the second downsampling module to be trained is 128×128. The sampling ratio 128×128 is four times the sampling ratio 64×64.

[0153] In some examples, the first initial downsampling module includes at least one sampling module cascaded together, the first initial upsampling module includes at least one upsampling module cascaded together, and the second initial downsampling module includes multiple downsampling modules cascaded together. Since the sampling module with the highest output resolution is usually one of the most computationally intensive modules, adjusting the sampling rate of this module can optimize the model's computational efficiency while maintaining output quality. Therefore, the electronic device adjusts the sampling rate of the sampling module with the highest output resolution among the aforementioned cascaded sampling modules.

[0154] For example, such as Figure 6 As shown, the electronic device adjusts the downsampling rate of DownSample 1, the downsampling module with the highest output resolution in the first initial upsampling module corresponding to the first initial encoder block SD Encoder Block_1. The electronic device also adjusts the upsampling rate of Upsample 1, the upsampling module with the highest output resolution in the first initial upsampling module corresponding to the first initial decoder block SD Decoder Block_1. Finally, the electronic device adjusts the upsampling rate of Upsample 1, the upsampling module with the highest output resolution in the second initial downsampling module corresponding to the second initial encoder block SD Decoder Block_1.

[0155] In some examples, the electronic device in this application embodiment can increase the sampling rate of the sampling module in a variety of ways. For example, the electronic device can increase the sampling rate of the sampling module by adjusting ordinary convolution parameters, using dilated convolution, using pooling layers, and interpolation methods, or it can combine the above methods to increase the sampling rate of the sampling module.

[0156] For example, an electronic device can increase the sampling rate of the sampling module by setting parameters such as kernel size (k), stride (s), and padding (p) according to the following formula.

[0157]

[0158] For example, the upsampling module UpSample 1, using a 3×3 convolutional kernel with a stride of 2 and padding of 1, can achieve a 2x upsampling. To enable the trained model to handle image inputs with a resolution of 1024×1024, the electronic device needs to increase the upsampling factor of UpSample 1 from 2 to 4. Optionally, the electronic device can adopt... Figure 7 The method increases the upsampling factor of the upsampling module UpSample1 from 2 to 4.

[0159] For example, Figure 7 This is a flowchart illustrating a method for increasing the upsampling factor provided in an embodiment of this application. It should be noted that this method does not rely on... Figure 7 The specific order described below is a limitation. It should be understood that in other embodiments, the order of some steps in the method can be interchanged according to actual needs, or some steps can be omitted or deleted. For example... Figure 7 As shown, the method includes the following steps S3031 to S3033:

[0160] S3031, Electronic device sets the convolution kernel size of the upsampling module UpSample1.

[0161] In some examples, the electronic device determines the kernel size of the upsampling module UpSample 1 based on the input size and target output size. The electronic device then sets the kernel according to this kernel size. For example, if the electronic device determines the kernel size k to be 5, it sets the kernel according to this kernel size k = 5.

[0162] S3032, The electronic device sets the stride s of the convolution kernel of the upsampling module UpSample1.

[0163] In some examples, the stride is the distance the convolutional kernel slides across the input feature map. The electronic device determines this stride based on the input size and the target output size. For example, to enable the upsampling module UpSample1 to achieve a 4x upsampling, the electronic device can set the stride 's' of the convolutional kernel to a value greater than 1. For instance, the electronic device could set this stride 's' to 4.

[0164] S3033, Electronic device sets the padding p of the upsampling module UpSample1.

[0165] In some examples, padding is the process by which the model being trained adds pixel values ​​around the boundaries of the input feature map to increase its size. Electronic devices can make the output feature map reach the target output size by adding sufficient padding to the input feature map. For example, to enable the upsampling module UpSample 1 to achieve a 4x upsampling, the electronic device sets the padding p to 1, making the target output size after the convolution operation four times the input size.

[0166] In this way, by setting the kernel size to 5, the stride to 4, and the padding to 1, the electronic device can increase the upsampling factor of the upsampling module UPSample Block from 2 to 4.

[0167] In other examples, the electronic device can also achieve a higher downsampling rate by replacing ordinary convolutions in the sampling module with dilated convolutions, based on the following formula. For example, the electronic device can achieve 4x downsampling by setting the kernel size k of the upsampling module UpSample1 to 3, the dilation rate d to 2, the padding p to 2, and the stride s to 4.

[0168]

[0169] In the above example, if the resolution of the image to be enhanced by the model to be trained is other than 1024×1024, the electronic device can adjust the sampling rate of the sampling module according to the corresponding scaling factor. For example, if the resolution of the image to be enhanced by the model to be trained is not 1024×1024, but 2048×2048, the electronic device can enable the model to be trained to process images with a resolution of 2048×2048 by increasing the downsampling rate of the downsampling module DownSample 1 to 8 and the upsampling rate of the upsampling module UpSample 1 to 8.

[0170] In this way, the electronic device obtains the model to be trained by adjusting the sampling rate of the sampling module in the initial model. This ensures that the obtained model to be trained can effectively process images at the first resolution. At the same time, increasing the sampling rate of the sampling module in the model to be trained increases the extent to which the feature map size of the model can be reduced, thus significantly reducing the computational load and computation time, thereby improving training and inference efficiency.

[0171] S304. The electronic device uses the first image to train the base model to be trained until the output value of the loss function of the base model to be trained is less than or equal to a preset threshold, thereby obtaining a fixed base model.

[0172] In some examples, the fixed-base model exhibits stable generative capabilities during subsequent training. To avoid overfitting or training instability between the base model and the conditional generation network during model training, the electronic device first trains the base model. Once the base model demonstrates stable generative capabilities, it is designated as the fixed-base model. The electronic device then trains the current model, allowing it to learn more quickly how to adjust the generation process based on the control signals generated by the conditional generation network.

[0173] In some examples, electronic devices use a first image to train a base model to be trained until the output value of the loss function of the base model to be trained is less than or equal to a preset threshold, thereby enabling the base model to be trained to have a stable generation capability, and the trained base model is determined as a fixed base model.

[0174] To better understand the training process of the base model to be trained in the embodiments of this application, the following is combined with... Figure 8 The introduction describes the structure of the trainable model, including the base model to be trained, and combines it with... Figure 9 This section introduces the structure of the trainable model, including the fixed-base model, and combines... Figure 10 This paper introduces methods for training base models.

[0175] For example, Figure 8 This is a schematic diagram of the structure of a trainable model, including a base model to be trained, provided in an embodiment of this application. For example... Figure 8 As shown, the model to be trained includes a base model to be trained and a conditionally controlled generative network to be trained. The base model to be trained includes multiple first encoder blocks to be trained, first intermediate blocks to be trained, and multiple first decoder blocks to be trained. Each first encoder block to be trained includes a first downsampling module to be trained, and each first decoder block to be trained includes a first upsampling module to be trained. The conditionally controlled generative network to be trained includes multiple second encoder blocks to be trained, second intermediate blocks to be trained, and multiple zero-convolutional layers. Each second encoder block to be trained includes a second downsampling module to be trained.

[0176] For example, with Figure 6 Taking the model to be trained as an example, multiple first encoder blocks to be trained correspond to Figure 6 The left half of the model shows SD Encoder Blocks 1, 2, 3, and 4. The first encoder block to be trained corresponds to... Figure 6 The left half of the model shows the SD MiddleBlok. Multiple first-stage decoder blocks to be trained correspond to... Figure 6 The left half of the model shows SD Decoder Blcok_4, SD Decoder Blcok_3, SD Decoder Blcok_2, and SD Decoder Blcok_1.

[0177] For example, Figure 9 This is a schematic diagram of the structure of a training model including a fixed-base model, provided in an embodiment of this application. For example... Figure 9 As shown, the model to be trained includes a fixed-base model and a conditional control generative network to be trained. The fixed-base model is based on the electronic device's... Figure 8 The training base model shown is obtained by training. The fixed base model includes multiple first fixed encoder blocks, multiple first fixed intermediate blocks, and multiple first fixed decoder blocks. Figure 9 The model to be trained shown still uses... Figure 8 The conditional generator network to be trained is shown. The parameters in the multiple second encoder blocks, the second intermediate blocks, and the multiple zero convolutional layers in the conditional generator network to be trained remain unchanged.

[0178] In some embodiments, the electronic device uses a first image to train a base model to obtain a fixed base model. The following is an example... Figure 10 This method is described. For example, Figure 10 This is a schematic diagram illustrating a method for training a base model on an electronic device. It should be noted that this method does not rely on... Figure 10 The specific order described below is a limitation. It should be understood that in other embodiments, the order of some steps in the method can be interchanged according to actual needs, or some steps can be omitted or deleted. Figure 10 As shown, the method includes the following steps S1001 to S1006:

[0179] S1001: The electronic device inputs the first image data to be trained into the first encoder block of the base model to be trained after N training iterations.

[0180] In some examples, the first training image data is a first image with added noise or data where the electronic device encodes the first image into a latent space representation and then adds noise to the latent space representation corresponding to the first image. In some examples, N is a positive integer. The image generation capability of the training base model after N training iterations can be obtained in step S1005 by calculating the output value of the loss function.

[0181] For example, such as Figure 6As shown, the first training image data can be the data after the Variational Autoencoder (VAE-Encoder) encodes the first image (HR Image) into a latent space representation, and then adds Gaussian noise to the latent space representation corresponding to the first image. Optionally, the electronic device can input the first initial image data into the first initial encoder block through a convolutional layer. The convolutional layer can perform processing on the first initial image data, such as feature extraction, enhancing noise robustness, filtering noise and enhancing features, and reducing data dimensionality, thereby promoting the accurate recognition and rapid processing of the first initial image data by the first initial encoder block.

[0182] S1002: The electronic device uses the first encoder block to be trained to downsample and extract features from the first image data to be trained, and obtains the first feature map to be trained.

[0183] In some examples, the first encoder block to be trained learns to process a first image with a first resolution and a first sharpness during the process of downsampling and feature extraction of the first image data to be trained. The first feature map to be trained includes features extracted from the first image by the first encoder block to be trained.

[0184] S1003: The electronic device uses the first intermediate block to be trained to extract features from the first feature map to be trained, and obtains the first intermediate feature map to be trained.

[0185] In some examples, the first training intermediate feature map includes features extracted from the first training intermediate block by the first training feature map.

[0186] S1004: The electronic device uses the first training decoder block to upsample the first training intermediate feature map and then generates the second output image.

[0187] In some examples, the first trainable decoder block recovers the second output image from the latent spatial representation using upsampling of the first trainable intermediate feature map. During this process, the first trainable decoder learns how to generate the second output image with higher clarity and more detail. The noise added to the first trainable image data enhances the robustness of the first trainable decoder block to noise. In this process, the first trainable decoder learns how to better remove the noise added to the first trainable image data, thus generating a cleaner second output image.

[0188] S1005: The electronic device calculates the output value of the base model loss function of the base model to be trained based on the second output image.

[0189] In some examples, the electronic device measures the difference between the second output image generated by the base model and the first image by calculating the output value of the base model loss function of the base model to be trained. This difference reflects the image enhancement capability of the model to be trained. Once this output value reaches a preset threshold, the electronic device considers the model to be trained to have achieved its training objective.

[0190] In some examples, the base model loss function can be the mean squared error loss function, the output value of the base model loss function can be the output value of the mean squared error loss function, and the preset threshold of the base model loss function can be the preset threshold of the mean squared error loss function. Electronic devices use the mean squared error loss function to constrain the base model under training, which can quantify the prediction error of the base model under training, optimize the parameters of the base model under training, improve the image enhancement capability of the base model under training, and evaluate the stability of the base model under training.

[0191] S1006: The electronic device trains the base model to be trained until the output value of the base model loss function is less than or equal to the preset threshold of the base model loss function, thus obtaining a fixed base model.

[0192] In some examples, when the output value of the base model loss function decreases to less than or equal to a preset threshold, the electronic device determines that the performance of the base model to be trained has reached the target level, and further training may not significantly improve its performance. At this point, the electronic device determines that the base model to be trained has completed sufficient training and stops training it. The electronic device then designates the trained base model as a fixed base model. That is, if the output value of the base model loss function is not greater than the preset threshold, the electronic device uses the base model to be trained after N iterations as a fixed base model.

[0193] In some examples, if the output value of the base model loss function is greater than a preset threshold, the electronic device adjusts the parameters of the base model to be trained. The electronic device trains the base model to be trained using another first image until the output value of the base model loss function is less than or equal to the preset threshold, thus obtaining a fixed base model.

[0194] S305. With the fixed base model fixed, the electronic device uses multiple image pairs to train the model to be trained until the output value of the loss function of the model to be trained is less than or equal to a preset threshold, thereby obtaining the target generative model.

[0195] In some examples, the electronic device trains a model to be trained using multiple image pairs, with a fixed base model that has been trained in place. The electronic device uses the third image data corresponding to the first image (HR Image) and the second image (LR Image) to train the modules in the conditional control generative network to be trained, so that the model to be trained learns how to reconstruct the target image (first image) from the second image.

[0196] The following is passed Figure 11 Step S305 will be described in detail below. For example, Figure 11 This is a schematic diagram illustrating a method for training a model on an electronic device. It should be noted that this method does not rely on... Figure 11 The specific order described below is a limitation. It should be understood that in other embodiments, the order of some steps in the method can be interchanged according to actual needs, or some steps can be omitted or deleted. For example, Figure 11 As shown, the method includes the following steps S1101 to S1109:

[0197] S1101, The electronic device inputs the second image data to be trained into the first fixed encoder block of the fixed base model.

[0198] In some examples, the second training image data is data in which the electronic device adds noise to the first image in an image pair, or data in which the electronic device adds noise to the latent space representation corresponding to the first image after encoding the first image in an image pair into a latent space representation.

[0199] For example, Figure 12 This is a schematic diagram illustrating the data processing flow of a model to be trained, including a fixed-base model. For example... Figure 12 As shown, the second training image data can be data obtained by Variational Autoencoder (VAE-Encoder) encoding the first image (HR Image) in an image pair into a latent space representation, and then adding Gaussian noise to the latent space representation corresponding to that image pair. Optionally, the model to be trained can pass the second training image data to the first fixed encoder block through a convolutional layer. The convolutional layer can perform processing on the second training image data, such as feature extraction, noise robustness enhancement, noise filtering, feature enhancement, and data dimensionality reduction, thereby promoting the accurate recognition and rapid processing of the second training image data by the base model to be trained.

[0200] S1102. The electronic device uses the first fixed encoder block in the fixed base model to downsample and extract features from the second image data to be trained, thereby obtaining the second feature map to be trained.

[0201] In some examples, the first fixed encoder block of the fixed-base model is able to extract features from the second training image data relatively well. The second training feature map includes the features extracted from the second training image data by the first fixed encoder block. The second training feature map is used in subsequent image denoising and generation processes.

[0202] S1103. The electronic device uses the second encoder block in the training condition control generator network, which has been trained M times, to downsample and extract features from the third image data input to the second encoder block to obtain a third feature map. The third image data includes the second image data.

[0203] In some examples, the third training image data is the data corresponding to an image pair after adding noise to the first image in an image pair, or the data corresponding to an image pair after the electronic device encodes an image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image in the image pair. As mentioned above, the second training image data is the data after the electronic device adds noise to the first image in an image pair, or the data after the electronic device encodes the first image in an image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image. Therefore, the third training image data includes the second training image data. Additionally, the third training image data also includes the second image in an image pair, or the data after the electronic device encodes the second image in an image pair into a latent space representation.

[0204] For example, such as Figure 12 As shown, Figure 12 SD Encoder Block 1 is the second encoder block to be trained. The third training image data includes the first data input to SD Encoder Block 1 from the convolutional layer (conv layer) on the left side of the training model. This first data is the data after the variational autoencoder (VAE-Encoder) on the left side of the training model encodes the first image (HR Image) in an image pair into a latent space representation, and then adds Gaussian noise to the latent space representation corresponding to that image pair. The third training image data also includes the second data input to SD Encoder Block 1 from the zero convolutional layer (zero conv) on the right side of the training model. This second data is the data after the variational autoencoder (VAE-Encoder) on the right side of the training model encodes the second image (LR Image) in an image pair into a latent space representation.

[0205] In some examples, the third feature map to be trained includes features extracted from the third image data by the second encoder block to be trained. This feature can introduce control signals to the fixed-base model, helping it learn to reconstruct the target image (the first image) from the second image.

[0206] In some examples, M is a positive integer. The image generation capability of the trainable model, consisting of the conditional control generative network and the fixed base model, after M training iterations, can be obtained in step S1108 by calculating the output value of the loss function.

[0207] S1104. The electronic device uses the second intermediate block to be trained to extract features from the third feature map to be trained, and obtains the second intermediate feature map to be trained. The second intermediate feature map to be trained is used to introduce control signals for the fixed base model.

[0208] In some examples, the second training intermediate feature map includes features extracted from the third training feature map by the second training intermediate block. This feature can introduce control signals for the fixed-base model, helping it to restore the sharpness of the second image to that of the first image.

[0209] S1105, the electronic device passes the third feature map to be trained to the first fixed decoder block of the fixed base model through a zero convolutional layer. The third feature map to be trained is used to introduce control signals to the fixed base model.

[0210] In some examples, zero convolutional layers are used to connect a second encoder block to be trained and a second intermediate block to be trained to a first fixed decoder block.

[0211] S1106. The electronic device uses the first fixed intermediate block to extract features from the second training intermediate feature map based on the control signal provided by the second training intermediate feature map, and obtains the target intermediate feature map.

[0212] In some examples, the second intermediate feature map to be trained includes features extracted by the second intermediate block from the third feature map to be trained. Thus, the control signals provided by the second intermediate feature map to be trained can assist the first fixed intermediate block in extracting features from the second feature map to be trained.

[0213] S1107. The electronic device uses the first fixed decoder block to upsample and decode the target intermediate feature map based on the control signal provided by the third feature map to be trained, and generates the third output image.

[0214] In some examples, the third feature map to be trained includes features extracted from the third image data by the second encoder block to be trained. The control signals provided by these features help the first fixed decoder block to reconstruct the target intermediate feature map into the third output image.

[0215] S1108. The electronic device calculates the output value of the loss function of the model to be trained based on the third output image.

[0216] In some examples, electronic devices measure the performance of the training model by calculating the loss function of the training model to measure the difference between the third output image generated by the training model and the first image.

[0217] S1109. The electronic device trains the model to be trained until the output value of the loss function of the model to be trained is less than or equal to the preset threshold of the loss function of the model to be trained, thereby obtaining the target generative model.

[0218] In some examples, if the output value of the loss function of the model to be trained is not greater than the preset threshold of the loss function, the electronic device uses the base model to be trained after M iterations as the target generative model. If the output value of the loss function of the model to be trained is greater than the preset threshold, the electronic device adjusts the parameters of the generative network controlled by the training conditions. The electronic device inputs another pair of images from multiple image pairs into the model to be trained and continues to train the model until the output value of the loss function of the model to be trained is less than or equal to the preset threshold, thus obtaining the target generative model.

[0219] In some examples, when the output value of the loss function of the model to be trained is not greater than the preset threshold of the loss function, the electronic device determines that the model to be trained has converged and uses the model after M training iterations as the target generative model. In other examples, when the output value of the loss function of the model to be trained is greater than the preset threshold, the electronic device determines that the model to be trained has not yet converged and the parameters of the generative network need to be adjusted again. The electronic device inputs another pair of images from multiple image pairs into the model to be trained and continues to train the model until the output value of the loss function of the model to be trained is less than or equal to the preset threshold, then training stops, and the target generative model is obtained.

[0220] In this way, electronic devices can train a higher-performing model by minimizing the loss function of the model to be trained, enabling it to enhance high-resolution images more effectively.

[0221] In some embodiments, in step S1108, the loss function used by the electronic device to train the model to be trained includes the mean squared error loss function.

[0222] In other embodiments, the loss function used by the electronic device to train the model includes a perceptual loss function. The electronic device calculates the output value of the perceptual loss function based on a third output image. If the output value of the perceptual loss function is greater than a preset threshold, the electronic device adjusts the parameters of the network to be trained.

[0223] In some examples, electronic devices evaluate the quality of the third output image generated by the model under training by using a perceptual loss function to measure the difference in feature space between the third output image generated by the model under training and the target image (i.e., the first image).

[0224] For example, Figure 13 This is a schematic diagram illustrating a method used by an electronic device to train a model using a loss function. For example... Figure 13 As shown, the electronic device uses a perceptual loss function to measure the difference between the visual features of the output image generated by the model under test and the visual features of the target image. The electronic device minimizes the output value of the perceptual loss function, making the output image generated by the model under test more closely resemble the target image. This process is described in detail below:

[0225] The calculation of the perceptual loss function requires two sets of images: a third output image and a target image (i.e., the first image). The third output image is generated by the model under test after enhancing the second image. The training objective of the model under test is to make the third output image more closely resemble the target image. The electronic device feeds these two sets of images into a pre-trained deep neural network. This pre-trained deep neural network can be a model such as VGG that has been pre-trained on a large amount of data. Through the pre-training process, it has learned a rich hierarchical structure of visual features.

[0226] In this pre-trained deep neural network, the third output image and / or the target image are processed layer by layer. The electronic device captures the visual features of the third output image and / or the target image based on the outputs of certain specific layers in the pre-trained deep neural network. These visual features constitute the feature representation of the third output image and / or the target image in the pre-trained deep neural network. This feature representation contains high-level semantic information and visual texture information of the third output image and / or the target image.

[0227] After obtaining the visual features of the third output image and the target image, the electronic device calculates the difference between these two sets of visual features. The electronic device quantifies this difference by calculating the Euclidean distance or Manhattan distance between the two sets of visual features in the feature space. Euclidean distance is the most intuitive distance metric, measuring the straight-line distance between feature points. Manhattan distance, on the other hand, is the total distance traveled along the coordinate axes in a grid-like space. The perceptual loss based on Euclidean distance can be expressed as:

[0228]

[0229] This represents the feature extractor of the i-th layer of the pre-trained network. represents the generated third output image, and x represents the target image.

[0230] In this way, the electronic device reduces the difference between the visual features corresponding to the third output image and the target image by minimizing the output value of the perceptual loss function. During this process, the model to be trained can learn how to adjust the generated third output image so that its visual features are as close as possible to those of the target image.

[0231] In some examples, the parameters of the conditional generation network to be trained include global parameters, local parameters, and other parameters. Global parameters include the learning rate and number of iterations. Local parameters include convolutional layer parameters, intensity parameters, and detail parameters. Other parameters include regularization parameters and optimizer parameters. This application does not limit the parameters of the conditional generation network to be trained by the electronic device. In some examples, when adjusting the parameters of the conditional generation network to be trained, the electronic device may adjust the global parameters first, and then adjust the local parameters. This application also does not limit this approach.

[0232] Thus, by using a perceptual loss function to constrain the model being trained, electronic devices can not only improve the pixel-level similarity between the third output image and the target image, but also enhance the visual realism of the third output image. Ultimately, this allows the third output image generated by the model to be more visually and stylistically similar to the target image.

[0233] In some embodiments, the loss function used by an electronic device to train a model includes an adversarial loss function. For example, such as... Figure 13 As shown, the model to be trained also includes a discriminator, which is used to compete against the generator, which consists of a fixed-base model and a conditionally controlled generative network to be trained. Electronic devices allow the generator and discriminator to compete against each other, minimizing the output values ​​of the first adversarial loss function corresponding to the discriminator and the first adversarial loss function corresponding to the generator, respectively, to improve the quality and realism of the images generated by the model to be trained. This process is described in detail below:

[0234] In some examples, the adversarial loss is divided into two parts: the generator's loss and the discriminator's loss. The discriminator aims to maximize its output probability for the target image while minimizing its output probability for the fourth output image generated by the model under training. The first loss function for the discriminator can be expressed as:

[0235]

[0236] Where x is from the real data distribution p data The target image, z, is derived from the noise distribution p. z The random noise vector sampled in (z). G(z) is the fourth output image generated by the generator. D(x) is the output probability of the discriminator for the target image x, and D(G(z)) is the output probability of the discriminator for the generated fourth output image G(z).

[0237] The generator's goal is to minimize the discriminator's output probability for the fourth output image, that is, to make the discriminator consider the fourth output image to be real. The generator's second loss function can be expressed as:

[0238] During training, the generator and discriminator alternately optimize their respective loss functions. For example, Figure 14 This is a schematic diagram illustrating a method for training a model on an electronic device based on an adversarial loss function. It should be noted that this method does not rely on... Figure 14 The specific order described below is a limitation. It should be understood that in other embodiments, the order of some steps in the method can be interchanged according to actual needs, or some steps can be omitted or deleted. For example, Figure 14 As shown, the method includes the following steps S1401 to S1406:

[0239] S1401, the parameters of the fixed generator of the electronic device are input into the model to be trained from a pair of multiple image pairs, and the fourth output image is obtained from the first fixed decoder block.

[0240] S1402. The electronic device calculates the output value of the first adversarial loss function of the discriminator based on the fourth output image.

[0241] S1403. The electronic device trains the model to be trained until the output value of the first adversarial loss function is less than or equal to the preset threshold of the first adversarial loss function.

[0242] In some examples, the parameters of the generator are fixed in the electronic device, minimizing the output value of the first adversarial loss function. The parameters in the discriminator are updated so that the discriminator can distinguish between the target image and the fourth output image.

[0243] For example, if the output value of the first adversarial loss function is greater than a preset threshold of the first adversarial loss function, the electronic device adjusts the parameters of the discriminator. The electronic device inputs another pair of images from multiple image pairs into the model to be trained, and trains the model until the output value of the first adversarial loss function is less than or equal to the preset threshold of the first adversarial loss function.

[0244] S1404. The electronic device fixes the parameters of the discriminator, inputs another pair of images from multiple image pairs into the model to be trained, and obtains the fifth output image from the first fixed decoder block.

[0245] S1405. The electronic device calculates the output value of the generator's second adversarial loss function based on the fifth output image.

[0246] S1406. The electronic device trains the model to be trained until the output value of the second adversarial loss function is less than or equal to the preset threshold of the second adversarial loss function.

[0247] In some examples, the electronic device fixes the parameters of the discriminator and sets the output value of the second adversarial loss function. When the training reaches the preset threshold of the second adversarial loss function, the discriminator considers the fifth output image to be very similar to the target image. During this process, the generator continuously updates its parameters to make the generated fifth output image even closer to the target image.

[0248] For example, if the output value of the second adversarial loss function is greater than a preset threshold of the second adversarial loss function, the electronic device adjusts the parameters of the generator. The electronic device inputs another pair of images from multiple image pairs into the model to be trained, and trains the model until the output value of the second adversarial loss function is less than or equal to the preset threshold of the second adversarial loss function.

[0249] In this way, the electronic device forces the generator and discriminator to compete with each other. The generator learns to produce images that are difficult for the discriminator to distinguish, while the discriminator continuously improves its ability to differentiate between generated and target images. Ultimately, this enables the model under training to generate high-quality, highly realistic images.

[0250] In some embodiments, the electronic device can also determine the performance of the model under training by calculating the similarity between the currently generated output image and the previously generated output image. If the similarity between the current output image and the previous output image is less than or equal to a preset similarity threshold, the electronic device considers that further training can no longer significantly improve the performance of the base model under training. At this point, the electronic device determines that the base model under training has completed sufficient training and stops training the base model under training. The electronic device defines the trained base model under training as a fixed base model. The similarity between the current output image and the previous output image can be the Euclidean distance or cosine similarity between the image features in the current output image and the previous output image.

[0251] The above combination Figures 1-14 The model generation method provided in the embodiments of this application is described in detail below. Figure 15 and Figure 16 This application provides a detailed description of the electronic device provided in its embodiments.

[0252] Figure 15 This is a schematic diagram of the hardware structure of the electronic device 100 provided in the embodiments of this application.

[0253] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, and subscriber identification module (SIM) card interface 195, etc.

[0254] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0255] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0256] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0257] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0258] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0259] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0260] Internal memory 121 can be used to store executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located within the processor.

[0261] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0262] Figure 16 This is a software structure block diagram of the electronic device provided in the embodiments of this application.

[0263] See Figure 16 As shown, the layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0264] The application layer can include a series of application packages.

[0265] like Figure 16 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0266] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0267] like Figure 16 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0268] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0269] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0270] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0271] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0272] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0273] The foregoing primarily describes the solutions provided by the embodiments of this application from the perspective of electronic devices. It is understood that, in order to achieve the aforementioned functions, the electronic device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the steps of a model generation method described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by software-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0274] This application embodiment can divide the above-described electronic device into functional modules or functional units according to the above method examples. For example, each function can be divided into a separate functional module or functional unit, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0275] This application also provides a chip system. Figure 17 This is a block diagram of a chip system provided in an embodiment of this application.

[0276] See Figure 17As shown, the chip system includes at least one processor 1701 and at least one interface circuit 1702. The processor 1701 and the interface circuit 1702 are interconnected via lines. For example, the interface circuit 1702 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 1702 can be used to send signals to other devices. Exemplarily, the interface circuit can read instructions stored in the memory and send those instructions to the processor 1701. When the instructions are executed by the processor 1701, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete components, which are not specifically limited in this application embodiment.

[0277] This application also provides a computer-readable storage medium, which includes computer instructions that, when the computer instructions are used in the aforementioned electronic device (such as...), Figure 15 or Figure 16 When the electronic device 100 shown is run, it causes the electronic device to perform the various functions or steps performed by the mobile phone in the above method embodiment.

[0278] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.

[0279] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, the division of the above functional modules is only used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0280] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0281] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0282] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0283] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0284] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A model generation method characterized by comprising: The method includes: An initial model is obtained, which includes an initial base model and an initial conditional control generation network. The initial base model is a pre-trained diffusion model of the U-Net architecture, and the initial conditional control generation network is a branch network of ControlNet. The initial conditional control generation network is used to introduce control signals to the initial base model. Multiple image pairs are constructed, each image pair including a first image and a second image. The first image is a training image with a first resolution and a first sharpness, and the second image is a training image with a first resolution and a second sharpness. The first image and the second image are training images with different sharpnesses based on the same scene, and the first sharpness is higher than the second sharpness. The output size of the sampling module of the initial model is adjusted to match the first resolution to obtain the model to be trained. The model to be trained includes a base model to be trained and a conditionally controlled generative network to be trained. Using the first image, the base model to be trained is trained until the output value of the loss function of the base model to be trained is less than or equal to a preset threshold, thereby obtaining a fixed base model; With the fixed base model fixed, the target generative model is obtained by using the multiple image pairs to train the model to be trained until the output value of the loss function of the model to be trained is less than or equal to a preset threshold.

2. The method of claim 1, wherein, The initial base model includes multiple first initial encoder blocks, a first initial intermediate block, and multiple first initial decoder blocks; The first initial encoder block is used to downsample and extract features from the first initial image data to obtain a first initial feature map, and then pass the first initial feature map to the first initial intermediate block. The first initial image data is the first image after adding noise or the data after the electronic device encodes the first image into a latent space representation and adds noise to the latent space representation corresponding to the first image. The first initial encoder block includes a first initial downsampling module, which is used to reduce the resolution of the first initial image data and extract features from the first initial image data. The first initial intermediate block is located between the first initial encoder block and the first initial decoder block. The first initial intermediate block is used to extract features from the received first initial feature map to obtain a first initial intermediate feature map. The first initial decoder block is used to upsample the first initial intermediate feature map to generate the first output image; The first initial decoder block further includes a first upsampling module, which is used to upsample the first initial intermediate feature map.

3. The method of claim 2, wherein, The initial condition control generation network includes multiple second initial encoder blocks, a second initial intermediate block, and multiple zero convolutional layers; The second initial encoder block is used to downsample and extract features from the second initial image data input to the second initial encoder block to obtain a second initial feature map, pass the second initial feature map to the second initial intermediate block, and pass the second initial feature map to the first initial decoder block of the initial base model through the zero convolutional layer. The second initial feature map is used to introduce control signals to the initial base model. The second initial image data is the data corresponding to the image pair after adding noise to the first image in the image pair, or the data corresponding to the image pair after the electronic device encodes the image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image in the image pair. The second initial intermediate block is used to extract features from the second initial feature map to obtain a second intermediate feature map. The second intermediate feature map is then passed to the first intermediate block of the initial base model through a zero convolutional layer. The second intermediate feature map is used to introduce control signals into the initial base model. The second initial encoder block further includes a second initial downsampling module, which is used to reduce the resolution of the second image data and extract features from the second image data.

4. The method of claim 3, wherein, Adjusting the output size of the sampling module of the initial model to match the first resolution includes: The sampling rates of the first initial downsampling module and the first initial upsampling module corresponding to the first initial downsampling module in the initial base model are adjusted to match the first resolution. The sampling rate of the second initial downsampling module in the initial condition control generation network is adjusted to match the first resolution.

5. The method of claim 4, wherein, The output resolution of the first initial downsampling module, the first initial upsampling module, and the second initial downsampling module is all the second resolution, which is smaller than the first resolution.

6. The method of claim 5, wherein, The step of adjusting the sampling rate of the first initial downsampling module and the first initial upsampling module corresponding to the first initial downsampling module in the initial base model to match the first resolution, and adjusting the sampling rate of the second initial downsampling module in the initial condition control generation network to match the first resolution, includes: Based on the ratio of the first resolution to the second resolution, the downsampling factor of the first initial downsampling module, the upsampling factor of the first initial upsampling module, and the downsampling factor of the second initial downsampling module are increased to obtain the first training downsampling module, the first training upsampling module, and the second training downsampling module.

7. The method according to claim 6, characterized in that, The base model to be trained includes multiple first encoder blocks to be trained, multiple intermediate blocks to be trained, and multiple decoder blocks to be trained. The first encoder block to be trained includes the first downsampling module to be trained, and the first decoder block to be trained includes the first upsampling module to be trained. The fixed base model includes multiple first fixed encoder blocks, a first fixed intermediate block, and multiple first fixed decoder blocks; The conditional control generator network to be trained includes multiple second encoder blocks to be trained, a second intermediate block to be trained, and multiple zero convolutional layers. The second encoder block to be trained includes a second downsampling module to be trained.

8. The method according to claim 7, characterized in that, The step of using the first image to train the base model to be trained until the output value of the loss function of the base model to be trained is less than or equal to a preset threshold, thereby obtaining a fixed base model, includes: The first training image data is input into the first training encoder block of the training base model after N training iterations. The first training image data is the first image after adding noise or the data after the electronic device encodes the first image into a latent space representation and adds noise to the latent space representation corresponding to the first image. N is a positive integer. The first encoder block to be trained is used to downsample and extract features from the first image data to be trained to obtain the first feature map to be trained. The first training intermediate block is used to further extract features from the first training feature map to obtain the first training intermediate feature map. The first intermediate feature map to be trained is upsampled using the first decoder block to be trained to generate the second output image; Based on the second output image, the output value of the base model loss function of the base model to be trained is calculated.

9. The method according to claim 8, characterized in that, The method further includes: if the output value of the base model loss function is not greater than a preset threshold of the base model loss function, the base model to be trained after N training iterations is used as the fixed base model.

10. The method according to claim 9, characterized in that, The method further includes: If the output value of the base model loss function is greater than the preset threshold of the base model loss function, the parameters of the base model to be trained are adjusted. The base model to be trained is trained using another of the first images until the output value of the base model loss function is less than or equal to a preset threshold of the base model loss function, thus obtaining a fixed base model.

11. The method according to claim 10, characterized in that, The step of training the model to be trained until the output value of the loss function of the model to be trained is less than or equal to a preset threshold, while fixing the fixed base model, to obtain the target generation model, includes: The second training image data is input into the first fixed encoder block of the fixed base model. The second training image data is the data after the electronic device adds noise to the first image in the image pair, or the data after the electronic device encodes the first image in the image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image. The first fixed encoder block in the fixed base model is used to downsample and extract features from the second image data to be trained, thereby obtaining the second feature map to be trained. The second encoder block in the conditional generator network to be trained after M training iterations downsamples and extracts features from the third image data to be trained input to the second encoder block to obtain a third feature map to be trained. The third image data to be trained includes the second image data to be trained. The third initial image data is the data corresponding to the first image in the image pair after adding noise, or the data corresponding to the image pair after the electronic device encodes the image pair into a latent space representation and adds noise to the latent space representation corresponding to the first image in the image pair. M is a positive integer. The second intermediate block to be trained extracts features from the third feature map to be trained to obtain a second intermediate feature map to be trained. The second intermediate feature map to be trained is used to introduce control signals for the fixed base model. The third feature map to be trained is passed to the first fixed decoder block of the fixed base model through the zero convolutional layer, and the third feature map to be trained is used to introduce control signals to the fixed base model. The first fixed intermediate block performs feature extraction on the second training intermediate feature map based on the control signal provided by the second training intermediate feature map to obtain the target intermediate feature map; The first fixed decoder block performs upsampling and decoding processing on the target intermediate feature map based on the control signal provided by the third training feature map, and generates a third output image; Based on the third output image, calculate the output value of the loss function of the model to be trained.

12. The method according to claim 11, characterized in that, The method further includes: if the output value of the loss function of the model to be trained is not greater than a preset threshold of the loss function to be trained, the base model to be trained after M training iterations is used as the target generation model.

13. The method according to claim 12, characterized in that, The method further includes: if the output value of the loss function of the model to be trained is greater than a preset threshold of the loss function to be trained, adjusting the parameters of the training condition control generation network; Another pair of images from the plurality of image pairs is input into the model to be trained, and the model to be trained is trained until the output value of the loss function of the model to be trained is less than or equal to a preset threshold of the loss function of the model to be trained, thereby obtaining the target generation model.

14. The method according to claim 13, characterized in that, If the output value of the loss function of the model to be trained is greater than a preset threshold of the loss function, the parameters of the training condition control network are adjusted, including: Based on the third output image, the output value of the perceptual loss function is calculated. If the output value of the perceptual loss function is greater than the preset threshold of the perceptual loss function, the parameters of the training condition control generation network are adjusted.

15. The method according to claim 14, characterized in that, The model to be trained further includes a discriminator, which is used to adversarially compete with the generator composed of the fixed-base model and the conditionally controlled generative network to be trained. Both the generator and the discriminator include an adversarial loss function. The step of training the model to be trained until the output value of the loss function of the model to be trained is less than or equal to a preset threshold of the loss function of the model to be trained, thereby obtaining the target generative model, further includes: With the parameters of the generator fixed, a pair of images from the plurality of image pairs are input into the model to be trained, and a fourth output image is obtained from the first fixed decoder block; Based on the fourth output image, the output value of the first adversarial loss function of the discriminator is calculated; If the output value of the first adversarial loss function is greater than the preset threshold of the first adversarial loss function, the parameters of the discriminator are adjusted. Another pair of images from the plurality of image pairs is input into the model to be trained, and the model to be trained is trained until the output value of the first adversarial loss function is less than or equal to a preset threshold of the first adversarial loss function. By fixing the parameters of the discriminator, another pair of images from the plurality of image pairs is input into the model to be trained, and the fifth output image is obtained from the first fixed decoder block; Based on the fifth output image, calculate the output value of the generator's second adversarial loss function; If the output value of the second adversarial loss function is greater than the preset threshold of the second adversarial loss function, the parameters of the generator are adjusted. Another pair of images from the plurality of image pairs is input into the model to be trained, and the model to be trained is trained until the output value of the second adversarial loss function is less than or equal to a preset threshold of the second adversarial loss function.

16. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 15.

17. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 15.