Image generation method and device, equipment, storage medium and computer program product
By dividing the image generation process into inference process of multiple noise level and resolution stages, the existing image generation methods are solved, and efficient image generation is achieved.
Patent Information
- Application Number
- CN202510245473.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-04
AI Technical Summary
Although the existing diffusion model based on Transformer architecture can generate high-quality images, the inference is time-consuming and takes up a lot of video memory, resulting in high application costs.
The image generation process is divided into multiple inference stages from the two dimensions of noise level and image resolution. Image inference is performed in sequence through preset image generation models, and gradually denoising and upsampling are generated to generate the target image.
It reduces the computational complexity and memory usage, reduces the inference time, and thus reduces the application cost of image generation models.
Smart Images

Figure CN120259458A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to an image generation method, apparatus, device, storage medium, and computer program product. Background Art
[0002] Currently, the AIGC (Artificial Intelligence Generated Content) text-to-image generation technology, as an application of the most popular generative artificial intelligence in the field of artificial intelligence in recent years, has extremely wide related application scenarios. With the model iteration, the current more mainstream and better-performing text-to-image model is the diffusion model based on the Transformer architecture (i.e., the DiT (Diffusion Transformers) structure model).
[0003] However, although the mainstream diffusion models based on the Transformer architecture (such as SD3, Flux, etc.) can generate high-quality images, the inference time is very long and the video memory occupancy is also large, resulting in a high application cost. Summary of the Invention
[0004] The main purpose of this application is to provide an image generation method, apparatus, device, storage medium, and computer program product, aiming to solve the technical problem that although the existing image generation methods can generate high-quality images, the inference time is long and the video memory occupancy is large, resulting in a high application cost.
[0005] To achieve the above purpose, this application provides an image generation method, and the image generation method includes:
[0006] In response to an image generation prompt, input the image generation prompt into a preset image generation model;
[0007] Divide the image generation process into multiple inference stages from a first dimension and a second dimension through the preset image generation model, where the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution;
[0008] Perform image inference through the multiple inference stages in sequence based on the image generation prompt to generate a target image.
[0009] Optionally, the performing image inference through the multiple inference stages in sequence based on the image generation prompt to generate a target image includes:
[0010] In each inference stage, perform image inference according to the image generation prompt and the output of the previous inference stage to generate the image of the current inference stage;
[0011] The output of the current inference stage is upsampled and used as the input for the next inference stage until a target image with the target resolution is generated.
[0012] Optionally, each of the inference stages corresponds to a different noise level range and a different image resolution. During the sequential execution of the multiple inference stages, the noise is gradually reduced from high noise to no noise, and the resolution is gradually upsampled from low resolution to high resolution. During the sequential execution of the multiple inference stages, the number of image generation modules used in each inference stage increases sequentially.
[0013] Optionally, the image generation module used in each inference stage is the core generation module of the target model, where the target model is a diffusion model based on the Transformer architecture, and the image generation module is a Transformer module.
[0014] Optionally, the step of inputting the image generation prompt into a preset image generation model in response to the image generation prompt includes:
[0015] In response to the image generation prompt, mapping the image generation prompt to the latent space to obtain a latent space representation, and inputting the latent space representation into the preset image generation model;
[0016] Correspondingly, the step of performing image inference through the multiple inference stages in sequence based on the image generation prompt to generate a target image includes:
[0017] Performing image inference through the multiple inference stages in sequence based on the latent space representation to generate a target image.
[0018] Optionally, the step of mapping the image generation prompt to the latent space to obtain a latent space representation, and inputting the latent space representation into the preset image generation model in response to the image generation prompt includes:
[0019] In response to the image generation prompt, extracting semantic information of the image generation prompt through a text encoder, where the text encoder supports texts in different languages;
[0020] Based on the semantic information, mapping the image generation prompt to a latent space matching the preset image generation model through the text encoder to obtain a latent space representation, and inputting the latent space representation into the preset image generation model, where texts in different languages are mapped to the same latent space after being encoded.
[0021] Optionally, before inputting the image generation prompt into the preset image generation model in response to the image generation prompt, it further includes:
[0022] Dividing the image generation process into multiple training stages from the first dimension and the second dimension through an initial image generation model;
[0023] Train the initial image generation model through the multiple training stages in sequence based on the image samples to obtain a preset image generation model.
[0024] Optionally, the training the initial image generation model through the multiple training stages in sequence based on the image samples to obtain a preset image generation model includes:
[0025] Train through the multiple training stages in sequence based on the image samples, and obtain the loss value of each training stage;
[0026] Train the initial image generation model based on the loss values of the respective training stages to obtain a preset image generation model.
[0027] Optionally, the training the initial image generation model based on the loss values of the respective training stages to obtain a preset image generation model includes:
[0028] Perform weighted summation on the loss values of the respective training stages to obtain a total loss value;
[0029] Adjust the model parameters of the initial image generation model based on the total loss value to obtain a preset image generation model.
[0030] Optionally, the dividing the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model includes:
[0031] Obtain the model deployment requirements, and determine the number of stage divisions according to the model deployment requirements;
[0032] Divide the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model based on the number of stage divisions.
[0033] Optionally, the dividing the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model based on the number of stage divisions includes:
[0034] Divide multiple noise level ranges and multiple resolutions according to the number of stage divisions;
[0035] Divide the image generation process into multiple training stages through the initial image generation model based on the multiple noise level ranges and the multiple resolutions.
[0036] In addition, to achieve the above object, the present application also proposes an image generation device, and the image generation device includes:
[0037] An information input module, configured to input the image generation prompt into a preset image generation model in response to the image generation prompt;
[0038] A process division module, configured to divide the image generation process into multiple inference stages from a first dimension and a second dimension through the preset image generation model, where the first dimension corresponds to a noise level and the second dimension corresponds to an image resolution;
[0039] An image inference module, configured to perform image inference through the multiple inference stages in sequence based on the image generation prompt to generate a target image.
[0040] Optionally, the image inference module is further configured to, in each inference stage, perform image inference according to the image generation prompt and the output of the previous inference stage to generate an image of the current inference stage; upsample the output of the current inference stage and use it as the input of the next inference stage until a target image with a target resolution is generated.
[0041] Optionally, each of the inference stages corresponds to a different noise level range and a different image resolution. During the sequential execution of the multiple inference stages, the noise is gradually reduced from high noise to no noise, and the resolution is gradually upsampled from low resolution to high resolution. During the sequential execution of the multiple inference stages, the number of image generation modules used in each inference stage increases sequentially.
[0042] Optionally, the image generation module used in each inference stage is the core generation module of the target model, where the target model is a diffusion model based on the Transformer architecture, and the image generation module is a Transformer module.
[0043] Optionally, the information input module is further configured to map the image generation prompt to a latent space in response to the image generation prompt to obtain a latent space representation, and input the latent space representation into the preset image generation model;
[0044] Correspondingly, the image inference module is further configured to perform image inference through the multiple inference stages in sequence based on the latent space representation to generate a target image.
[0045] Optionally, the information input module is further configured to extract semantic information of the image generation prompt through a text encoder in response to the image generation prompt, where the text encoder supports texts in different languages; map the image generation prompt to a latent space matching the preset image generation model based on the semantic information through the text encoder to obtain a latent space representation, and input the latent space representation into the preset image generation model, where texts in different languages are mapped to the same latent space after encoding.
[0046] In addition, to achieve the above object, the present application further provides an image generation device, which includes a memory, a processor, and an image generation program stored on the memory and executable on the processor. The image generation program is configured to implement the image generation method as described above.
[0047] In addition, to achieve the above object, the present application further provides a storage medium, on which an image generation program is stored. When the image generation program is executed by a processor, it implements the image generation method as described above.
[0048] In addition, to achieve the above object, the present application further provides a computer program product, which includes an image generation program. When the image generation program is executed by a processor, it implements the image generation method as described above.
[0049] One or more technical solutions proposed by the present application have at least the following technical effects:
[0050] In the present application, in response to an image generation prompt, the image generation prompt is input into a preset image generation model. The preset image generation model divides the image generation process into multiple inference stages from a first dimension and a second dimension. Among them, the first dimension corresponds to the noise level, and the second dimension corresponds to the image resolution. Based on the image generation prompt, multiple inference stages are sequentially performed for image inference to generate a target image. Since the present application divides the image generation stage into multiple inference stages from the noise level dimension and the image resolution dimension through the preset image generation model, and performs image generation in stages, the computational complexity is reduced, the inference time and the occupied video memory are reduced, and further the application cost of the preset image generation model is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0052] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a schematic flowchart of the first embodiment of the image generation method of the present application;
[0054] Figure 2 It is a schematic flowchart of the second embodiment of the image generation method of the present application;
[0055] Figure 3The structural diagram of the spatio-temporal pyramid model for an embodiment of the image generation method of this application;
[0056] Figure 4 The schematic flowchart of the third embodiment of the image generation method of this application;
[0057] Figure 5 The schematic module structure diagram of the image generation device for an embodiment of this application;
[0058] Figure 6 The schematic device structure diagram of the hardware operating environment involved in the image generation method for an embodiment of this application.
[0059] The implementation, functional features, and advantages of the purpose of this application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0060] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0061] To better understand the technical solutions of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0062] Currently, as an application of the most popular generative artificial intelligence in the field of artificial intelligence in recent years, the AIGC (Artificial Intelligence Generated Content) text-to-image generation technology has extremely wide related application scenarios. With the model iteration, the current more mainstream and better-performing text-to-image model is the diffusion model based on the Transformer architecture (i.e., the DiT (Diffusion Transformers) structure model). Among them, the diffusion model based on the Transformer architecture generates high-quality images through an iterative denoising process. Its core is to replace the U-Net of the traditional diffusion model with a Transformer module and use the attention mechanism to capture global information.
[0063] However, although the mainstream diffusion models based on the Transformer architecture (such as SD3, Flux, etc.) can generate high-quality images, the inference time is very long and the video memory occupation is also large, resulting in a relatively high application cost.
[0064] Therefore, in order to overcome the above defects, the present application provides a solution, which includes: in response to an image generation prompt, inputting the image generation prompt into a preset image generation model, and dividing the image generation process into multiple inference stages from a first dimension and a second dimension by the preset image generation model, where the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution, and performing image inference through multiple inference stages in sequence based on the image generation prompt to generate a target image; since the present application divides the image generation stage into multiple inference stages from the noise level dimension and the image resolution dimension by the preset image generation model, and performs image generation in stages, the computational complexity is reduced, the inference time and the occupied video memory are reduced, and further the application cost of the preset image generation model is reduced.
[0065] It should be noted that the execution subject of this embodiment may be an image generation device with data processing, network communication, and program running functions, such as a server, a computer, etc., or other electronic devices that can implement the same or similar functions. This embodiment does not limit this.
[0066] Based on this, an embodiment of the present application provides an image generation method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the image generation method of the present application.
[0067] In the first embodiment, the image generation method includes:
[0068] Step S10: In response to an image generation prompt, input the image generation prompt into a preset image generation model.
[0069] It should be understood that the image generation prompt may refer to a text prompt used to guide the preset image generation model to generate specific content or style, such as "a picture of a black motorcycle". In a specific implementation, the user inputs an image generation prompt, such as "a picture of a black motorcycle", and the image generation device receives the image generation prompt and inputs the image generation prompt into the preset image generation model.
[0070] Step S20: Divide the image generation process into multiple inference stages from a first dimension and a second dimension by the preset image generation model, where the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution.
[0071] It can be understood that the preset image generation model can be a spatio-temporal pyramid model. Among them, the spatio-temporal pyramid model can refer to a multi-stage generation model that combines the time dimension (noise level) and the space dimension (image resolution). By gradually generating images from low resolution to high resolution in stages, the computational cost can be reduced and the efficiency can be improved. Among them, the noise level can be a quantization index representing the denoising process in the diffusion model, and the noise level can vary continuously from 1 (fully noisy) to 0 (fully denoised). The image resolution can refer to the number of pixels contained in a unit length in the image, which is an important index to measure the clarity and detail performance of the image.
[0072] In a specific implementation, the image generation process is divided into multiple inference stages by combining the spatio-temporal pyramid with the noise level and the image resolution. Each inference stage corresponds to a different noise level range and a different image resolution.
[0073] Step S30: Based on the image generation prompt, perform image inference in the multiple inference stages in sequence to generate a target image.
[0074] In a specific implementation, in each inference stage, image inference is performed according to the image generation prompt and the output of the previous inference stage to generate the image of the current inference stage. Then, the image of the current inference stage is upsampled (or further processed) as the input of the next inference stage. Through gradual inference and generation, a high-quality target image that meets the user's prompt is finally obtained.
[0075] Furthermore, in order to facilitate the preset image generation model to perform image generation and improve the image generation efficiency, in this embodiment, the image generation prompt is first converted into a latent space representation, and then the latent space representation is input into the preset image generation model for image inference. The step S10 includes: in response to the image generation prompt, map the image generation prompt to the latent space to obtain the latent space representation, and input the latent space representation into the preset image generation model; correspondingly, the step S30 includes: based on the latent space representation, perform image inference in the multiple inference stages in sequence to generate a target image.
[0076] It can be understood that the latent space can refer to a high-dimensional mathematical space used to represent the latent features or representations of input data (such as image generation prompts). In this space, similar input data will be mapped to nearby points or vectors. The latent space representation can refer to the low-dimensional feature representation of the image after compression encoding, such as the low-resolution tensor output by the VAE (Variational Autoencoder) encoder, which is used to reduce the computational complexity.
[0077] In a specific implementation, a text encoder (such as BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), etc.) is used to convert the image generation prompt into a vector representation in the latent space. Then, this latent space representation is passed as input to a preset image generation model.
[0078] Furthermore, to process inputs in different languages, the step of mapping the image generation prompt to the latent space in response to the image generation prompt, obtaining a latent space representation, and inputting the latent space representation into a preset image generation model includes: in response to the image generation prompt, extracting semantic information of the image generation prompt through a text encoder, where the text encoder supports texts in different languages; based on the semantic information, mapping the image generation prompt to a latent space matching the preset image generation model through the text encoder, obtaining a latent space representation, and inputting the latent space representation into the preset image generation model, where after texts in different languages are encoded, they are mapped to the same latent space.
[0079] It should be understood that a text encoder can refer to a neural network module that converts text input into a semantic vector in the latent space. Its core function is to extract the semantic information of the text and map it to a latent space representation compatible with the image generation model. For example, CLIP multilingual version and XLM-R. Among them, the CLIP multilingual version can refer to a multilingual text-image alignment model based on contrastive learning, which supports text encoding in multiple languages. XLM-R. XLM-R can refer to a pre-trained model dedicated to multilingual understanding (such as an extended version of RoBERTa), which can extract cross-language shared semantic features. The Text Encoder needs to meet the following two points: 1. Multilingual compatibility: Support input of texts in different languages (such as Chinese, English, Spanish, etc.); 2. Latent space alignment: After texts in different languages are encoded, they need to be mapped to the same latent space (that is, the latent vectors corresponding to semantically similar texts in different languages are close in distance).
[0080] In a specific implementation, after the user inputs an image generation prompt, the text encoder encodes the prompt, extracts the semantic information therein, and the text encoder maps the extracted semantic information to the latent space to obtain a latent space representation. Then, the latent space representation is input into a preset image generation model.
[0081] In this embodiment, a preset image generation model divides the image generation stage into multiple inference stages from the dimensions of noise level and image resolution, and performs image generation in stages, thereby reducing the computational complexity, reducing the inference time and the video memory occupied, and further reducing the application cost of the preset image generation model.
[0082] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the second embodiment of the image generation method of this application. Based on the first embodiment shown above Figure 1 , the second embodiment of the image generation method of this application is proposed.
[0083] In the second embodiment, step S30 includes:
[0084] Step S301: In each inference stage, perform image inference according to the image generation prompt and the output of the previous inference stage to generate an image of the current inference stage.
[0085] It should be understood that each inference stage corresponds to a different noise level range and a different image resolution. During the sequential execution of multiple inference stages, the noise is gradually reduced from high noise to no noise, and the resolution is gradually upsampled from low resolution to high resolution. During the sequential execution of multiple inference stages, the number of image generation modules used in each inference stage increases sequentially.
[0086] It should be noted that the image generation module used in each inference stage is the core generation module of the target model, where the target model is a diffusion model based on the Transformer architecture (such as SD3, Flux, etc.), and the image generation module is a Transformer module (such as the MM-DiT module of SD3, etc.).
[0087] Step S302: Upsample the output of the current inference stage and use it as the input of the next inference stage until a target image with a target resolution is generated.
[0088] For ease of understanding, refer to Figure 3 for illustration, but it does not limit this application. Figure 3 which is a structural diagram of the spatio-temporal pyramid model of an embodiment of the image generation method of this application. As an example, Figure 3Among them, it is assumed that the image generation module used by the spatio-temporal pyramid model in each inference stage is the MM-DiT module of SD3. Among them, MM-DiT (Multi-Modal Diffusion Transformer) is the core module of the SD3 model, supporting multi-modal inputs (such as text, images). The multi-modal diffusion Transformer adopted in the SD3 model generates images by fusing text semantics (output of the Text Encoder) and image latent variables (Latent). The Noise level (from 1 to 0) corresponds to the time dimension, and it can be evenly or manually divided among each stage. The Model Blocks correspond to the model space dimension and are divided in a way of stacking by resolution.
[0089] Assume that the image generation prompt is "a picture of a black motorcycle". The image generation process is divided into multiple stages (such as 3 stages) by the spatio-temporal pyramid model. Each stage corresponds to a specific resolution (256→512→1024) and a range of noise levels (such as Noise Level 1→0.6, 0.6→0.3, 0.3→0). At the starting stage of inference, the input is pure noise with a resolution of 256. Inference is performed under the block and noise level of the first stage, and then the upsampled inference result is used as the input for the second stage. Similarly, the inference of the second and third stages is carried out in turn, and finally an image with a resolution of 1024 is obtained. Performing image inference through multiple inference stages in sequence based on the image generation prompt specifically includes:
[0090] 1. Stage 1 (256x256 resolution), using 1 MM-DiT module:
[0091] Input: 256x256 pure noise Latent + text encoding vector ("a picture of a black motorcycle").
[0092] Processing: Denoise within the range of Noise Level 1→0.6 to generate a rough motorcycle contour (such as the black body, wheel positions).
[0093] Time consumption: Approximately 1 second (fast calculation at low resolution), using 2 MM-DiT modules.
[0094] 2. Stage 2 (512x512 resolution):
[0095] Input: The 256x256 Latent output from Stage 1 upsampled to 512x512 + text encoding vector.
[0096] Processing: Refine details (such as headlights, tire textures) within the range of Noise Level 0.6→0.3.
[0097] Time consumption: about 2 seconds (increased computational load for medium resolution).
[0098] 3. Stage 3 (1024x1024 resolution), using 3 MM-DiT modules:
[0099] Input: Upsample the 512x512 Latent output from Stage 2 to 1024x1024 + text encoding vector.
[0100] Processing: Generate high-definition details (such as metal reflections, background environment) within the range of Noise Level 0.3 → 0.
[0101] Time consumption: about 4 seconds (highest computational load for high resolution).
[0102] Traditional DiT models (such as SD3) take about 10 seconds to directly generate 1024x1024 images and occupy 12GB of video memory. However, in this application, the image generation process of the traditional DiT model is divided into multiple inference stages through a spatio-temporal pyramid model, and image generation is performed in stages. The total time consumption is about 7 seconds (1 + 2 + 4), and only 6GB of video memory is required (releasing intermediate results in stages).
[0103] In each inference stage of this embodiment, image inference is performed according to the image generation prompt and the output of the previous inference stage to generate the image of the current inference stage. The output of the current inference stage is upsampled and used as the input of the next inference stage until the target image with the target resolution is generated, so that a high-quality target image that meets the user's prompt can be obtained through step-by-step inference and generation.
[0104] Refer to Figure 4 , Figure 4 is the flowchart of the third embodiment of the image generation method of this application. Based on the above embodiments, the third embodiment of the image generation method of this application is proposed.
[0105] In the third embodiment, before the step S10, it further includes:
[0106] Step S01: Divide the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model.
[0107] It should be understood that, in order to improve the reliability of the preset image generation model, in this embodiment, the image generation process is also divided into multiple training stages from the first dimension and the second dimension through the initial image generation model, and the initial image generation model is trained based on image samples in multiple training stages in sequence to obtain the preset image generation model.
[0108] It can be understood that the initial image generation model can refer to an image generation model that is untrained or incompletely trained, which is the starting point of the image generation process. In this embodiment, the initial image generation model can refer to the image generation module in the spatio-temporal pyramid model that is untrained or incompletely trained. Among them, the image generation module is the core generation module of the target model, and the target model is a diffusion model based on the Transformer architecture (such as SD3, Flux, etc.), and the image generation module is a Transformer module (such as the MM-DiT module of SD3, etc.). The training stage can refer to the process in which the model learns how to generate high-quality images based on input data (such as image samples). In this stage, the model optimizes its performance by continuously adjusting its internal parameters.
[0109] Further, in order to self-adjust the number of stages, step S01 includes: obtaining the model deployment requirements, and determining the number of stage divisions according to the model deployment requirements; based on the number of stage divisions, dividing the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model.
[0110] It should be understood that dividing the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model based on the number of stage divisions includes: dividing multiple noise level ranges and multiple resolutions according to the number of stage divisions; based on the multiple noise level ranges and multiple resolutions, dividing the image generation process into multiple training stages through the initial image generation model.
[0111] Step S02: Based on the image samples, sequentially perform the multiple training stages to train the initial image generation model to obtain a preset image generation model.
[0112] In a specific implementation, in each training stage, the model is trained using image samples that match the noise level and resolution of that stage. By sequentially performing multiple training stages, the model gradually learns how to generate high-quality images.
[0113] It should be understood that training the initial image generation model based on the image samples sequentially through multiple training stages to obtain a preset image generation model can be based on the image samples sequentially through multiple training stages, and obtaining the loss value of each training stage; training the initial image generation model based on the loss value of each training stage to obtain a preset image generation model. In a specific implementation, training the initial image generation model based on the loss value of each training stage to obtain a preset image generation model can be adjusting the parameters of the image generation module used in each training stage based on the loss value of each training stage to obtain a preset image generation model.
[0114] Further, to optimize all stage model parameters simultaneously, in this embodiment, the model parameters of the initial image generation model can also be adjusted based on the total loss value of each training stage. Training the initial image generation model based on the loss values of each training stage to obtain a preset image generation model includes: performing weighted summation on the loss values of each training stage to obtain a total loss value; adjusting the model parameters of the initial image generation model based on the total loss value to obtain a preset image generation model.
[0115] In a specific implementation, when this embodiment is incorporated into the sd3 model, the training of the MM-DiT module is changed to be performed under a spatio-temporal pyramid. Taking the overall training being divided into 3 stages in terms of time and space and generating an image with a resolution of 1024 as an example, the corresponding resolution for each stage is 256, 512, and 1024. The Noise level (from 1 to 0) corresponds to the time dimension, and it can be evenly or manually divided among each stage. The model blocks correspond to the model space dimension and are divided in a way of stacking according to the resolution. The input at the starting stage of training is the latent of an image with a resolution of 256. The starting point of each subsequent stage is the latent after upsampling the end point of the previous stage. The target of each stage is the latent of an image with the resolution of this stage corresponding to the noise level. The losses of the three stages are summed up and trained simultaneously. In practical applications, the DiT structure is not limited, and the number of stages can be adjusted by oneself.
[0116] The training stage includes the following steps: 1. Stage division: Divide the training into N stages (such as 3 stages), and each stage corresponds to a specific resolution (256 → 512 → 1024) and a noise range (Noise Level segmentation). Example: When generating a 1024px image, stage 1 processes 256px (Noise Level 1 → 0.6), stage 2 processes 512px (0.6 → 0.3), and stage 3 processes 1024px (0.3 → 0). 2. Input and target: Input of stage 1: The latent variable of 256px (downsampled after the original image is encoded by VAE). Input of stage 2: The latent variable of 256px generated in stage 1 is upsampled to 512px. Input of stage 3: The latent variable of 512px generated in stage 2 is upsampled to 1024px. Target: The latent variable generated in each stage needs to approximate the real latent variable corresponding to the resolution and noise level. 3. Joint training: All stage model parameters are optimized simultaneously, and the total loss is the sum of the losses of each stage (such as diffusion loss + reconstruction loss).
[0117] For ease of understanding, the following is an example, but it does not limit this application. As an example, assume that the image generation process is divided into multiple stages (such as 3 stages) through a spatio-temporal pyramid model. Each stage corresponds to a specific resolution (256→512→1024) and a range of noise levels (such as Noise Level 1→0.6, 0.6→0.3, 0.3→0). The model training steps include:
[0118] 1. Data preparation:
[0119] Collect high-definition images (1024px) of "black motorcycles", generate 1024px latent variables through VAE encoding, and downsample to obtain 512px and 256px latent variables.
[0120] 2. Stage 1 training (256px):
[0121] Input: 256px latent variable (adding noise with Noise Level 1→0.6).
[0122] Objective: The model needs to recover the 256px motorcycle contour (such as the body shape and wheel position) from the noise.
[0123] Loss calculation: Mean Square Error (MSE) between the predicted noise and the true noise.
[0124] 3. Stage 2 training (512px):
[0125] Input: The 256px latent variable generated in Stage 1 is upsampled to 512px (adding noise with Noise Level 0.6→0.3).
[0126] Objective: Generate 512px details (such as headlights and tire textures).
[0127] Loss calculation: The same as above, but based on the 512px latent variable.
[0128] 4. Stage 3 training (1024px):
[0129] Input: The 512px latent variable generated in Stage 2 is upsampled to 1024px (adding noise with Noise Level 0.3→0).
[0130] Objective: Generate high-definition details (such as metal reflections and background environment).
[0131] Loss calculation: Based on the 1024px latent variable.
[0132] 5. Joint optimization:
[0133] Total loss = Stage 1 Loss + Stage 2 Loss + Stage 3 Loss, and all parameters are updated by backpropagation.
[0134] In this embodiment, the initial image generation model divides the image generation process into multiple training stages from the first dimension and the second dimension, and trains the initial image generation model through multiple training stages in sequence based on image samples to obtain a preset image generation model, thereby improving the reliability of the preset image generation model.
[0135] In summary, the beneficial effects brought by this application are as follows: 1. The spatio-temporal pyramid structure is designed, which can achieve high-quality text-to-image effects based on the mainstream DiTs. 2. Due to the optimization in time steps, model space, and input tokens, the spatio-temporal pyramid structure can save costs during training and inference. 3. Combining different text encoders can realize image generation in different languages.
[0136] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image generation method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0137] This application also provides an image generation device. Please refer to Figure 5 , the image generation device includes:
[0138] An information input module 10, configured to input the image generation prompt into a preset image generation model in response to an image generation prompt;
[0139] A process division module 20, configured to divide the image generation process into multiple inference stages from the first dimension and the second dimension through the preset image generation model, where the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution;
[0140] An image inference module 30, configured to perform image inference through the multiple inference stages in sequence based on the image generation prompt to generate a target image.
[0141] The image generation device provided by this application adopts the image generation method in the above embodiment, which can solve the technical problem that although the existing image generation methods can generate high-quality images, the inference takes a long time and occupies a large amount of video memory, resulting in a high application cost. Compared with the prior art, the beneficial effects of the image generation device provided by this application are the same as those of the image generation method provided by the above embodiment, and other technical features in the image generation device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0142] The present application provides an image generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image generation method in the first embodiment above.
[0143] Reference is made below Figure 6 , which shows a schematic structural diagram of an image generation device suitable for implementing the embodiments of the present application. The image generation device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions, tablet computers), PMPs (Portable Media Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The image generation device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0144] As Figure 6 shown, the image generation device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can execute various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the image generation device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image generation device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image generation device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.
[0145] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0146] The image generation device provided by the present application adopts the image generation method in the above embodiments, and can solve the technical problem that although the existing image generation methods can generate high-quality images, the inference takes a long time and occupies a large amount of video memory, resulting in a relatively high application cost. Compared with the prior art, the beneficial effects of the image generation device provided by the present application are the same as those of the image generation method provided by the above embodiments, and other technical features in the image generation device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0147] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0148] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0149] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image generation method in the above embodiments.
[0150] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0151] The above computer-readable storage medium can be included in an image generation device; it can also exist independently and not be assembled into the image generation device.
[0152] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by an image generation device, the image generation device is caused to execute the above image generation method.
[0153] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or WAN (Wide Area Network), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0155] The modules described in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0156] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above image generation method, and can solve the technical problem that although the existing image generation methods can generate high-quality images, the inference time is long and the video memory occupation is large, resulting in a relatively high application cost. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the image generation method provided by the above embodiments, and will not be elaborated here.
[0157] The present application also provides a computer program product, including a computer program, which when executed by a processor implements the image generation method as described above.
[0158] The computer program product provided by the present application can solve the technical problem that although the existing image generation methods can generate high-quality images, the inference time is long and the video memory occupation is large, resulting in a relatively high application cost. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the image generation method provided by the above embodiments, and will not be elaborated here.
[0159] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
[0160] A1. An image generation method, the image generation method comprising:
[0161] In response to an image generation prompt, input the image generation prompt into a preset image generation model;
[0162] Divide the image generation process into multiple inference stages from a first dimension and a second dimension by the preset image generation model, wherein the first dimension corresponds to a noise level and the second dimension corresponds to an image resolution;
[0163] Based on the image generation prompt, perform image inference in sequence through the multiple inference stages to generate a target image.
[0164] A2. The image generation method as described in A1, the performing image inference in sequence through the multiple inference stages based on the image generation prompt to generate a target image, comprising:
[0165] In each inference stage, perform image inference according to the image generation prompt and the output of the previous inference stage to generate an image of the current inference stage;
[0166] Use the output of the current inference stage after upsampling as the input of the next inference stage until a target image with a target resolution is generated.
[0167] A3. The image generation method as described in A2, each inference stage corresponds to a different noise level range and a different image resolution. During the sequential execution of the multiple inference stages, the noise is gradually reduced from high noise to no noise, and the resolution is gradually upsampled from low resolution to high resolution. During the sequential execution of the multiple inference stages, the number of image generation modules used in each inference stage increases in sequence.
[0168] A4. The image generation method as described in A3, the image generation module used in each inference stage is the core generation module of the target model, wherein the target model is a diffusion model based on the Transformer architecture, and the image generation module is a Transformer module.
[0169] A5. The image generation method as described in any one of A1 to A4, the in response to an image generation prompt, input the image generation prompt into a preset image generation model, comprising:
[0170] In response to an image generation prompt, map the image generation prompt to a latent space to obtain a latent space representation, and input the latent space representation into a preset image generation model;
[0171] Correspondingly, the performing image inference in sequence through the multiple inference stages based on the image generation prompt to generate a target image, comprising:
[0172] Based on the latent space representation, perform the multiple inference stages in sequence for image inference to generate a target image.
[0173] A6. The image generation method as described in A5, which, in response to an image generation prompt, maps the image generation prompt to a latent space to obtain a latent space representation, and inputs the latent space representation into a preset image generation model, includes:
[0174] In response to an image generation prompt, extract the semantic information of the image generation prompt through a text encoder, where the text encoder supports texts in different languages;
[0175] Based on the semantic information, map the image generation prompt to a latent space matching the preset image generation model through the text encoder to obtain a latent space representation, and input the latent space representation into the preset image generation model, where texts in different languages are mapped to the same latent space after encoding.
[0176] A7. The image generation method as described in any one of A1 to A4, before inputting the image generation prompt into the preset image generation model in response to the image generation prompt, further includes:
[0177] Divide the image generation process into multiple training stages from the first dimension and the second dimension through an initial image generation model;
[0178] Based on image samples, perform the multiple training stages in sequence to train the initial image generation model to obtain a preset image generation model.
[0179] A8. The image generation method as described in A7, which includes training the initial image generation model based on image samples through the multiple training stages in sequence to obtain a preset image generation model, including:
[0180] Based on image samples, perform the multiple training stages in sequence and obtain the loss value of each training stage;
[0181] Train the initial image generation model based on the loss values of the respective training stages to obtain a preset image generation model.
[0182] A9. The image generation method as described in A8, which includes training the initial image generation model based on the loss values of the respective training stages to obtain a preset image generation model, including:
[0183] Perform a weighted sum of the loss values of the respective training stages to obtain a total loss value;
[0184] Adjust the model parameters of the initial image generation model based on the total loss value to obtain a preset image generation model.
[0185] A10. The image generation method as described in A7, wherein the image generation process is divided into multiple training stages from the first dimension and the second dimension through an initial image generation model, including:
[0186] Obtain the model deployment requirements, and determine the number of stage divisions according to the model deployment requirements;
[0187] Based on the number of stage divisions, divide the image generation process into multiple training stages from the first dimension and the second dimension through an initial image generation model.
[0188] A11. The image generation method as described in A10, wherein dividing the image generation process into multiple training stages from the first dimension and the second dimension through the initial image generation model based on the number of stage divisions includes:
[0189] Divide multiple noise level ranges and multiple resolutions according to the number of stage divisions;
[0190] Based on the multiple noise level ranges and the multiple resolutions, divide the image generation process into multiple training stages through an initial image generation model.
[0191] The present application also discloses B12. An image generation device, the image generation device includes:
[0192] An information input module, configured to input the image generation prompt into a preset image generation model in response to an image generation prompt;
[0193] A process division module, configured to divide the image generation process into multiple inference stages from the first dimension and the second dimension through the preset image generation model, wherein the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution;
[0194] An image inference module, configured to perform image inference in sequence for the multiple inference stages based on the image generation prompt to generate a target image.
[0195] B13. The image generation device as described in B12, wherein the image inference module is further configured to, in each inference stage, perform image inference according to the image generation prompt and the output of the previous inference stage to generate an image of the current inference stage; upsample the output of the current inference stage and use it as the input of the next inference stage until a target image with a target resolution is generated.
[0196] B14. The image generation device as described in B13, wherein each inference stage corresponds to a different noise level range and a different image resolution. During the sequential execution of the multiple inference stages, the noise is gradually reduced from high noise to no noise, and the resolution is gradually upsampled from low resolution to high resolution. During the sequential execution of the multiple inference stages, the number of image generation modules used in each inference stage increases sequentially.
[0197] B15. The image generation device as described in B14, wherein the image generation modules used in each inference stage are the core generation modules of the target model. Among them, the target model is a diffusion model based on the Transformer architecture, and the image generation module is a Transformer module.
[0198] B16. The image generation device as described in any one of B12 to B15, wherein the information input module is further configured to, in response to an image generation prompt, map the image generation prompt to the latent space, obtain a latent space representation, and input the latent space representation into a preset image generation model.
[0199] Correspondingly, the image inference module is further configured to perform image inference through the multiple inference stages based on the latent space representation to generate a target image.
[0200] B17. The image generation device as described in B16, wherein the information input module is further configured to, in response to an image generation prompt, extract semantic information of the image generation prompt through a text encoder, where the text encoder supports texts in different languages; map the image generation prompt to a latent space matching the preset image generation model based on the semantic information through the text encoder, obtain a latent space representation, and input the latent space representation into the preset image generation model, wherein texts in different languages are mapped to the same latent space after being encoded.
[0201] The present application also discloses C18. An image generation device, the image generation device includes: a memory, a processor, and an image generation program stored on the memory and executable on the processor. When the image generation program is executed by the processor, it implements the image generation method as described above.
[0202] The present application also discloses D19. A storage medium, on which an image generation program is stored. When the image generation program is executed by a processor, it implements the image generation method as described above.
[0203] The present application also discloses E20. A computer program product, the computer program product includes an image generation program. When the image generation program is executed by a processor, it implements the image generation method as described above.
Claims
1. An image generation method, characterized in that, The described image generation method includes: In response to an image generation prompt, input the image generation prompt into a preset image generation model; Divide the image generation process into multiple inference stages from a first dimension and a second dimension through the preset image generation model, where the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution; Based on the image generation prompt, sequentially perform the multiple inference stages for image inference to generate a target image.
2. The image generation method according to claim 1, wherein The sequentially performing the multiple inference stages for image inference based on the image generation prompt to generate a target image includes: In each inference stage, perform image inference according to the image generation prompt and the output of the previous inference stage to generate an image of the current inference stage; Use the output of the current inference stage after upsampling as the input of the next inference stage until a target image with a target resolution is generated.
3. The image generation method according to claim 2, wherein Each of the inference stages corresponds to a different noise level range and a different image resolution. During the sequential execution of the multiple inference stages, the noise is gradually reduced from high noise to no noise, and the resolution is gradually upsampled from low resolution to high resolution. During the sequential execution of the multiple inference stages, the number of image generation modules used in each inference stage increases sequentially.
4. The image generation method according to claim 3, wherein The image generation module used in each inference stage is the core generation module of the target model, where the target model is a diffusion model based on the Transformer architecture, and the image generation module is a Transformer module.
5. The image generation method according to any one of claims 1 to 4, characterized in that, The step of, in response to an image generation prompt, inputting the image generation prompt into a preset image generation model includes: In response to an image generation prompt, map the image generation prompt to a latent space to obtain a latent space representation, and input the latent space representation into a preset image generation model; Correspondingly, the step of sequentially performing the multiple inference stages for image inference based on the image generation prompt to generate a target image includes: Based on the latent space representation, sequentially perform the multiple inference stages for image inference to generate a target image.
6. The image generation method according to claim 5, wherein The step of, in response to an image generation prompt, mapping the image generation prompt to a latent space to obtain a latent space representation, and inputting the latent space representation into a preset image generation model includes: In response to an image generation prompt, extract semantic information of the image generation prompt through a text encoder, where the text encoder supports texts in different languages; Based on the semantic information, map the image generation prompt to a latent space matching the preset image generation model through the text encoder to obtain a latent space representation, and input the latent space representation into a preset image generation model, where texts in different languages are mapped to the same latent space after encoding.
7. An image generation device, characterized in that, The described image generation device includes: An information input module, configured to input the image generation prompt into a preset image generation model in response to an image generation prompt; A process division module, configured to divide the image generation process into multiple inference stages from a first dimension and a second dimension through the preset image generation model, where the first dimension corresponds to the noise level and the second dimension corresponds to the image resolution; An image inference module, configured to perform image inference by sequentially performing the multiple inference stages based on the prompts generated from the image, and generate a target image.
8. An image generation device, characterized in that, The image generation device includes: a memory, a processor, and an image generation program stored on the memory and executable on the processor. When the image generation program is executed by the processor, the image generation method according to any one of claims 1 to 6 is implemented.
9. A storage medium, characterized in that, An image generation program is stored on the storage medium. When the image generation program is executed by a processor, the image generation method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that, The computer program product includes an image generation program. When the image generation program is executed by a processor, the image generation method according to any one of claims 1 to 6 is implemented.