Image generation method, device, equipment and program product

By compressing and characterizing the original image, and combining a denoising module and a deep compression autoencoder, the structure of the image generation model is optimized, solving the problem of high computational complexity of the diffusion model. This achieves efficient and real-time image generation, making it suitable for scenarios such as virtual digital human video generation.

CN120976050APending Publication Date: 2025-11-18ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511105292.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing diffusion models suffer from high computational complexity and slow generation speed in image generation tasks, failing to meet the real-time and efficiency requirements of scenarios such as virtual digital human video generation.

Method used

The original image is compressed and characterized using a compression factor, and inverse diffusion is performed through a denoising module. The structure and training process of the image generation model are optimized by combining a deep compression autoencoder and a decoder. Neural network architectures such as U-Net and lightweight Swing Transformer are used to reduce computation and improve image generation speed.

Benefits of technology

While ensuring image quality, it significantly improves the speed and efficiency of image generation, enabling high real-time performance and high efficiency in various work scenarios, and meeting the needs of long-duration, high-frame-rate video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976050A_ABST
    Figure CN120976050A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an image generation method and device, equipment and a program product. The method comprises the following steps: inputting an original image into an image generation model; through an image generation model, compression and characterization processing are carried out on an original image based on a compression factor to obtain a first image feature, a denoising module obtained through adjustment based on the compression factor carries out inverse diffusion processing on the first image feature to obtain a second image feature, and image reconstruction and decompression processing are carried out based on the second image feature. A target image output by an image generation model is obtained, and the image generation model comprises a denoising module. The image generation speed can be increased, and the real-time performance and the high efficiency of image generation in various operation scenes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to an image generation method, apparatus, device, and program product. Background Technology

[0002] With the development of Virtual Reality (VR) and Augmented Reality (AR) technologies, virtual digital humans have gradually become a research hotspot. Virtual digital humans are widely used in entertainment, education, and healthcare, among other fields, and one of their core technologies is the deep learning-based image generation model. Currently, diffusion models are an important framework for image generation, possessing advantages such as realistic generation effects and strong noise reduction capabilities. However, diffusion models often suffer from high computational complexity and slow generation speed in image generation tasks. Especially in scenarios such as virtual digital human video generation, a basic requirement for high-quality video generation includes a generation rate of 25 frames per second (fps). However, existing diffusion model-based algorithms take an average of 2-3 seconds to generate each frame, meaning that generating a 10-second virtual digital human video would take approximately 10 minutes, failing to meet the real-time and efficiency requirements for image generation in scenarios such as virtual digital human video generation. Summary of the Invention

[0003] Based on the defects and shortcomings of the existing technology, this application proposes an image generation method, apparatus, device and program product, which can improve the image generation speed and enhance the real-time performance and efficiency of image generation in various work scenarios.

[0004] According to a first aspect of this application, an image generation method is provided, comprising: inputting an original image into an image generation model; compressing and characterizing the original image based on a compression factor to obtain a first image feature through the image generation model; performing inverse diffusion processing on the first image feature by a denoising module adjusted based on the compression factor to obtain a second image feature; and performing image reconstruction and decompression processing based on the second image feature to obtain a target image output by the image generation model, wherein the image generation model includes the denoising module.

[0005] According to the image generation method provided in the first aspect of this application, the denoising module includes a feature processing layer, a bottleneck layer, and an output layer. The feature processing layer includes at least one convolutional layer. The number of channels in the first convolutional layer is adapted to the feature dimension of the first image feature. The first convolutional layer is used to receive the first image feature. The step of obtaining a second image feature by performing inverse diffusion processing on the first image feature using the denoising module adjusted based on the compression factor includes: inputting the first image feature into the feature processing layer, whereby the feature processing layer performs feature extraction and feature dimension adjustment on the first image feature, and outputs a first intermediate feature; inputting the first intermediate feature into the bottleneck layer, whereby the bottleneck layer performs global feature modeling on the first intermediate feature, and outputs a second intermediate feature; and inputting the second intermediate feature into the output layer, whereby the output layer adjusts the feature dimension of the second intermediate feature, and outputs a second image feature, wherein the feature dimensions of the first image feature and the second image feature are the same.

[0006] According to the image generation method provided in the first aspect of this application, the bottleneck layer includes a hierarchical visual transformer module.

[0007] According to the image generation method provided in the first aspect of this application, the image generation model includes a deep compression autoencoder; the step of compressing and characterizing the original image based on a compression factor to obtain a first image feature includes: inputting the original image into the deep compression autoencoder, and having the deep compression autoencoder compress and characterize the original image to obtain the first image feature output by the deep compression autoencoder, wherein the deep compression autoencoder is an encoder trained based on the compression factor.

[0008] According to the image generation method provided in the first aspect of this application, the image generation model includes a decoder; the step of performing image reconstruction and decompression processing based on the second image features to obtain the target image output by the image generation model includes: inputting the second image features into the decoder, and having the decoder perform image reconstruction and decompression processing on the second image features to obtain the target image output by the decoder, wherein the original image and the target image have the same resolution, and the decoder is obtained by synchronous training based on the compression factor and the deep compression autoencoder.

[0009] According to the image generation method provided in the first aspect of this application, the training process of the image generation model is as follows: The original encoder and the original decoder are trained synchronously using first sample data. When the first training metric meets the first requirement, the internal parameters of the original encoder and the original decoder are fixed to obtain the deep compression autoencoder and the decoder. Based on the compression and feature processing of the second sample data using the deep compression autoencoder, and the image reconstruction and decompression processing of the compressed sample results output by the original generation model using the decoder, the original denoising module is trained using the second sample data. When the second training metric meets the second requirement, the denoising module is obtained.

[0010] According to the image generation method provided in the first aspect of this application, the denoising module includes a U-shaped convolutional neural network.

[0011] According to a second aspect of this application, an image generation apparatus is provided, comprising: an input module for inputting an original image into an image generation model; and an output module for compressing and characterizing the original image based on a compression factor using the image generation model to obtain a first image feature, performing inverse diffusion processing on the first image feature by a denoising module adjusted based on the compression factor to obtain a second image feature, and performing image reconstruction and decompression processing based on the second image feature to obtain a target image output by the image generation model, wherein the image generation model includes the denoising module.

[0012] According to a third aspect of this application, an electronic device is provided, comprising: a memory and a processor; the memory is connected to the processor and is used to store a program; the processor is used to implement the image generation method as described in the first aspect by running the program in the memory.

[0013] According to a fourth aspect of this application, a computer program product is provided, including computer program instructions; the computer program instructions, when executed by a processor, cause the processor to perform the image generation method as described in the first aspect.

[0014] In this application, the original image is input into an image generation model. The image generation model compresses and characterizes the original image based on a compression factor to obtain first image features. A denoising module, adjusted based on the compression factor, performs inverse diffusion processing on the first image features to obtain second image features. Image reconstruction and decompression are then performed based on the second image features to obtain the target image output by the image generation model. The image generation model includes a denoising module. In this process, the first image features are not only the characterized image data but also compressed based on a compression factor. When the denoising module processes the compressed first image features, the amount of data required for inverse diffusion processing is reduced, thus improving the processing speed of the denoising module. Furthermore, the denoising module is adjusted based on the compression factor, achieving compatibility between the denoising module and the first image features and avoiding redundant processing, further improving the image generation speed. Furthermore, image reconstruction and decompression processing on the second image features ensures the visibility of the target image. While ensuring image visibility, the image generation speed is further improved, guaranteeing high real-time performance and efficiency in various operational scenarios. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is one of the flowcharts illustrating an image generation method provided in an embodiment of this application;

[0017] Figure 2 An example diagram of image compression based on variational autoencoder provided in this application embodiment;

[0018] Figure 3 An example diagram of image compression based on a deep compression autoencoder provided in this application embodiment;

[0019] Figure 4 This is a second schematic flowchart illustrating an image generation method provided in an embodiment of this application.

[0020] Figure 5 A block diagram of an image generation apparatus provided in an embodiment of this application;

[0021] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] Exemplary methods

[0024] In scenarios where high image generation speed is required, such as virtual digital human generation, this application provides an image generation method to improve image generation speed. This method can be implemented in the form of a software algorithm. The software algorithm implementing this method can run on any device with data processing capabilities, such as smart mobile devices, local computers, remote servers, etc. The device type can be determined based on the actual working scenario and specific needs. The scope of protection of this application is not limited by the specific type of device.

[0025] In one embodiment, such as Figure 1 As shown, the process steps for implementing the image generation method include:

[0026] Step 101: Input the original image into the image generation model.

[0027] In this embodiment, the original image refers to the initial image required for the image generation task. The image generation model is a pre-defined model for the image generation task, which can complete the image generation task based on the original image to obtain the final generated target image. Optionally, the image generation model can be obtained in advance through multiple stages such as structure setting, parameter adjustment and / or training.

[0028] Step 102: The original image is compressed and characterized by an image generation model based on a compression factor to obtain a first image feature. The first image feature is then subjected to inverse diffusion processing by a denoising module adjusted based on the compression factor to obtain a second image feature. The second image feature is then used for image reconstruction and decompression to obtain the target image output by the image generation model. The image generation model includes a denoising module.

[0029] In this embodiment, a crucial component of the image generation model is a denoising module. This module primarily performs inverse diffusion processing on the features to achieve reverse generation, thereby obtaining new second image features based on the known first image features. Optionally, the denoising module can improve and train the image generation model based on the diffusion model. Diffusion models are a type of generative model that learns the distribution of data by adding noise during the training phase, and then generates new data samples from the noise through a reverse generation process after the model training is complete.

[0030] In this embodiment, the original image is compressed according to a compression factor based on its original resolution to obtain a first image feature. This first image feature has a significantly reduced resolution compared to the original image. For example, if the original image size is 512×512 pixels and the compression factor is 32, then the size of the first image feature is 16×16 pixels. It should be noted that the specific value of the compression factor is set according to the actual situation and needs. During image generation, compared to the original image, performing reverse diffusion processing on the first image feature to generate the second image feature can significantly reduce the computational load and improve the image generation speed. Furthermore, the larger the compression factor, the faster the image generation speed. However, in practical applications, it is necessary to reasonably configure the compression factor based on the specific scenario and actual situation to ensure the overall effect and balance of image generation.

[0031] In this embodiment, based on the compression factor used when obtaining the first image features, the image compression structure in the denoising module is rationally simplified when constructing the denoising module based on the diffusion model. This avoids redundant parts in the denoising module, making the first image features and the denoising module more compatible. After the first image features are processed by the denoising module, the second image features are obtained. These second image features are new image features generated based on the information contained in the first image features. Optionally, after simplifying the denoising module, the initial generation module containing the denoising module needs to be further trained. Only after training is completed can the image generation model be obtained.

[0032] In this embodiment, since the denoising module is a modified version of the diffusion model, the second image feature obtained by the denoising module has the same resolution as the first image feature, as is known from the basic principle of image generation based on the diffusion model. Because the resolution of the second image feature is the same as the first image feature, it means that the resolution of the second image feature is low. This low resolution not only affects the visibility of the second image feature, but also fails to meet the requirements of higher resolution image processing steps if further image processing is needed based on the second image feature. Therefore, the second image feature is decompressed to improve the resolution of the decompressed target image and enhance its visibility. Optionally, the second image feature can be decompressed using the compression factor used when obtaining the first image feature, which is beneficial for meeting the higher resolution requirements in subsequent processing of the target image.

[0033] Optionally, the compression processing of the original image, the structural adjustment of the denoising module, and the decompression processing of the second image features are all implemented based on the same compression factor. This not only ensures the compatibility between the first image features and the denoising module, accelerating image generation, but also guarantees the high resolution of the target image. The accelerated image generation speed and the high resolution of the target image not only ensure high real-time performance and efficiency in image generation under various operational scenarios, but also guarantee higher usability of the target image.

[0034] In one embodiment, the denoising module includes a U-shaped convolutional neural network (U-Net) model.

[0035] In this embodiment, U-Net is a deep learning architecture commonly used for image segmentation tasks. Its U-shaped structure includes a downsampling path (shrinking path) and an upsampling path (expanding path). Skip connections combine low-level feature information with high-level semantic information, achieving high-precision pixel-level prediction. The denoising module primarily handles denoising and data reconstruction, often modified from a diffusion model. The diffusion model is a generative model whose core idea is to progressively add noise to the data and then train a neural network to reverse this process (i.e., denoising), thereby generating new samples. The training process of the diffusion model can be divided into two stages: forward diffusion, where Gaussian noise is progressively added to the input data until it becomes indistinguishable; and reverse generation, where a neural network is trained to predict and remove noise, progressively reconstructing clear data from random noise. For the diffusion model, U-Net is a commonly used neural network architecture for implementing the denoising task in the reverse generation process. U-Net's skip connections combine low-level detail information with high-level semantic information, which is crucial for generating high-quality images. Its symmetric encoder-decoder structure is well-suited for pixel-level tasks, such as the image generation task in this application. In other words, the diffusion model provides the framework for image generation models, defining how to model data distribution through forward diffusion and backward generation; while U-Net provides the implementation tools, serving as the core neural network in the diffusion model, responsible for performing the denoising task. The combination of U-Net's powerful expressive capabilities and the theoretical framework of the diffusion model results in excellent performance in image generation models in terms of image quality and diversity.

[0036] In one embodiment, the denoising module includes a feature processing layer, a bottleneck layer, and an output layer. The feature processing layer includes at least one convolutional layer. The number of channels in the first convolutional layer is adapted to the feature dimension of the first image feature. The first convolutional layer is used to receive the first image feature.

[0037] The denoising module, based on compression factor adjustment, performs inverse diffusion processing on the first image features to obtain the second image features. This includes: inputting the first image features into a feature processing layer, where the feature processing layer extracts features and adjusts the feature dimensions of the first image features, and outputs a first intermediate feature; inputting the first intermediate feature into a bottleneck layer, where the bottleneck layer performs global feature modeling on the first intermediate feature, and outputs a second intermediate feature; and inputting the second intermediate feature into an output layer, where the output layer adjusts the feature dimensions of the second intermediate feature, and outputs the second image feature. The first image features and the second image features have the same feature dimensions.

[0038] In this embodiment, the internal structure of the denoising module is divided into three parts: a feature processing layer, a bottleneck layer, and an output layer. The feature processing layer is used to initially extract features from the first image features. Optionally, the feature processing layer uses a commonly used encoder. The bottleneck layer is used to further extract global features based on the initially extracted first intermediate features. The output layer is used to decode the extracted features and output the second image features. Optionally, the output layer uses a commonly used decoder.

[0039] In this embodiment, taking the denoising module as an example, which includes a U-shaped convolutional neural network (U-Net) model, the U-Net model is simplified and modified to obtain the denoising module, and then the image generation model is obtained. Optionally, the U-Net model uses MobileNetV3 as the network backbone. MobileNetV3 is an efficient convolutional neural network architecture that can achieve high-performance image classification and other vision tasks with limited computing resources. MobileNetV3 uses inverted residuals, where the expand layer first increases the number of channels and then reduces the number of channels through depthwise separable convolutions. This structure helps to improve the efficiency of information flow while maintaining low computational overhead. MobileNetV3 introduces new non-symmetric activation functions, such as Hard-Swish and ReLU6, which are easier to implement and more computationally efficient on mobile devices. MobileNetV3 can also introduce a Squeeze-and-Excitation Module (SE module) in some layers. This module enhances the model's attention to different features and improves its representation ability by adaptively reweighting channel features. The SE module helps capture global information and improves the model's generalization ability.

[0040] The original MobileNetV3 architecture is a lightweight network primarily designed for targeted processing of color images (i.e., RGB images, which are 3-channel images). The input is an image of size (H, W, 3), where H represents the image height, W represents the image width, and 3 represents 3 channels. For 3-channel images, the first layer in the original MobileNetV3 architecture is an initial convolutional layer that increases the number of channels from 3 to 16. The following two layers, bottleneck2 and bottleneck3, perform low-level feature extraction on these 16 channels and then convert them back to 24 channels. The bottleneck layer is one of its core components, primarily using depthwise separable convolution and 1x1 convolution to reduce computation and improve efficiency. bottleneck2 and bottleneck3 are different types of bottleneck layers.

[0041] For the image generation method provided in this embodiment, the first image feature input to the denoising module is a compressed first image feature. The image size and number of channels have already been adjusted based on the compression factor. Therefore, there is no need to perform an initial convolution step or preliminary Bottleneck layer processing within the denoising module. Thus, the denoising module, based on the original MobileNetV3 architecture, removes the first initial convolutional layer and the first few Bottleneck layers, and directly extracts deep semantic features. Optionally, the actual number of Bottleneck layers removed is determined by the compression factor. That is, based on the specific number of channels of the first image feature after compression factor processing, the structure of the feature processing layer in the image generation model is adjusted so that the number of channels of the first image feature is consistent with the initial number of image channels that the feature processing layer can process. This allows the feature processing layer to directly extract features from the input first image feature without additional convolution operations. This adjustment to the image generation model structure reduces the data computation of the image generation model and improves the image generation speed.

[0042] In one embodiment, the bottleneck layer includes a hierarchical visual transformer (Swin Transformer) module.

[0043] In this embodiment, the Swin Transformer is a Transformer-based visual model that processes image data through a hierarchical shifted window mechanism, thereby effectively capturing local and global features. The lightweight Swin Transformer includes the following key features:

[0044] First, the layered shift window mechanism:

[0045] Local window partitioning divides the image into non-overlapping local windows and applies a self-attention mechanism within each window. This approach reduces computational complexity and enables the model to efficiently capture local features. Shifted windows introduce shifted windows at different levels to enhance the model's ability to perceive cross-window regions, thereby capturing a wider range of contextual information. This helps improve the model's representation ability and generalization performance.

[0046] Second, lightweight design:

[0047] Simplifying the network structure reduces the number of model parameters and computational overhead by reducing the number of network layers and channels. For example, fewer Transformer blocks or smaller hidden dimensions can be used. Efficient attention mechanisms, such as sparse self-attention or other optimization techniques, further reduce computational costs. For instance, self-attention calculations within a local window only need to consider the pixels within that window, rather than all pixels in the entire image. Low-rank decomposition reduces the number of parameters and speeds up inference by performing low-rank decomposition on the weight matrix. This approach can significantly reduce the model size without significantly sacrificing performance.

[0048] Third, multi-scale feature fusion:

[0049] Cross-level feature interaction: By introducing a cross-level feature fusion module, features at different resolutions can complement each other, improving the overall performance of the model. For example, high-resolution features can provide more detailed information, while low-resolution features can provide a wider range of spatial context. Multi-scale input: Supports multi-scale input images to better adapt to different application scenarios and hardware conditions.

[0050] Fourth, accelerate reasoning:

[0051] Knowledge distillation extracts useful knowledge from large pre-trained models and transfers it to lightweight models, which not only accelerates the inference process but also maintains high accuracy. Quantization and pruning combine techniques such as model quantization and pruning to further compress model size and improve running efficiency.

[0052] In this embodiment, the first intermediate feature obtained by the feature processing layer through feature extraction of the first image features mainly consists of the latent features contained in the first image features. In the specific implementation of the original U-Net model, the bottleneck layer is generally composed of several convolutional layers, used to further process the features obtained from the feature processing layer and generate the necessary input for the output layer. However, the Swin Transformer module replaces the traditional convolution operation, which can optimize the ability to capture global contextual information. In order to further improve the image generation quality and enable the model to capture long-distance dependencies in low-dimensional space, a lightweight Swin Transformer is embedded in the bottleneck layer to capture global features in the first image features, improve the understanding of distant pixels, and enhance the denoising effect.

[0053] In one specific embodiment, the denoising module is based on the U-Net model, with the lightweight MobileNetV3 structure as the network backbone, and the structure is adjusted based on the compression factor.

[0054] In this embodiment, taking an image with the first image feature (H, W, 32) as an example, the principle and process of the denoising module after U-Net structure adjustment processing the first image feature are as follows:

[0055] First, the first image feature (H,W,32) is input to the feature processing layer. After compression, the first image feature includes latent features.

[0056] Then, the feature processing layer extracts features, as follows:

[0057] 3x3 Conv+ReLU:

[0058] Depthwise Separable Conv(H / 2,W / 2,C1)

[0059] Depthwise Separable Conv(H / 4,W / 4,C2)

[0060] Bottleneck Block (H / 8, W / 8, C3)

[0061] Bottleneck Block(H / 16,W / 16,C4)

[0062] The feature processing layer employs a 4-layer convolution (Conv) with 3x3 kernels, using a rectified linear unit (ReLU) as the activation function. The 4 downsampling layers correspond to image heights of H / 2, H / 4, H / 8, and H / 16, and image widths of W / 2, W / 4, W / 8, and W / 16, respectively. The number of channels doubles from 32, denoted as C1, C2, C3, and C4. Depthwise Separable Conv represents depthwise separable convolution, primarily used for feature extraction to improve computational efficiency. The bottleneck block refers to a special structure in ResNet that reduces computational complexity by decreasing the dimensionality of intermediate layers while maintaining high model performance.

[0063] Next, the bottleneck layer performs global feature modeling. Optionally, the bottleneck layer embeds a Swin Transformer module to improve the denoising effect, as follows:

[0064] Swin Transformer(Window Attention)

[0065] The Swin Transformer module, based on the window-based self-attention (Window Attention) mechanism, processes data through a layered shifted window mechanism.

[0066] Finally, the output layer performs feature decoding to output the final denoising latent features, namely the second image features. The image size of the second image features is (H, W, 32), which is the same as the size of the first image features, as detailed below:

[0067] PixelShuffle or ConvTranspose2d(H / 8,W / 8,C3)

[0068] Skip Connection+Concat(H / 8,W / 8,C3⊕C3)

[0069] PixelShuffle or ConvTranspose2d(H / 4,W / 4,C2)

[0070] Skip Connection+Concat(H / 4,W / 4,C2⊕C2)

[0071] PixelShuffle or ConvTranspose2d(H / 2,W / 2,C1)

[0072] Skip Connection+Concat(H / 2,W / 2,C1⊕C1)

[0073] 1x1 Conv+Sigmoid / Tanh

[0074] Pixel Shuffle (Sub-Pixel Convolutional Neural Network) is a deep learning-based upsampling method that transforms low-resolution feature maps into high-resolution images by rearranging pixels in the feature map, effectively enlarging the scaled-down feature map. ConvTranspose2d is a 2D transposed convolutional layer that enlarges the spatial dimensions (width and height) of the input feature map through convolution operations, typically used for upsampling tasks. In the U-Net network, ConvTranspose2d is used to progressively reconstruct the details of the input image by enlarging the feature map and combining it with skip connections. U-Net uses skip connections to concatenate the high-resolution feature maps of the feature processing layer with the low-resolution feature map set of the output layer, thus preserving more detailed information. Low-level features help recover the latent features after denoising, ensuring that the final output image retains detailed information. ⊕ indicates a connection operation. Finally, a 1×1 Conv convolution kernel is used for channel adjustment to ensure that the output remains (H, W, 32). The sigmoid function or the hyperbolic tangent function Tanh is used as the activation function to maintain a stable numerical range. During the output layer processing, after three upsampling operations and one channel adjustment, the image height is restored to H / 8, H / 4, H / 2, and H respectively, the image width is restored to W / 8, W / 4, W / 2, and W respectively, and the number of channels is restored from C4 to C3, C2, and C1 (i.e., 32).

[0075] In this embodiment, the image generation model obtained by adjusting U-Net can significantly reduce the time required for the denoising process; using MobileNetV3 as the network backbone can reduce the amount of computation and improve efficiency; embedding a lightweight Swin Transformer in the bottleneck layer of U-Net allows the model to capture long-distance dependencies in low-dimensional space, thereby improving the quality of image generation.

[0076] In one embodiment, the process of compressing the original image based on the compression factor can be accomplished using any of the following compression techniques: arithmetic coding, Huffman coding, deep learning-based image compression, or other image compression techniques. The chosen compression technique is sufficient to satisfy the image generation process provided in this application.

[0077] In one embodiment, the image generation model includes a deep compression autoencoder.

[0078] The first image feature is obtained by compressing and characterizing the original image based on the compression factor, including: inputting the original image into a deep compression autoencoder, and performing compression and characterization on the original image by the deep compression autoencoder to obtain the first image feature output by the deep compression autoencoder, wherein the deep compression autoencoder is an encoder trained based on the compression factor.

[0079] In this embodiment, the Deep Compressed Autoencoder (DC-AE) is a neural network architecture that combines autoencoder and model compression techniques. It is primarily used for efficient feature extraction, data dimensionality reduction (including reducing the resolution of the original image to the resolution of a first image feature), and data reconstruction. An autoencoder is an unsupervised learning model whose core idea is to compress input data into a low-dimensional latent space representation using an encoder, and then reconstruct the original data from this latent space using a decoder. Model compression techniques aim to reduce the number of parameters and computational complexity of neural networks while maintaining model performance as much as possible. DC-AE combines these two approaches, learning not only a compact representation of the data during training but also optimizing the model structure to reduce storage and computational costs. DC-AE can extract more representative features; compared to traditional autoencoders, DC-AE has fewer model parameters, making it suitable for resource-constrained scenarios (such as mobile devices or embedded systems); due to the compression mechanism, DC-AE is more robust to noise and overfitting.

[0080] For common image generation processes, a common approach is to configure a Variational Autoencoder (VAE) to compress the original image directly during the denoising process. However, VAEs can only achieve low compression ratios, and in specific cases, such as... Figure 2 As shown, variational autoencoders can only compress an original image of size (H, W) to a maximum of 1 / 8 of the original image before performing other image processing steps such as feature extraction during denoising. Common image generation methods perform denoising in the pixel space of the original image, resulting in extremely high computational costs.

[0081] In this embodiment, a deep compression autoencoder is pre-trained independently, allowing it to learn the mapping relationship from the original high-dimensional data to low-dimensional latent variables. This makes the compression process of the original image independent of the image feature processing. The deep compression autoencoder pre-trains an efficient latent space, mapping the high-dimensional image (i.e., the original image) to a low-dimensional compact representation (i.e., the first image feature), reducing computational load. The denoising module then performs denoising processing on the first image feature in the latent space, instead of directly processing the high-dimensional original image. For example,... Figure 3 As shown, the deep compression autoencoder can compress an original image of size (H, W) to a maximum of 1 / 64 of the original image size. A denoising module then performs feature extraction and other image processing steps. The resolution of the image requiring denoising is significantly reduced, thus greatly reducing the computational load for denoising, while still maintaining the high quality of the second image features, resulting in a significant improvement in image generation speed.

[0082] In one embodiment, the process of decompressing the second image features based on the compression factor can be accomplished using any of the following decompression techniques: Huffman decoding, deep learning-based decompression, PNG decoding, or other image decompression techniques. The chosen decompression technique is sufficient to satisfy the image generation process provided in this application.

[0083] In one embodiment, the image generation model includes a decoder.

[0084] Image reconstruction and decompression based on second image features to obtain the target image output by the image generation model includes: inputting the second image features into the decoder, and having the decoder perform image reconstruction and decompression on the second image features to obtain the target image output by the decoder. The original image and the target image have the same resolution, and the decoder is obtained by synchronous training based on the compression factor and the deep compression autoencoder.

[0085] In this embodiment, to improve the resolution of the second image features, a decoder trained synchronously with a deep compression autoencoder is used to decompress the second image features. The decompression ratio of the decoder is adjusted according to the compression factor, so that the resolution of the decompressed target image is consistent with that of the original image, thereby improving the visibility of the target image.

[0086] In this embodiment, as Figure 4 As shown, the image generation module includes a deep compression autoencoder, a denoising module, and a decoder. The image generation process is as follows:

[0087] Taking an original image with dimensions (H, W, C) as an example, where H represents the image height, W represents the image width, and C represents the number of image channels, the original image is input into a deep compression autoencoder. After compression by the compression layer in the deep compression autoencoder, the first image feature is output, with image dimensions (H / p, W / p, p). 2 C) Establish a latent space including the features of the first image. The principle of image compression is as follows:

[0088]

[0089] Where p is the compression factor.

[0090] The denoising module processes the first image features in the latent space to generate a second image feature. The image size of the second image feature is (H / p, W / p, p). 2 C).

[0091] The decoder decompresses the second image features, and the principle is as follows:

[0092]

[0093] The decoder outputs the target image, whose image size is (H, W, C).

[0094] In one embodiment, the training process of the image generation model is as follows: the original encoder and the original decoder are trained synchronously using the first sample data. When the first training metric meets the first requirement, the internal parameters of the original encoder and the original decoder are fixed to obtain a deep compression autoencoder and a decoder. Based on the compression and feature processing of the second sample data based on the deep compression autoencoder, and the image reconstruction and decompression processing of the compressed sample results output by the original generation model based on the decoder, the original denoising module is trained using the second sample data. When the second training metric meets the second requirement, the denoising module is obtained.

[0095] In this embodiment, the image generation model needs to be further trained after initial structural adjustments and parameter settings. During training, the original encoder and decoder need to be trained synchronously, while the denoising module is trained asynchronously based on the deep compression autoencoder and decoder. Optionally, the training process of the deep compression autoencoder and decoder is an end-to-end learning process, which can be trained using loss functions such as mean squared error, perceptual loss, and / or adversarial loss. The internal parameters of the original encoder are adjusted, and the loss function is optimized so that the deep compression autoencoder can efficiently obtain the first image features while preserving the quality and detail information of the original image. Simultaneously, the decoder is trained, and the internal parameters of the original decoder are adjusted so that the decoder can efficiently decompress the image.

[0096] After the deep compression autoencoder and decoder are trained synchronously, the original denoising module, with its structure adjusted and initial parameters set, is trained based on the deep compression autoencoder and decoder with fixed internal parameters. Specifically, to improve the adaptability of the denoising module with the deep compression autoencoder and decoder, the image compression and decompression process is implemented using the pre-trained deep compression autoencoder and decoder. The original denoising module is trained by compressing the second sample data based on the deep compression autoencoder and decompressing the compressed sample results output by the original generation model based on the decoder. This allows the original denoising module to learn denoising capabilities based on the first image features in the latent space, thus completing the training of the denoising module and ultimately obtaining the image generation model.

[0097] It should be noted that the first and second sample data can be from the same batch of samples. For example, in a virtual digital human generation scenario, the first and second sample data could be image data from the same batch of videos. Alternatively, the first and second sample data can be different samples, thereby increasing the number of samples and further improving the training effect.

[0098] In this embodiment, after the image generation model is trained for the first time, in order to further improve the image processing effect of the image generation model, a third batch of sample data can be used to optimize the training of the deep compression autoencoder, decoder and denoising module, and the internal parameters of the deep compression autoencoder, decoder and / or denoising module can be fine-tuned again. For example, in the virtual digital human generation scenario, a small number of ultra-high-definition videos can be used for optimization training, thereby improving the quality and effect of image generation.

[0099] In one embodiment, this method is combined with virtual digital human technology in scenarios such as real-time interaction, animation production, and virtual anchors. This method can reduce the time to generate each frame of the target image to sub-second levels, achieving high real-time performance and efficiency, meeting the needs of long-duration, high-frame-rate video generation. The compression factor can be dynamically adjusted to ensure that the generated virtual digital human images maintain the original high quality level in terms of detail and visual effects, exhibiting strong versatility and broad application prospects. By optimizing the image generation process, hardware resource consumption and operating costs are reduced, improving the cost-effectiveness of image generation.

[0100] In this application, the original image is input into an image generation model. The image generation model compresses and characterizes the original image based on a compression factor to obtain first image features. A denoising module, adjusted based on the compression factor, performs inverse diffusion processing on the first image features to obtain second image features. Image reconstruction and decompression are then performed based on the second image features to obtain the target image output by the image generation model. The image generation model includes a denoising module. In this process, the first image features are not only the characterized image data but also compressed based on a compression factor. When the denoising module processes the compressed first image features, the amount of data required for inverse diffusion processing is reduced, thus improving the processing speed of the denoising module. Furthermore, the denoising module is adjusted based on the compression factor, achieving compatibility between the denoising module and the first image features and avoiding redundant processing, further improving the image generation speed. Furthermore, image reconstruction and decompression processing on the second image features ensures the visibility of the target image. While ensuring image visibility, the image generation speed is further improved, guaranteeing high real-time performance and efficiency in various operational scenarios.

[0101] Exemplary device

[0102] Accordingly, embodiments of this application also provide an image generation apparatus, such as... Figure 5 As shown, the device may include:

[0103] Input module 501 is used as an input module to input the original image into the image generation model;

[0104] The processing module 502 is used to compress and characterize the original image based on the compression factor through the image generation model to obtain the first image feature, and to perform inverse diffusion processing on the first image feature by the denoising module adjusted based on the compression factor to obtain the second image feature, and to perform image reconstruction and decompression processing based on the second image feature to obtain the target image output by the image generation model. The image generation model includes the denoising module.

[0105] In one embodiment, the denoising module includes a feature processing layer, a bottleneck layer, and an output layer. The feature processing layer includes at least one convolutional layer. The number of channels in the first convolutional layer is adapted to the feature dimension of the first image feature. The first convolutional layer is used to receive the first image feature.

[0106] The processing module 502 is used to input the first image features into the feature processing layer, where the feature processing layer performs feature extraction and feature dimension adjustment on the first image features and outputs the first intermediate features; input the first intermediate features into the bottleneck layer, where the bottleneck layer performs global feature modeling on the first intermediate features and outputs the second intermediate features; input the second intermediate features into the output layer, where the output layer adjusts the feature dimension of the second intermediate features and outputs the second image features, wherein the feature dimensions of the first image features and the second image features are the same.

[0107] In one embodiment, the bottleneck layer includes a hierarchical visual transformer module.

[0108] In one embodiment, the image generation model includes a deep compression autoencoder;

[0109] The processing module 502 is used to input the original image into the deep compression autoencoder, and the deep compression autoencoder performs compression and feature processing on the original image to obtain the first image feature output by the deep compression autoencoder. The deep compression autoencoder is an encoder trained based on the compression factor.

[0110] In one embodiment, the image generation model includes a decoder;

[0111] The processing module 502 is used to input the second image features into the decoder, and the decoder performs image reconstruction and decompression processing on the second image features to obtain the target image output by the decoder. The original image and the target image have the same resolution. The decoder is obtained by synchronous training based on the compression factor and the deep compression autoencoder.

[0112] In one embodiment, the image generation apparatus further includes a training module for training the image generation model. The training process is as follows: the original encoder and the original decoder are trained synchronously using first sample data. When the first training metric meets the first requirement, the internal parameters of the original encoder and the original decoder are fixed to obtain a deep compression autoencoder and a decoder. Based on the compression and feature processing of the second sample data based on the deep compression autoencoder, and the image reconstruction and decompression processing of the compressed sample results output by the original generation model based on the decoder, the original denoising module is trained using the second sample data. When the second training metric meets the second requirement, the denoising module is obtained.

[0113] In one embodiment, the denoising module includes a U-shaped convolutional neural network.

[0114] The image generation apparatus provided in this embodiment belongs to the same concept as the image generation method provided in the above embodiments of this application. It can execute the image generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the specific processing content of the image generation method provided in the above embodiments of this application, and will not be repeated here.

[0115] Exemplary electronic devices

[0116] This application also provides an electronic device, such as... Figure 6 As shown, the electronic device includes a memory 600 and a processor 601.

[0117] The memory 600 is connected to the processor 601 and is used to store programs.

[0118] The processor 601 is used to implement the image generation method in the above embodiments by running the program stored in the memory 600.

[0119] Specifically, the aforementioned electronic device may also include: a communication interface 602, an input device 603, an output device 604, and a bus 605.

[0120] The processor 601, memory 600, communication interface 602, input device 603, and output device 604 are interconnected via a bus. Among them:

[0121] Bus 605 may include a pathway for transmitting information between various components of a computer system.

[0122] The processor 601 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0123] Processor 601 may include a main processor, as well as a baseband chip, modem, etc.

[0124] The memory 600 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 600 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0125] Input device 603 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0126] Output device 604 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0127] The communication interface 602 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0128] The processor 601 executes the program stored in the memory 600 and calls other devices, which can be used to implement the various steps of the image generation method provided in the above embodiments of this application.

[0129] Exemplary computer program products and storage media

[0130] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the image generation method described in the embodiments of this application.

[0131] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0132] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor of the steps in the image generation method described in the embodiments of this application.

[0133] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0134] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0135] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0136] The modules and sub-modules in the devices and terminals provided in the various embodiments of this application can be merged, divided, and deleted according to actual needs.

[0137] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0138] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0139] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0140] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0141] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0142] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0143] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image generation method, characterized in that, include: Input the original image into the image generation model; The image generation model compresses and characterizes the original image based on a compression factor to obtain a first image feature. A denoising module, adjusted based on the compression factor, performs inverse diffusion processing on the first image feature to obtain a second image feature. Based on the second image feature, image reconstruction and decompression are performed to obtain the target image output by the image generation model. The image generation model includes the denoising module.

2. The image generation method according to claim 1, characterized in that, The denoising module includes a feature processing layer, a bottleneck layer, and an output layer. The feature processing layer includes at least one convolutional layer. The number of channels in the first convolutional layer is adapted to the feature dimension of the first image feature. The first convolutional layer is used to receive the first image feature. The denoising module, adjusted based on the compression factor, performs inverse diffusion processing on the first image features to obtain the second image features, including: The first image features are input into the feature processing layer, which extracts features and adjusts feature dimensions of the first image features, and outputs the first intermediate features. The first intermediate feature is input into the bottleneck layer, which performs global feature modeling on the first intermediate feature and outputs the second intermediate feature. The second intermediate feature is input into the output layer, the output layer adjusts the feature dimension of the second intermediate feature, and outputs the second image feature, wherein the feature dimension of the first image feature and the second image feature are the same.

3. The image generation method according to claim 2, characterized in that, The bottleneck layer includes a hierarchical visual transformer module.

4. The image generation method according to claim 1, characterized in that, The image generation model includes a deep compression autoencoder; The step of compressing and characterizing the original image based on a compression factor to obtain the first image feature includes: The original image is input into the deep compression autoencoder, which performs compression and feature processing on the original image to obtain the first image feature output by the deep compression autoencoder. The deep compression autoencoder is an encoder trained based on the compression factor.

5. The image generation method according to claim 4, characterized in that, The image generation model includes a decoder; The step of performing image reconstruction and decompression based on the second image features to obtain the target image output by the image generation model includes: The second image feature is input into the decoder, which performs image reconstruction and decompression on the second image feature to obtain the target image output by the decoder. The original image and the target image have the same resolution. The decoder is obtained by synchronous training based on the compression factor and the deep compression autoencoder.

6. The image generation method according to claim 5, characterized in that, The training process of the image generation model is as follows: The original encoder and the original decoder are trained synchronously using the first sample data. When the first training metric meets the first requirement, the internal parameters of the original encoder and the original decoder are fixed to obtain the deep compression autoencoder and the decoder. Based on the compression and feature processing of the second sample data using the deep compression autoencoder, and the image reconstruction and decompression processing of the compressed sample results output by the original generative model using the decoder, the original denoising module is trained using the second sample data. When the second training metric meets the second requirement, the denoising module is obtained.

7. The image generation method according to any one of claims 1-6, characterized in that, The denoising module includes a U-shaped convolutional neural network.

8. An image generation apparatus, characterized in that, include: The input module is used to input the original image into the image generation model; The output module is used to compress and characterize the original image based on a compression factor to obtain a first image feature through the image generation model, and to perform inverse diffusion processing on the first image feature by a denoising module adjusted based on the compression factor to obtain a second image feature. Based on the second image feature, image reconstruction and decompression processing are performed to obtain the target image output by the image generation model, wherein the image generation model includes the denoising module.

9. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the image generation method as described in any one of claims 1-7 by running a program in the memory.

10. A computer program product, characterized in that, Includes computer program instructions; When the computer program instructions are executed by the processor, the processor causes the processor to perform the image generation method as described in any one of claims 1-7.