Model training method, image generation method, device, equipment and program product

By using image category vectors and feature vectors in image repair for image generation and training the model, the problems of poor image repair results and large computing overhead in the prior art are solved, and higher quality image repair and more efficient computing processes are achieved.

CN119919948APending Publication Date: 2025-05-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998298.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art has oversmoothing effects and image artifacts in image repair, resulting in poor repair results and huge computing overhead.

Method used

By obtaining the pre-constructed initial model and sample image pair, the image category vector and image feature vector of the degenerated sample image are determined, the image generation process is used for image generation, and the model is trained according to the loss value between the generated image and the target image to obtain an image generation model.

Benefits of technology

It reduces the illusion of text generation in the text-generating graphics model, reduces the significant content differences in repair images, improves the fidelity of generated images, and saves the calculation overhead in the image generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919948A_ABST
    Figure CN119919948A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, an image generation method and device, electronic equipment and a computer program product, and relates to the technical field of image processing. The method comprises the following steps: acquiring a pre-constructed initial model and a sample image pair, wherein the sample image pair comprises a degraded sample image and a target sample image; determining an image category vector and an image feature vector corresponding to the degraded sample image; performing image generation processing based on the image category vector and the image feature vector by the initial model to obtain a prediction generated image; and training the initial model according to a loss value between the predicted generated image and the target sample image to obtain an image generation model. According to the invention, the text generation illusion phenomenon of the text generation graph model can be reduced, the significant content difference occurring in the image restoration can be reduced, the fidelity of the generated image can be improved, and the calculation overhead in the image generation process can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a model training method, an image generating method, a model training device, an image generating device, an electronic device and a computer program product. Background Art

[0002] Image restoration is a technique that aims to restore or improve image quality and is widely used in image super-resolution, denoising, and deblurring. Traditional image restoration methods usually rely on image processing techniques and statistical models, such as interpolation, Fourier transform, and total variation regularization. These methods can achieve certain results when dealing with specific types of image damage, but they usually have certain limitations when dealing with complex damaged or low-quality images.

[0003] In recent years, deep learning techniques, especially generative adversarial networks and convolutional neural networks, have made significant progress in the field of image restoration. These methods can generate high-quality restoration results by synthesizing large-scale low-quality image-high-quality image (LQ-HQ) image data pairs for model training. However, due to the limitations of data synthesis methods and optimization goals, conventional deep learning methods may have over-smoothing effects and generate image artifacts when processing real-world images, resulting in poor restoration results. Summary of the invention

[0004] The present disclosure provides a model training method, an image generation method, a model training device, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product, to at least solve the problem that the image generation quality of the Chinese raw image model in the related art strongly depends on the completeness and accuracy of the text description, resulting in significant content differences in the image generation results, low fidelity, and huge computational overhead. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of the present disclosure, a model training method is provided, comprising: obtaining a pre-constructed initial model and a sample image pair, wherein the sample image pair comprises a degraded sample image and a target sample image; determining an image category vector and an image feature vector corresponding to the degraded sample image; performing image generation processing based on the image category vector and the image feature vector by the initial model to obtain a predicted generated image; and training the initial model according to a loss value between the predicted generated image and the target sample image to obtain an image generation model.

[0006] In an exemplary embodiment of the present disclosure, determining the image category vector and the image feature vector corresponding to the degraded sample image includes: acquiring a preconfigured image preprocessing model; performing image preprocessing on the degraded sample image based on the image preprocessing model to obtain a restored degraded sample image; performing encoding processing on the restored degraded sample image to obtain an encoded degraded image vector; determining the image category vector corresponding to the degraded sample image based on the encoded degraded image vector; and determining the image feature vector corresponding to the degraded sample image based on the encoded degraded image vector.

[0007] In an exemplary embodiment of the present disclosure, encoding the repaired degraded sample image to obtain the encoded degraded image vector includes: obtaining a pre-trained image encoder and determining a latent space dimension corresponding to the image encoder; and encoding the repaired degraded sample image based on the latent space dimension by the image encoder to obtain the encoded degraded image vector.

[0008] In an exemplary embodiment of the present disclosure, determining the image category vector corresponding to the degraded sample image based on the coded degraded image vector includes: acquiring a pre-trained category recognition model and an image category database, the image category database including a plurality of candidate image categories; determining an initial image category vector based on the coded degraded image vector by the category recognition model; and comparing the initial image category vector with candidate image category vectors corresponding to each of the plurality of candidate image categories to obtain the image category vector, the image category vector being used to identify an image category label corresponding to the degraded sample image.

[0009] In an exemplary embodiment of the present disclosure, determining the image feature vector corresponding to the degraded sample image based on the coded degraded image vector includes: obtaining a pre-trained image feature extraction model; and performing image feature extraction processing on the coded degraded image vector based on the image feature extraction model to obtain the image feature vector.

[0010] In an exemplary embodiment of the present disclosure, the initial model includes an image generation network and a control branch network, and the initial model performs image generation processing based on the image category vector and the image feature vector to obtain a predicted generated image, including: performing noise processing on the target sample image to obtain a noisy target sample image; generating an image generation reference condition based on the image category vector and the image feature vector by the first image generation module in the control branch network; using the noisy target sample image as an input of the image generation network, and the image generation network includes multiple second image generation modules; and performing image generation processing according to the noisy target sample image and the image generation reference condition by the multiple second image generation modules to obtain the predicted generated image.

[0011] In an exemplary embodiment of the present disclosure, the initial model is trained according to the loss value between the predicted generated image and the target sample image to obtain the image generation model, including: determining the image loss value between the predicted generated image and the target sample image; based on the image loss value, adjusting the model parameters of the first image generation module until the model training end condition is reached to obtain the image generation model.

[0012] According to a second aspect of the present disclosure, there is provided an image generation method, comprising: obtaining an image to be processed; obtaining a pre-trained image generation model, wherein the image generation model is trained based on a model training method; and performing image generation processing on the image to be processed by the image generation model to obtain a target generated image.

[0013] According to a third aspect of the present disclosure, a model training device is provided, comprising: a sample image acquisition module, used to acquire a pre-constructed initial model and a sample image pair, wherein the sample image pair comprises a degraded sample image and a target sample image; a vector determination module, used to determine an image category vector and an image feature vector corresponding to the degraded sample image; a predicted image generation module, used to perform image generation processing based on the image category vector and the image feature vector by the initial model to obtain a predicted generated image; and a model training module, used to train the initial model according to a loss value between the predicted generated image and the target sample image to obtain an image generation model.

[0014] In an exemplary embodiment of the present disclosure, the vector determination module includes a vector determination unit, which is used to: obtain a preconfigured image preprocessing model; based on the image preprocessing model, perform image preprocessing on the degraded sample image to obtain a repaired degraded sample image; perform encoding processing on the repaired degraded sample image to obtain an encoded degraded image vector; based on the encoded degraded image vector, determine an image category vector corresponding to the degraded sample image; based on the encoded degraded image vector, determine an image feature vector corresponding to the degraded sample image.

[0015] In an exemplary embodiment of the present disclosure, the vector determination unit includes an image encoding subunit, which is used to: obtain a pre-trained image encoder to determine the latent space dimension corresponding to the image encoder; and the image encoder encodes the repaired degraded sample image based on the latent space dimension to obtain the encoded degraded image vector.

[0016] In an exemplary embodiment of the present disclosure, the vector determination unit includes a category vector determination subunit, which is used to: obtain a pre-trained category recognition model and an image category database, the image category database including multiple candidate image categories; determine an initial image category vector based on the encoded degraded image vector by the category recognition model; compare the initial image category vector with candidate image category vectors corresponding to each of the multiple candidate image categories to obtain the image category vector, and the image category vector is used to identify the image category label corresponding to the degraded sample image.

[0017] In an exemplary embodiment of the present disclosure, the vector determination unit includes a feature vector determination subunit, which is used to: obtain a pre-trained image feature extraction model; and perform image feature extraction processing on the coded degraded image vector based on the image feature extraction model to obtain the image feature vector.

[0018] In an exemplary embodiment of the present disclosure, the initial model includes an image generation network and a control branch network, and the predicted image generation module includes a predicted image generation unit, which is used to: perform noise processing on the target sample image to obtain a noisy target sample image; generate an image generation reference condition based on the image category vector and the image feature vector by the first image generation module in the control branch network; use the noisy target sample image as input to the image generation network, and the image generation network includes multiple second image generation modules; perform image generation processing according to the noisy target sample image and the image generation reference condition by the multiple second image generation modules to obtain the predicted generated image.

[0019] In an exemplary embodiment of the present disclosure, the model training module includes a model training unit, which is used to: determine the image loss value between the predicted generated image and the target sample image; based on the image loss value, adjust the model parameters of the first image generation module until the model training end condition is reached, thereby obtaining the image generation model.

[0020] According to a fourth aspect of the present disclosure, an image generating device is provided, comprising: an image acquiring module for acquiring an image to be processed; a model acquiring module for acquiring a pre-trained image generating model, wherein the image generating model is trained based on a model training method; and an image generating module for performing image generating processing on the image to be processed by the image generating model to obtain a target generated image.

[0021] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to execute instructions to implement any one of the above-mentioned model training methods, or to implement the above-mentioned image generation method.

[0022] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the above-mentioned model training methods, or implement the above-mentioned image generation method.

[0023] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements any one of the above-described model training methods or the above-described image generation method.

[0024] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0025] On the one hand, using the image category vector as reference information in the image generation process can alleviate the text generation hallucination phenomenon of the text-based graph model, reduce the significant content differences in the repaired image, and improve the fidelity of the generated image. On the other hand, using the directly extracted image feature vector in the image generation process can save the computational overhead in the image generation process and improve processing efficiency.

[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0028] Figure 1 This is the model architecture diagram for image restoration based on MLLM and CLIP in the related XPSR scheme to generate high-quality images.

[0029] Figure 2 The figure is a flowchart of a model training method according to an exemplary embodiment.

[0030] Figure 3 It is an overall structural diagram of a model constructed for an image generation processing task according to an exemplary embodiment.

[0031] Figure 4 is a structural diagram of a DiT category image generation model according to an exemplary embodiment.

[0032] Figure 5 is a schematic diagram of an image feature vector extraction mechanism for a low-quality image according to an exemplary embodiment.

[0033] Figure 6 The figure is a flowchart of an image generating method according to an exemplary embodiment.

[0034] Figure 7 It is a schematic diagram showing the result of image restoration based on the image generation model of the present disclosure according to an exemplary embodiment.

[0035] Figure 8 It is a block diagram of a model training device according to an exemplary embodiment.

[0036] Fig. 9 The figure is a block diagram of an image generating device according to an exemplary embodiment.

[0037] Fig.10 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0038] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0040] Image generation technology based on Denoising Diffusion Probabilistic Models (DDPM) demonstrates excellent detail generation and reconstruction capabilities. Text-based image models represented by Stable Diffusion v1.5 (SDv1.5) and Stable Diffusion XL (SDXL) can convert a piece of text input by the user into a realistic and detailed image. This inspired the industry to introduce this generative model technology to achieve more natural and detailed image restoration.

[0041] In some related studies, a technical solution combining a control branch network (ControlFormer) and a text to image base model is usually used, that is, the text to image model (Text to Image) is used as the base model, and the ControlFormer network is added to introduce a low-quality image (LQ Image) as a conditional control base model for image generation, thereby achieving high-fidelity high-quality image restoration. Compared with conventional deep learning technology, this new method can give full play to the prior knowledge of the base model and significantly improve the detail generation ability and subjective experience of image restoration. However, the existing methods still have major limitations:

[0042] (1) Model inference costs are high and computational efficiency is low. Since the base model is a text-image model, it is usually necessary to use a multi-modal large language model (MLLM), such as Bootstrapping Language-Image Pre-training (BLIP), to extract the text description of LQ, and then send the text description to a text feature extraction model, such as Contrastive Language-Image Pre-Training (CLIP), to convert it into a text embedding vector, and finally send it to the base model. The computational overhead of the multi-modal large model is very huge, which affects the inference of the image restoration model.

[0043] (2) Affected by the text description of the Vincent-Tukey model, the fidelity of the restored image is poor. LQ usually has obvious image quality damage, which may trigger the "hallucination" phenomenon of text generation in the multimodal annotation model, such as: incorrect quantitative relationship (LQ has two people, and the generated text description is three people), making something out of nothing (LQ's selfie has no jewelry, and the generated text description has earrings and necklaces), and calling a deer a horse (LQ's animal is a dog, and the generated text description is a cat). These descriptions will affect the output of the entire restoration model, resulting in significant content differences between the restored HQ and LQ, and low fidelity.

[0044] Taking the image restoration enhancement (Cross-modal Priors for Super Resolution, XPSR) model based on diffusion model and cross-modal prior information as an example, refer to Figure 1 , Figure 1 This is a model architecture diagram for image restoration based on MLLM and CLIP in the related XPSR scheme to generate high-quality images. This method requires MLLM and CLIP to complete the conversion of "LQ->text description->text embedding vector" (such as Figure 1 The obtained embedding vector is combined with the feature vector of LQ extracted by the ControlNet network (as shown in the sub-figure above). Figure 1 The image is restored by the base model (as shown in the sub-figure below), and then returned to HQ.

[0045] This scheme has the following disadvantages: (1) The conversion process of "LQ->text description->text embedding vector" requires the intervention of MLLM, which has a huge computational overhead; (2) Affected by the image damage of LQ, the text description output by MLLM is usually accompanied by "hallucinations", which affects the fidelity of the restored HQ image.

[0046] Based on this, according to the embodiments of the present disclosure, a model training method, an image generation method, a model training device, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product are proposed.

[0047] Figure 2 is a flow chart of a model training method according to an exemplary embodiment. Figure 2As shown, the model training method can be used in a computer device, wherein the computer device described in the present disclosure may include mobile terminal devices such as mobile phones, tablet computers, laptops, PDAs, and fixed terminal devices such as desktop computers. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.

[0048] Step S210, obtaining a pre-built initial model and a sample image pair, where the sample image pair includes a degraded sample image and a target sample image;

[0049] Step S220, determining an image category vector and an image feature vector corresponding to the degraded sample image;

[0050] Step S230, performing image generation processing based on the image category vector and the image feature vector by the initial model to obtain a predicted generated image;

[0051] Step S240, training the initial model according to the loss value between the predicted generated image and the target sample image to obtain an image generation model.

[0052] According to the model training method in this example embodiment, on the one hand, using the image category vector as reference information in the image generation process can reduce the text generation hallucination phenomenon of the text-generated graph model, reduce the significant content differences in the repaired image, and improve the fidelity of the generated image. On the other hand, using the directly extracted image feature vector in the image generation process can save the computational overhead in the image generation process and improve processing efficiency.

[0053] The model training method in this example embodiment will be further described below.

[0054] In step S210, a pre-built initial model and a sample image pair are obtained, where the sample image pair includes a degraded sample image and a target sample image.

[0055] In an exemplary embodiment of the present disclosure, the initial model may be an image restoration network model based on a base model generated by a category image. The sample image pair may be an LQ-HQ image data pair used to train the initial model. The degraded sample image may be image data after image degradation, and the degraded sample image may be an image whose image quality index value is less than a predefined quality parameter threshold. The target sample image may be image data with good picture quality, such as an image whose image quality index value is greater than or equal to a predefined quality parameter threshold. Image quality indicators may include, but are not limited to, image resolution, color depth, signal to noise ratio (SNR), contrast, peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), etc.

[0056] Image degradation refers to the phenomenon that image quality decreases during the formation, recording, processing and transmission of images due to imperfections in the imaging system, recording equipment, transmission media and processing methods. In actual application scenarios, due to factors such as limitations of shooting equipment, compression encoding process, multiple uploads and downloads, the image quality is reduced. This part of the image can be regarded as a low-quality image, namely LQ sample data. The target sample image can be a high-quality image with a higher resolution, namely HQ sample data.

[0057] In order to overcome the problem that the content and quality of the generated image are strongly dependent on the completeness and accuracy of the text description in the image generation scheme based on the text base model, the present disclosure improves it into a base model based on category generation, and the pre-built initial model can be an image generation model implemented based on the base model based on category generation.

[0058] In step S220, an image category vector and an image feature vector corresponding to the degraded sample image are determined.

[0059] In an exemplary embodiment of the present disclosure, the image category vector may be a category vector corresponding to the image content contained in the degraded sample image. The image feature vector may be a feature vector obtained by performing feature extraction processing on the image content contained in the degraded sample image.

[0060] For the degraded sample image in the sample image pair, the content category contained in the degraded sample image can be identified and processed to determine the corresponding image category vector; in addition, the image detail features in the sample degraded image can be extracted and processed to obtain the corresponding image feature vector, so that the obtained vector can be used as the input of the control branch network for the image generation process.

[0061] In an exemplary embodiment of the present disclosure, for step S220, determining the image category vector and the image feature vector corresponding to the degraded sample image includes: obtaining a preconfigured image preprocessing model; performing image preprocessing on the degraded sample image based on the image preprocessing model to obtain a repaired degraded sample image; performing encoding processing on the repaired degraded sample image to obtain an encoded degraded image vector; determining the image category vector corresponding to the degraded sample image based on the encoded degraded image vector; and determining the image feature vector corresponding to the degraded sample image based on the encoded degraded image vector.

[0062] The image preprocessing model may be a model for performing image preprocessing on a sample degraded image. Image preprocessing may include a preliminary image restoration process. Image preprocessing may be a process for improving the image quality and clarity of a sample degraded image, or may be a process for improving the image stability of a sample degraded image. The restored degraded sample image may be a preliminary restored sample image obtained after performing image preprocessing on the sample degraded image. The coded degraded image vector may be a vector obtained after performing image coding processing on the coded degraded image vector.

[0063] refer to Figure 3 , Figure 3 FIG. 1 is an overall structural diagram of a model constructed for an image generation processing task according to an exemplary embodiment. Figure 3 In the method, before the sample degraded image is used as the input of the control branch network in the initial model, the sample degraded image can be preprocessed by a pre-configured image preprocessing model (Low-quality Image Enhanced, LqEnh). For example, the convolutional neural network model based on residual dense blocks (Residual-in-Residual Dense Block, RRDBNet) network structure in the LqEnh model is used to perform preliminary restoration processing on the sample degraded image to obtain a restored degraded sample image (Enh Image).

[0064] After obtaining the repaired degraded sample image, the repaired degraded sample image can be coded to obtain a coded degraded image vector corresponding to the repaired degraded sample image. The coded degraded image vector is then used as the input of the control branch network in the initial model, and the control branch network sends the coded degraded image vector as a constraint condition to the DiT block. The DiT block can determine the image feature vector corresponding to the degraded sample image based on the received input data.

[0065] In addition, the coded degraded image vector can be processed for image category recognition to determine the image category vector corresponding to the degraded sample image; the repaired degraded sample image can also be directly processed for image category recognition to obtain the corresponding image category vector. The image category vector and image feature vector determined by the above steps can be used as inputs of the control branch network to accurately guide the image generation results. Since the real-world LQ images may have suffered complex damage, by adding the LqEnh module, the LQ images can be initially repaired, thereby providing a more stable and clear image for the subsequent image restoration based on the category image generation base, which can improve the stability and clarity of the overall image restoration.

[0066] In an exemplary embodiment of the present disclosure, encoding processing is performed on a restored degraded sample image to obtain an encoded degraded image vector, including: obtaining a pre-trained image encoder and determining a latent space dimension corresponding to the image encoder; and the image encoder performs encoding processing on the restored degraded sample image based on the latent space dimension to obtain an encoded degraded image vector.

[0067] The image encoder may be an encoder model for encoding an image to obtain a latent vector corresponding to the image. For example, the image encoder may be a variational auto encoder (VAE). The latent space dimension may be the dimension of the latent vector output by the image encoder after encoding the input image.

[0068] For related text-based graph models, such as SDv1.5, SDXL, DiT, etc., the VAE-F8C4 model is used. The VAE-F8C4 encoder can represent a VAE encoder that downsamples the image space domain by 8 times and has a latent space dimension of 4 during the image encoding process. For example, the VAE-F8C4 encoder can compress the input image (image size is 3xHxW) into a latent vector (H / 8xW / 8x4). This large compression ratio will cause the original information of the input image to be lost and distorted, and is not suitable for image restoration tasks that focus more on image details and high-frequency textures.

[0069] To solve the above problems, continue to refer to Figure 3 , Figure 3 The network structure of the initial network includes a pre-trained image encoder. For example, the present disclosure can select a variational autoencoder VAE-F8C16 model as the image encoder. The present disclosure uses an open source image (Open Images) dataset and an internal massive (e.g., 20 million) high-quality dataset to retrain the VAE-F8C16 model, and uses the retrained VAE-F8C16 model as the image encoder of the present disclosure.

[0070] The VAE-F8C16 model can indicate that during the image encoding process, the image spatial domain is downsampled by 8 times, and the latent space dimension is increased to 16, that is, the number of image channels after encoding is 16. Compared with the VAE-F8C4 encoder, the texture distortion caused by the large compression ratio of VAE image encoding is significantly reduced, and the PSNR is improved by 4-6 decibels (dB). Using an image encoder with a smaller compression ratio can effectively retain the details of the image, which is more suitable for processing image restoration tasks.

[0071] In an exemplary embodiment of the present disclosure, an image category vector corresponding to a degraded sample image is determined based on a coded degraded image vector, including: obtaining a pre-trained category recognition model and an image category database, the image category database including a plurality of candidate image categories; determining an initial image category vector based on the coded degraded image vector by the category recognition model; and comparing the initial image category vector with candidate image category vectors corresponding to each of the plurality of candidate image categories to obtain an image category vector, the image category vector being used to identify an image category label corresponding to the degraded sample image.

[0072] The category recognition model may be a network model for identifying specific categories of image content in degraded sample images. The image category database may be a database for storing candidate image categories. The candidate image category may be a specific category to which the image content may correspond. The initial image category vector may be a category vector output by the category recognition model after performing recognition processing on the coded degraded image vector. The candidate image category vector may be an image category vector pre-stored in the image category database, that is, an image category vector to which the image content may correspond. The image category label may be a specific category label corresponding to the degraded sample image.

[0073] In view of the possible hallucination phenomenon in the Chinese image generation model of the related scheme, the DiT block of the present disclosure can use the image category as the generation condition, and integrate the category features into the DiT block through the adaptive layer normalization block (Adaptive Layer Normalization, AdaLN) control mechanism to guide the image generation process. Figure 4 , Figure 4 is a structural diagram of a DiT category image generation model according to an exemplary embodiment. The image category vector corresponding to the degraded sample image can be determined by the category recognition model.

[0074] For an input image, such as a degraded sample image x after restoration, the content category can first be encoded through the CLIP Image Encoder model to obtain an initial image category vector (embedding) of 1x768 dimensions; in addition, the encoded degraded image vector can be directly used as a CLIP Image Encoder model to obtain the corresponding initial image category vector.

[0075] Then, the nearest neighbor search is performed in the image category database. For example, the image category database can contain nearly 10,000 different candidate image categories. The features of the category database are K-Means clustered by CLIP features of 50 million public and internal high-quality data sets. The cluster centers are set to 10,000, and the dimension of each category vector is 768, so the category database features are 10,000x768.

[0076] After the nearest neighbor search process, the target category vector closest to the degraded sample image among multiple candidate image category vectors can be determined as the image category vector corresponding to the degraded sample image. The image category vector can identify the image category label corresponding to the degraded sample image, that is, the category label (Class Label) corresponding to the degraded sample image. Then, the identified image category vector is added to the Transformer-based diffusion model (Diffusion Image Transformer, DiT) network through the AdaLN module. The present disclosure changes the text image base model to a category image generation network, which can reduce the hallucination generation of the text image base model and improve the image generation effect.

[0077] In an exemplary embodiment of the present disclosure, based on the coded degraded image vector, an image feature vector corresponding to the degraded sample image is determined, including: obtaining a pre-trained image feature extraction model; based on the image feature extraction model, performing image feature extraction processing on the coded degraded image vector to obtain the image feature vector.

[0078] Among them, the image feature extraction model can be a network model used to extract image detail features.

[0079] In the related text graph model, the vector representation corresponding to the input image is determined by adopting the processing steps of "LQ->text description->text embedding vector". However, the above process usually requires the intervention of MLLM, and the computational overhead is very huge. In order to solve the above problem, the present disclosure adopts a pre-trained image feature extraction model (such as CLIP Image Encoder model) to directly extract the image embedding vector of the LQ image.

[0080] refer to Figure 5 , Figure 5is a schematic diagram of an image feature vector extraction mechanism for a low-quality image according to an exemplary embodiment. Figure 5 The specific process of image feature extraction using the CLIP Image Encoder model is shown: "LQ->Image Embedding Vector", that is, the model is directly used to extract image detail features, eliminating the MLLM step of converting images into text descriptions, greatly reducing the model's inference overhead, and retaining image detail features.

[0081] In step S230, the initial model performs image generation processing based on the image category vector and the image feature vector to obtain a predicted generated image.

[0082] In an exemplary embodiment of the present disclosure, the predicted generated image may be an image obtained after the initial model performs image generation processing based on the image category vector and the image feature vector.

[0083] Continue to refer Figure 3 After obtaining the image category vector and the image feature vector, the image category vector and the image feature vector can be used as the image generation reference conditions of the DiT model based on the control branch network in the initial model, so that the image generation network of the initial model can accurately perform image generation processing and obtain the predicted generated image.

[0084] In an exemplary embodiment of the present disclosure, for step S230, the initial model performs image generation processing based on the image category vector and the image feature vector to obtain a predicted generated image, including: performing noise processing on the target sample image to obtain a noisy target sample image; generating an image generation reference condition based on the image category vector and the image feature vector by the first image generation module in the control branch network; using the noisy target sample image as input to the image generation network, the image generation network includes multiple second image generation modules; and performing image generation processing by the multiple second image generation modules according to the noisy target sample image and the image generation reference condition to obtain a predicted generated image.

[0085] Among them, the image generation network can be a network structure for performing image generation processing according to the model input and the image generation reference condition. The control branch network (ControlFormer) can be a powerful tool for enhancing the controllability of the image generation process, allowing the user to accurately guide the generation result by providing a specific control image. The noisy target sample image can be an image obtained by adding random noise to the target sample image. The first image generation module can be a DiT block included in the control branch network. The image generation reference condition can be a reference condition based on which the image generation network generates a predicted generated image that matches the model input. The second image generation module can be a DiT block included in the image generation network.

[0086] Continue to refer Figure 3 In the present invention, a target sample image with added noise is used as the input of an image generation network. For a target sample image in a sample image pair, the target sample image is subjected to a noise adding process until input data completely containing Gaussian noise and incapable of identifying any image information is obtained, which is used as a noisy target sample image.

[0087] In addition, the image category vector and the image feature vector determined by the above steps can also be used as inputs of the control branch network, and the control branch network can use the image feature vector as input of the first image generation module, and perform image generation processing on the image feature vector through multiple first image generation modules respectively; at the same time, the control branch network can also perform feature extraction processing on the image category vector by a multilayer perceptron (MLP) as an input of multiple first image generation modules, and be used in the image generation process of the first image generation module.

[0088] The noisy target sample image is used as the input of the image generation network in the initial model. Figure 3 It can be seen that the image generation network includes multiple second image generation modules. The noisy target sample image is subjected to inverse diffusion processing through multiple second image generation modules, and the output results obtained through multiple first image generation modules can act on the second image generation module to guide the image generation process of the second image generation module. The image generation network in the initial model can perform the above image generation processing through multiple second image generation modules according to the noisy target sample image and the image generation reference conditions, and finally obtain the predicted generated image. Through the above process, the model can learn the detailed information in the low-quality image to guide the subsequent high-quality image generation task.

[0089] In step S240, the initial model is trained according to the loss value between the predicted generated image and the target sample image to obtain an image generation model.

[0090] In an exemplary embodiment of the present disclosure, the image generation model may be a network model for processing image generation tasks. The image generation model may be a network model that uses a category image generation network DiT as a base model, introduces an LQ image as a conditional constraint by adding a ControlFormer module, drives DiT to perform image restoration, and outputs an HQ image.

[0091] After obtaining the predicted sample image output by the initial model, the predicted generated image can be compared with the target sample image, and the loss value between the two can be calculated. Then, the initial model can be trained according to the calculated loss value until the model training end condition is reached to obtain the image generation model.

[0092] In an exemplary embodiment of the present disclosure, for step S240, the initial model is trained according to the loss value between the predicted generated image and the target sample image to obtain the image generation model, including: determining the image loss value between the predicted generated image and the target sample image; based on the image loss value, adjusting the model parameters of the first image generation module until the model training end condition is reached to obtain the image generation model.

[0093] The image loss value may be an indicator for measuring the difference between the predicted generated image and the target sample image. The model training end condition may be a judgment condition for indicating whether the model training process can be ended.

[0094] The loss value between the predicted generated image and the target sample image is used as the image loss value. Then, the model parameters of the first image generation module in the initial model are adjusted according to the image loss value, that is, when fine-tuning the model parameters, the pre-trained image enhancement model is kept unchanged, and only the first image generation module in the control branch network is trained. The first image generation module can be trained by a hybrid loss function, for example, the hybrid loss function can include an L1 norm loss function, a learned perceptual image patch similarity (LPIPS) loss function, etc.

[0095] The initial model is trained using the above loss function until the model training end condition is reached. The pre-configured model training end condition can be that the current model is in a convergence state, or that the current training times have reached a pre-defined training round threshold. Through the model training process, the model parameters can be adjusted to a more ideal state, and the model can learn how to generate high-quality images from low-quality images, and use it in subsequent image restoration tasks.

[0096] In summary, the model training method disclosed in the present invention obtains a pre-built initial model and a sample image pair, wherein the sample image pair includes a degraded sample image and a target sample image; determines an image category vector and an image feature vector corresponding to the degraded sample image; performs image generation processing based on the image category vector and the image feature vector by the initial model to obtain a predicted generated image; and trains the initial model according to the loss value between the predicted generated image and the target sample image to obtain an image generation model. On the one hand, using the image category vector as reference information for the image generation process can alleviate the text generation hallucination phenomenon of the text-generated image model, reduce the significant content difference in the repaired image, and improve the fidelity of the generated image. On the other hand, using the directly extracted image feature vector for the image generation process can save the computational overhead in the image generation process and improve the processing efficiency. On the other hand, by performing image preprocessing on the sample degraded image, the stability and clarity of the overall image restoration can be improved. On the other hand, using a variational autoencoder with a smaller compression ratio can retain more image details.

[0097] Figure 6 is a flowchart of an image generating method according to an exemplary embodiment. Figure 6 As shown, the image generation method can be used in a computer device. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.

[0098] Step S610, obtaining an image to be processed;

[0099] Step S620, obtaining a pre-trained image generation model, where the image generation model is trained based on a model training method;

[0100] Step S630: The image generation model performs image generation processing on the image to be processed to obtain a target generated image.

[0101] The image generation solution disclosed in the present invention can significantly improve image clarity, improve user viewing experience, and achieve user experience quality (Quality of Experience, QoE) benefits.

[0102] In an exemplary embodiment of the present disclosure, an image to be processed is obtained, and the image to be processed may be a low-quality initial image to be processed for image restoration. The image to be processed is input into a pre-trained image generation model, and the image generation model uses a category image generation network DiT as a base model, introduces an LQ image as a conditional constraint by adding a ControlFormer module, drives DiT to perform image restoration, and finally outputs a high-quality target generated image. The target generated image may be a high-quality image obtained after image restoration processing by the image generation model.

[0103] refer to Figure 7 , Figure 7 FIG. 1 is a schematic diagram showing the result of image restoration based on the image generation model disclosed in the present invention according to an exemplary embodiment. Figure 7 It can be seen from the figure that the image restoration result based on this scheme generates more detailed textures, and the character's hair is clearly visible, which can greatly improve the image clarity.

[0104] Figure 8 is a block diagram of a model training device according to an exemplary embodiment. Figure 8 The model training device 800 includes: a sample image acquisition module 810, a vector determination module 820, a prediction image generation module 830 and a model training module 840.

[0105] Specifically, the sample image acquisition module 810 is used to obtain a pre-built initial model and a sample image pair, the sample image pair including a degraded sample image and a target sample image; the vector determination module 820 is used to determine the image category vector and the image feature vector corresponding to the degraded sample image; the predicted image generation module 830 is used to perform image generation processing based on the image category vector and the image feature vector by the initial model to obtain a predicted generated image; the model training module 840 is used to train the initial model according to the loss value between the predicted generated image and the target sample image to obtain an image generation model.

[0106] In an exemplary embodiment of the present disclosure, the vector determination module 820 includes a vector determination unit, which is used to: obtain a preconfigured image preprocessing model; based on the image preprocessing model, perform image preprocessing on the degraded sample image to obtain a repaired degraded sample image; perform encoding processing on the repaired degraded sample image to obtain an encoded degraded image vector; based on the encoded degraded image vector, determine an image category vector corresponding to the degraded sample image; based on the encoded degraded image vector, determine an image feature vector corresponding to the degraded sample image.

[0107] In an exemplary embodiment of the present disclosure, the vector determination unit includes an image encoding subunit, which is used to: obtain a pre-trained image encoder and determine the latent space dimension corresponding to the image encoder; and the image encoder encodes the repaired degraded sample image based on the latent space dimension to obtain a coded degraded image vector.

[0108] In an exemplary embodiment of the present disclosure, the vector determination unit includes a category vector determination subunit, which is used to: obtain a pre-trained category recognition model and an image category database, the image category database includes multiple candidate image categories; determine an initial image category vector based on the encoded degraded image vector by the category recognition model; compare the initial image category vector with the candidate image category vectors corresponding to each of the multiple candidate image categories to obtain an image category vector, and the image category vector is used to identify the image category label corresponding to the degraded sample image.

[0109] In an exemplary embodiment of the present disclosure, the vector determination unit includes a feature vector determination subunit, which is used to: obtain a pre-trained image feature extraction model; and perform image feature extraction processing on the coded degraded image vector based on the image feature extraction model to obtain an image feature vector.

[0110] In an exemplary embodiment of the present disclosure, the initial model includes an image generation network and a control branch network, and the predicted image generation module 830 includes a predicted image generation unit, which is used to: perform noise processing on the target sample image to obtain a noisy target sample image; generate an image generation reference condition based on an image category vector and an image feature vector by the first image generation module in the control branch network; use the noisy target sample image as an input of the image generation network, and the image generation network includes multiple second image generation modules; and perform image generation processing according to the noisy target sample image and the image generation reference condition by the multiple second image generation modules to obtain a predicted generated image.

[0111] In an exemplary embodiment of the present disclosure, the model training module 840 includes a model training unit, which is used to: determine the image loss value between the predicted generated image and the target sample image; based on the image loss value, adjust the model parameters of the first image generation module until the model training end condition is reached to obtain the image generation model.

[0112] Fig. 9 is a block diagram of an image generating device according to an exemplary embodiment. Fig. 9 The image generating device 900 includes: an image acquiring module 910 , a model acquiring module 920 and an image generating module 930 .

[0113] Specifically, the image acquisition module 910 is used to acquire the image to be processed; the model acquisition module 920 is used to acquire a pre-trained image generation model, where the image generation model is trained based on a model training method; and the image generation module 930 is used to perform image generation processing on the image to be processed by the image generation model to obtain a target generated image.

[0114] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0115] Refer to the following Fig.10 hereinafter describes the electronic device 1000 according to such an embodiment of the present disclosure. Fig.10 The electronic device 1000 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0116] like Fig.10 As shown, the electronic device 1000 is in the form of a general computing device. The components of the electronic device 1000 may include but are not limited to: the at least one processing unit 1010, the at least one storage unit 1020, a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010), and a display unit 1040.

[0117] The storage unit stores program codes, which can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.

[0118] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1021 and / or a cache memory unit 1022 , and may further include a read-only memory unit (ROM) 1023 .

[0119] The storage unit 1020 may also include a program / utility 1024 having a set (at least one) of program modules 1025, such program modules 1025 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0120] Bus 1030 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0121] The electronic device 1000 may also communicate with one or more external devices 1070 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 1060. As shown, the network adapter 1060 communicates with other modules of the electronic device 1000 via a bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0122] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of the device to complete the above model training method and image generation method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0123] In an exemplary embodiment, a computer program product is also provided, including a computer program, which implements any one of the above-mentioned model training methods and image generation methods when executed by a processor.

[0124] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0125] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A model training method, characterized in that: include: Acquire a pre-built initial model and a sample image pair, wherein the sample image pair includes a degraded sample image and a target sample image; Determine an image category vector and an image feature vector corresponding to the degraded sample image; Performing image generation processing based on the image category vector and the image feature vector by the initial model to obtain a predicted generated image; The initial model is trained according to the loss value between the predicted generated image and the target sample image to obtain an image generation model.

2. The method according to claim 1, characterized in that The step of determining the image category vector and the image feature vector corresponding to the degraded sample image includes: Get pre-configured image preprocessing models; Based on the image preprocessing model, performing image preprocessing on the degraded sample image to obtain a restored degraded sample image; Performing encoding processing on the repaired degraded sample image to obtain an encoded degraded image vector; Based on the coded degraded image vector, determining an image category vector corresponding to the degraded sample image; Based on the encoded degraded image vector, an image feature vector corresponding to the degraded sample image is determined.

3. The method according to claim 2, characterized in that The encoding process is performed on the repaired degraded sample image to obtain an encoded degraded image vector, including: Obtain a pre-trained image encoder and determine a latent space dimension corresponding to the image encoder; The image encoder performs encoding processing on the restored degraded sample image based on the latent space dimension to obtain the encoded degraded image vector.

4. The method according to claim 2, characterized in that: The step of determining the image category vector corresponding to the degraded sample image based on the coded degraded image vector comprises: Obtaining a pre-trained category recognition model and an image category database, wherein the image category database includes a plurality of candidate image categories; Determining an initial image category vector by the category recognition model based on the coded degraded image vector; The initial image category vector is compared with candidate image category vectors corresponding to each of the plurality of candidate image categories to obtain the image category vector, where the image category vector is used to identify an image category label corresponding to the degraded sample image.

5. The method according to claim 2, characterized in that: The step of determining the image feature vector corresponding to the degraded sample image based on the coded degraded image vector comprises: Get a pre-trained image feature extraction model; Based on the image feature extraction model, image feature extraction processing is performed on the coded degraded image vector to obtain the image feature vector.

6. The method according to claim 1, characterized in that The initial model includes an image generation network and a control branch network. The initial model performs image generation processing based on the image category vector and the image feature vector to obtain a predicted generated image, including: Performing noise processing on the target sample image to obtain a noisy target sample image; The first image generation module in the control branch network generates an image generation reference condition based on the image category vector and the image feature vector; Using the noisy target sample image as an input of the image generation network, the image generation network includes a plurality of second image generation modules; The plurality of second image generation modules perform image generation processing according to the noisy target sample image and the image generation reference condition to obtain the predicted generated image.

7. The method according to claim 6, characterized in that The step of training the initial model according to the loss value between the predicted generated image and the target sample image to obtain an image generation model comprises: Determining an image loss value between the predicted generated image and the target sample image; Based on the image loss value, the model parameters of the first image generation module are adjusted until the model training end condition is reached to obtain the image generation model.

8. An image generation method, characterized in that: include: Get the image to be processed; Obtain a pre-trained image generation model, wherein the image generation model is trained based on the model training method according to any one of claims 1 to 7; The image generation model performs image generation processing on the image to be processed to obtain a target generated image.

9. A model training device, characterized in that: include: A sample image acquisition module, used to acquire a pre-built initial model and a sample image pair, wherein the sample image pair includes a degraded sample image and a target sample image; A vector determination module, used to determine an image category vector and an image feature vector corresponding to the degraded sample image; A predicted image generation module, used to perform image generation processing based on the image category vector and the image feature vector using the initial model to obtain a predicted generated image; The model training module is used to train the initial model according to the loss value between the predicted generated image and the target sample image to obtain an image generation model.

10. An image generating device, characterized in that: include: An image acquisition module, used for acquiring an image to be processed; A model acquisition module, used to acquire a pre-trained image generation model, wherein the image generation model is trained based on the model training method according to any one of claims 1 to 7; The image generation module is used to perform image generation processing on the image to be processed by the image generation model to obtain a target generated image.

11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the model training method as described in any one of claims 1 to 7, or to implement the image generation method as described in claim 8.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the model training method as described in any one of claims 1 to 7, or implements the image generation method as described in claim 8.