Model training method, image generation method, device, equipment and program product

By using semantic correlation vectors and feature characterization vectors in image repair technology to train the image generation model, the problem of illusion and insufficient fidelity in image generation in the prior art is solved, and a clearer and more detailed image generation effect is achieved.

CN119942560APending Publication Date: 2025-05-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411998592.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the existing image repair technology based on diffusion model, the generated images have hallucinations, and the fidelity of the repaired images and the detailed generation effect are not clear enough.

Method used

By obtaining the pre-constructed initial model and sample image pair, the semantic correlation vector of the degenerated sample image and the initial input image feature representation vector of the initial model are determined, and the image generation process is performed based on these vectors to obtain the predicted generated image, and the initial model is trained based on the predicted generated image and the target sample image to obtain the image generation model.

Benefits of technology

The semantic consistency between the generated image and the input image is achieved, and the generated image with better clarity and more detailed textures is output, solving the problems of illusion and insufficient fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942560A_ABST
    Figure CN119942560A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, an image generation method and device, electronic equipment and a computer program product, and relates to the technical field of image processing. The method comprises the following steps: acquiring a pre-constructed initial model and a sample image pair, wherein the sample image pair comprises a degraded sample image and a target sample image; determining a semantic association vector corresponding to the degraded sample image; determining an initial input image of the initial model and a feature representation vector of the initial input image; performing image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a prediction generated image; and training the initial model according to the prediction generation image and the target sample image to obtain an image generation model. According to the image generation model, the generated image with better definition and more detail textures can be output, and the generated image and the input image keep semantic consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a model training method, an image generating method, a model training device, an image generating device, an electronic device and a computer program product. Background Art

[0002] In the field of video processing and enhancement, generative technology based on diffusion models has become a key trend in technological development. The diffusion model, with its unique step-by-step generation process, demonstrates excellent detail generation and reconstruction capabilities, which enables it to perform well in multiple tasks such as video restoration, noise reduction, super-resolution, and image quality improvement. The introduction of generative model technology can achieve more natural and detailed visual effects, greatly improve the visual quality of video images, and significantly enhance the user's consumption experience.

[0003] The mainstream image restoration framework based on diffusion model usually adopts the technical solution of combining control network and text image base model, that is, using text image model (Text to Image) as the base model, for example, using stable diffusion model (Stable Diffusion model), and adding control branch network (ControlNet network) on the network base as low-quality image condition control to guide the generated image, perform image restoration and generate high-quality image with high fidelity. More advanced technical solutions include pixel-aware stable diffusion (PASD) algorithm, semantic-aware SR (SeeSR) framework, image restoration enhancement algorithm based on diffusion model and cross-modal prior information (Cross-modal Priors for Super Resolution, XPSR), etc. These solutions are generally designed and adjusted from the perspective of prompt word engineering, reference condition introduction, image semantic alignment, etc., to achieve a balance between detail generation richness and fidelity. Summary of the invention

[0004] The present disclosure provides a model training method, an image generation method, a model training device, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product, to at least solve the problem that the Chinese raw image model in the related art causes the generated image to have hallucinations, and the fidelity of the restored image is insufficient and the detail generation effect is not clear enough. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of the present disclosure, a model training method is provided, comprising: obtaining a pre-constructed initial model and a sample image pair, the sample image pair comprising a degraded sample image and a target sample image; determining a semantic association vector corresponding to the degraded sample image; determining an initial input image of the initial model and a feature representation vector of the initial input image; performing image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image; and training the initial model according to the predicted generated image and the target sample image to obtain an image generation model.

[0006] In an exemplary embodiment of the present disclosure, the semantic association vector includes a semantic category vector and an image semantic vector, and determining the semantic association vector corresponding to the degraded sample image includes: obtaining a preconfigured image category recognition network; performing category recognition processing on the degraded sample image by the image category recognition network to obtain the semantic category vector; or performing pixel feature extraction processing and semantic feature extraction processing on the degraded sample image by the image category recognition network to obtain a pixel feature vector and a semantic feature vector; and generating the image semantic vector based on the pixel feature vector and the semantic feature vector.

[0007] In an exemplary embodiment of the present disclosure, determining the initial input image of the initial model and the feature representation vector of the initial input image includes: performing noise processing on the target sample image and using the target sample image after noise processing as the initial input image; and performing block and encoding processing on the initial input image to obtain the feature representation vector.

[0008] In an exemplary embodiment of the present disclosure, the initial model includes a first initial model, and the initial model performs image generation processing based on the semantic association vector and the feature representation vector to obtain a predicted generated image, including: taking the feature representation vector as an input of the first initial model, the first initial model includes multiple first image generation modules; obtaining a semantic category vector of the degraded sample image; and performing image generation processing based on the semantic category vector and the feature representation vector by the multiple first image generation modules to obtain the predicted generated image.

[0009] In an exemplary embodiment of the present disclosure, the first image generation module includes a first hybrid expert model layer and an adaptive normalization layer, the first hybrid expert model layer includes a first gating network and multiple first expert models; the multiple first image generation modules perform image generation processing based on the semantic category vector and the feature representation vector to obtain the predicted generated image, including: through the adaptive normalization layer, using the semantic category vector as the first image generation control condition of the first hybrid expert model layer; because the first gating network is based on the feature representation vector, a specified number of first target expert models are selected from the multiple first expert models; the first target expert model denoises the feature representation vector based on the first image generation control condition to obtain a first predicted feature vector; and the first predicted feature vector is decoded to obtain the predicted generated image.

[0010] In an exemplary embodiment of the present disclosure, the initial model includes a second initial model, and the initial model performs image generation processing based on the semantic association vector and the feature representation vector to obtain a predicted generated image, and also includes: obtaining an image semantic vector of the degraded sample image; using the image semantic vector and the feature representation vector as inputs of the second initial model, the second initial model includes multiple second image generation modules; and the multiple second image generation modules perform image generation processing based on the image semantic vector and the feature representation vector to obtain the predicted generated image.

[0011] In an exemplary embodiment of the present disclosure, the multiple second image generation modules perform image generation processing based on the image semantic vector and the feature representation vector to obtain the predicted generated image, including: the cross-attention module of the second image generation module performs feature extraction processing on the image semantic vector and the feature representation vector to obtain fused image features; the feature representation vector is input into the second hybrid expert model layer of the second image generation module, the second hybrid expert model layer includes a second gating network and multiple second expert models; since the second gating network is based on the feature representation vector, a specified number of second target expert models are selected from the multiple second expert models; the second target expert model performs denoising processing on the feature representation vector based on the fused image features to obtain a second predicted feature vector; and the second predicted feature vector is decoded to obtain the predicted generated image.

[0012] In an exemplary embodiment of the present disclosure, the predicted generated image includes a plurality of noisy images under different prediction states, and the initial model is trained according to the predicted generated image and the target sample image to obtain the image generation model, including: determining the noisy images under the plurality of different prediction states based on the target sample image, and the noise data corresponding to each of the noisy images; constructing a prediction speed function based on the prediction speed at which the noisy image approaches the target sample image; determining the noise image difference according to the difference between the noise data and the target sample image; constructing a model loss function according to the prediction speed function and the noise image difference, and training the initial model based on the model loss function to obtain the image generation model.

[0013] According to a second aspect of the present disclosure, there is provided an image generation method, comprising: obtaining an image to be processed; obtaining a pre-trained image generation model, wherein the image generation model is trained based on a model training method; and performing image generation processing on the image to be processed by the image generation model to obtain a target generated image.

[0014] According to a third aspect of the present disclosure, a model training device is provided, comprising: an image pair acquisition module, used to acquire a pre-constructed initial model and a sample image pair, wherein the sample image pair comprises a degraded sample image and a target sample image; a semantic vector determination module, used to determine a semantic association vector corresponding to the degraded sample image; an image vector determination module, used to determine an initial input image of the initial model and a feature representation vector of the initial input image; a predicted image generation module, used to perform image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image; and a model training module, used to train the initial model according to the predicted generated image and the target sample image to obtain an image generation model.

[0015] In an exemplary embodiment of the present disclosure, the semantic association vector includes a semantic category vector and an image semantic vector, and the semantic vector determination module includes a semantic vector determination unit, which is used to: obtain a preconfigured image category recognition network; perform category recognition processing on the degraded sample image by the image category recognition network to obtain the semantic category vector; or perform pixel feature extraction processing and semantic feature extraction processing on the degraded sample image by the image category recognition network to obtain a pixel feature vector and a semantic feature vector; generate the image semantic vector based on the pixel feature vector and the semantic feature vector.

[0016] In an exemplary embodiment of the present disclosure, the image vector determination module includes an image vector determination unit, which is used to: perform noise processing on the target sample image and use the target sample image after noise processing as the initial input image; and perform block and encoding processing on the initial input image to obtain the feature representation vector.

[0017] In an exemplary embodiment of the present disclosure, the initial model includes a first initial model, and the predicted image generation module includes a first image generation unit, which is used to: use the feature representation vector as an input of the first initial model, and the first initial model includes multiple first image generation modules; obtain a semantic category vector of the degraded sample image; and perform image generation processing based on the semantic category vector and the feature representation vector by the multiple first image generation modules to obtain the predicted generated image.

[0018] In an exemplary embodiment of the present disclosure, the first image generation module includes a first hybrid expert model layer and an adaptive normalization layer, the first hybrid expert model layer includes a first gating network and multiple first expert models; the first image generation unit includes a first image generation subunit, which is used to: through the adaptive normalization layer, use the semantic category vector as the first image generation control condition of the first hybrid expert model layer; because the first gating network is based on the feature representation vector, a specified number of first target expert models are selected from the multiple first expert models; the first target expert model denoises the feature representation vector based on the first image generation control condition to obtain a first predicted feature vector; and decodes the first predicted feature vector to obtain the predicted generated image.

[0019] In an exemplary embodiment of the present disclosure, the initial model includes a second initial model, and the predicted image generation module includes a second image generation unit, which is used to: obtain an image semantic vector of the degraded sample image; use the image semantic vector and the feature representation vector as inputs of the second initial model, and the second initial model includes multiple second image generation modules; and perform image generation processing based on the image semantic vector and the feature representation vector by the multiple second image generation modules to obtain the predicted generated image.

[0020] In an exemplary embodiment of the present disclosure, the second image generation unit includes a second image generation subunit, which is used to: perform feature extraction processing on the image semantic vector and the feature representation vector by a cross-attention module of the second image generation module to obtain a fused image feature; input the feature representation vector into a second hybrid expert model layer of the second image generation module, the second hybrid expert model layer including a second gating network and a plurality of second expert models; since the second gating network is based on the feature representation vector, a specified number of second target expert models are selected from the plurality of second expert models; perform denoising processing on the feature representation vector by the second target expert model based on the fused image feature to obtain a second predicted feature vector; and perform decoding processing on the second predicted feature vector to obtain the predicted generated image.

[0021] In an exemplary embodiment of the present disclosure, the predicted generated image includes multiple noisy images under different prediction states, and the model training module includes a model training unit, which is used to: determine the multiple noisy images under different prediction states based on the target sample image, and the noise data corresponding to each of the noisy images; construct a prediction speed function based on the predicted speed at which the noisy image approaches the target sample image; determine the noise image difference based on the difference between the noise data and the target sample image; construct a model loss function based on the prediction speed function and the noise image difference, and train the initial model based on the model loss function to obtain the image generation model.

[0022] According to a fourth aspect of the present disclosure, an image generating device is provided, comprising: an image acquiring module for acquiring an image to be processed; a model acquiring module for acquiring a pre-trained image generating model, wherein the image generating model is trained based on a model training method; and an image generating module for performing image generating processing on the image to be processed by the image generating model to obtain a target generated image.

[0023] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to execute instructions to implement any one of the above-mentioned model training methods, or to implement the above-mentioned image generation method.

[0024] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the above-mentioned model training methods, or implement the above-mentioned image generation method.

[0025] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements any one of the above-mentioned model training methods or the above-mentioned image generation method.

[0026] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0027] On the one hand, the semantic association vector of the image is used in the image generation process, so that the model can learn the semantic features of the image, ensuring that the final generated image is semantically consistent with the input image. On the other hand, the degraded image is used as the model input, without using text information to generate the image, so that the model can learn more image detail features, and finally output a generated image with better clarity and more detailed textures.

[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0030] Figure 1 The figure is a flowchart of a model training method according to an exemplary embodiment.

[0031] Figure 2 is a flowchart of the reasoning of LPM-DiT according to an exemplary embodiment.

[0032] Figure 3 is a flow chart of reasoning of LPM-MMDiT according to an exemplary embodiment.

[0033] Figure 4 is a detailed structural diagram of a DiT Block in LPM-DiT according to an exemplary embodiment.

[0034] Figure 5 is a detailed structural diagram of the MMDiT Block in the LPM-MMDiT according to an exemplary embodiment.

[0035] Figure 6 The figure is a flowchart of an image generating method according to an exemplary embodiment.

[0036] Figure 7 It is a simplified diagram of the reasoning process according to an exemplary embodiment.

[0037] Figure 8It is a block diagram of a model training device according to an exemplary embodiment.

[0038] Fig. 9 The figure is a block diagram of an image generating device according to an exemplary embodiment.

[0039] Fig.10 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0040] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0042] Common text-based graph models use text or categories as control conditions for image generation. For large models with image processing as their task, text information is not as dense as the information of the input image itself for modeling low-quality input. Specifically, in the mainstream diffusion model-based image restoration framework, mature text-based graph models are usually used as base models for development. Typical representatives include Stable Diffusion v1.5 (SDv1.5), SDXL (StableDiffusion XL), SD3, Flux.1, Kolors, etc. These models extract text features through language models such as Contrastive Language-Image Pre-Training (CLIP), large language model T5, and generative language model, and interact text and image features at each layer of the graph generation model to obtain text-controlled generation results.

[0043] The text-based graph model uses text description as the conditional input of the text-based graph model. For image restoration tasks, it is necessary to manually design text prompts or multimodal large language models (MLLMs) to generate text prompts (Captions) for low-quality images to achieve the conversion of "image (Image) → text description (Caption) → vector (Embedding)". Limited by the image quality damage of low-quality images, the Caption has incomplete coverage of content elements, resulting in the "hallucination" problem in the text annotation model in the large language model (LLM), resulting in the text description generated by the model for low-quality images may be ambiguous or inaccurate, which in turn affects the accuracy of texture generation of the base model. Therefore, the text-based graph model uses text as a condition and does not match the form of low-quality images as input in the processing model.

[0044] Based on this, according to the embodiments of the present disclosure, a model training method, an image generation method, a model training device, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product are proposed.

[0045] Figure 1 is a flow chart of a model training method according to an exemplary embodiment. Figure 1 As shown, the model training method can be used in a computer device, wherein the computer device described in the present disclosure may include mobile terminal devices such as mobile phones, tablet computers, laptops, PDAs, and fixed terminal devices such as desktop computers. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.

[0046] Step S110, obtaining a pre-built initial model and a sample image pair, where the sample image pair includes a degraded sample image and a target sample image;

[0047] Step S120, determining a semantic association vector corresponding to the degraded sample image;

[0048] Step S130, determining an initial input image of an initial model and a feature representation vector of the initial input image;

[0049] Step S140, performing image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image;

[0050] Step S150, training the initial model according to the predicted generated image and the target sample image to obtain an image generation model.

[0051] According to the model training method in this example embodiment, on the one hand, the semantic association vector of the image is used in the image generation process, so that the model can learn the semantic features of the image, ensuring that the final generated image is semantically consistent with the input image. On the other hand, the degraded image is used as the model input, and there is no need to use text information to generate the image, so that the model can learn more image detail features and finally output a generated image with better clarity and more detailed textures.

[0052] The model training method in this example embodiment will be further described below.

[0053] In step S110, a pre-built initial model and a sample image pair are obtained, where the sample image pair includes a degraded sample image and a target sample image.

[0054] In an exemplary embodiment of the present disclosure, the initial model may be a pre-built image generative model based on a diffusion model. Unlike the image-generating model, the condition for generating the model in the present disclosure is an image, i.e., an image-generating model. The sample image pair may be an image data pair used to train the initial model, including an image pair consisting of a low-quality image (LQ Image) and a high-quality image (HQ Image).

[0055] A pre-built initial model and a sample image pair for model training are obtained, wherein the degraded sample image in the sample image pair may be image data after image degradation, and the degraded sample image may be an image whose image quality index value is less than a pre-defined quality parameter threshold. The target sample image may be image data with good picture quality, such as an image whose image quality index value is greater than or equal to a pre-defined quality parameter threshold. Image quality indicators may include, but are not limited to, image resolution, color depth, signal to noise ratio (SNR), contrast, peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM).

[0056] Image degradation refers to the phenomenon that image quality decreases during the formation, recording, processing and transmission of images due to imperfections in the imaging system, recording equipment, transmission media and processing methods. In actual application scenarios, due to factors such as limitations of shooting equipment, compression encoding process, multiple uploads and downloads, the image quality is reduced. This part of the image can be regarded as a low-quality image, namely LQ sample data. The target sample image can be a high-quality image with a higher resolution, namely HQ sample data.

[0057] The present disclosure pre-constructs an image generation model based on a diffusion model. The condition for generating the model is an image, that is, an image-generated-image model. Through model training, the model can learn relevant semantic features and image features in the image for processing subsequent image generation tasks.

[0058] In step S120, a semantic association vector corresponding to the degraded sample image is determined.

[0059] In an exemplary embodiment of the present disclosure, the semantic association vector may be a vector related to the semantic features contained in the degraded sample image.

[0060] Semantic features are extracted from the degraded sample images to obtain the corresponding semantic association vector, which is then used as the input of the initial model for the training process of the initial model.

[0061] In an exemplary embodiment of the present disclosure, for step S120, determining the semantic association vector corresponding to the degraded sample image includes: obtaining a preconfigured image category recognition network; performing category recognition processing on the degraded sample image by the image category recognition network to obtain a semantic category vector; or performing pixel feature extraction processing and semantic feature extraction processing on the degraded sample image by the image category recognition network to obtain a pixel feature vector and a semantic feature vector; generating an image semantic vector based on the pixel feature vector and the semantic feature vector.

[0062] The semantic category vector may be a vector corresponding to a specific category to which the image content in the degraded sample image belongs, and the image semantic vector may be a vector corresponding to a semantic feature contained in the degraded sample image.

[0063] After obtaining the degraded sample image, the present disclosure uses an image category recognition network (such as a text-image pair pre-training model) to extract semantic features from the degraded sample image to obtain a corresponding semantic association vector. For example, a contrastive language-image pre-training (CLIP) model or a sigmoid loss for language image pre-training model (SigLIP model) based on a sigmoid loss function is used to extract relevant semantic features in the degraded sample image to obtain a semantic association vector. The model structure of the CLIP model and the SigLIP model includes a text recognition network structure for recognizing text input, and an image recognition network structure for performing category recognition and semantic recognition on image input.

[0064] The present disclosure uses a pre-trained image category recognition network to perform category recognition processing on the degraded sample image to obtain the semantic category vector corresponding to the degraded sample image. Figure 2 , Figure 2 It is a flow chart of LPM-DiT reasoning according to an exemplary embodiment. For the input image x (i.e., the degraded sample image), the image encoder (Image Encoder) model of CLIP or SigLIP is first used for category recognition processing to obtain a vector (embedding) of 1x768 dimensions as the initial semantic category vector; wherein, the features of the category database are K-Means clustered by the CLIP features of the pre-configured 50 million high-quality data sets. Depending on the model version, the cluster center can be set to N. In the specific implementation, N can be set to integer values ​​such as 1000, 5000, 10000, and the feature vector of the corresponding category database is Nx768.

[0065] Then, a nearest neighbor search is performed in the category database, and a candidate semantic category vector closest to the initial semantic category vector is selected from the category database as the semantic category vector corresponding to the degraded sample image. The semantic category vector determined subsequently is Figure 2 The semantic category label (Class Label) in is added to the initial model as an image generation condition for model training of the initial model.

[0066] In addition, the present disclosure also provides an implementation scheme for extracting image semantic vectors of degraded sample images, referring to Figure 3 , Figure 3 is a flow chart of reasoning of LPM-MMDiT according to an exemplary embodiment. Figure 3For the input image x (i.e., the degraded sample image), the image encoder model in CLIP or SigLIP is first used to extract pixel features and semantic features of the sample degraded image, respectively, to obtain a 256x768-dimensional pixel feature vector containing the image pixel features and a 1x768-dimensional semantic feature vector (semantic embedding) containing the image semantic features. Then, the two embeddings are concatenated to obtain a 257x768-dimensional feature as the image semantic vector. The obtained image semantic vector is then used as the input of the initial model and as a condition for model generation.

[0067] Different from the text-based graph model whose input is mainly text, the present disclosure uses images as the input of the initial model. Before the image is input into the initial model, the CLIP and SigLIP models can be used to extract the semantic association vector of the image, and use it as the generation condition for the subsequent model to generate the image.

[0068] In step S130, an initial input image of an initial model and a feature representation vector of the initial input image are determined.

[0069] In an exemplary embodiment of the present disclosure, the initial input image may be an image obtained by performing noise addition processing on the target sample image, and the feature representation vector may be a representation vector obtained by performing feature extraction on the initial input image.

[0070] In the process of training the initial model, an image with random noise is used as the input of the initial model. For example, the present disclosure can introduce noise into the target sample image through the forward diffusion process of the diffusion model to obtain a noisy image sample as the initial input image. In addition, image features are extracted from the initial input image to obtain a corresponding feature representation vector.

[0071] In an exemplary embodiment of the present disclosure, for step S130, the initial input image of the initial model and the feature representation vector of the initial input image are determined, including: performing noise processing on the target sample image and using the target sample image after noise processing as the initial input image; performing block and encoding processing on the initial input image to obtain the feature representation vector.

[0072] Obtain a target sample image, introduce noise into the target sample image based on the forward diffusion process of the diffusion model, and obtain a noisy image sample as the initial input image. After obtaining the initial input image, a pre-trained variational autoencoder (VAE) can be used to extract image features. The VAE encoder can perform block processing and encoding processing on the initial input image, and the image is encoded in the latent space as a feature image representation (Imagetoken) as a feature representation vector. VAE can include a VAE encoder (VAE-Encoder) and a VAE decoder (VAE-Decoder). The VAE encoder (VAE-Encoder) is used to compress the input image into a latent variable distribution, and the VAE decoder (VAE-Decoder) is used to decode the latent variable into an image.

[0073] The VAE-Encoder used in the related text graph model or category generation model compresses the input image 3×H×W (3 channels, height H and width W) to the feature dimension 4×H / 8×W / 8, that is, 4 channels, and the width and height become 1 / 8 of the original VAE-C4 model. The VAE encoder has a large compression rate for the feature space, which will cause loss and distortion of the original information of the input image. In order to overcome the above defects, the present disclosure adopts a VAE model architecture similar to the third version of the diffusion model (Stable Diffusion version3, SDv3), which increases the 4 channels after compression encoding to 16 channels, that is, the VAE-C16 model is used to encode and decode the image.

[0074] Moreover, the present disclosure retrained the VAE model. The VAE-C16 in this embodiment is trained with an open source image dataset (Open Images) and an additional 20 million high-quality datasets. Compared with the structure of VAE-C4, the PSNR of the VAE in the present disclosure in the reconstruction result of the real image is improved by 4-6 decibels (dB) compared with the VAE model of the 1.5 diffusion model (Stable Diffusion version 1.5, SDv1.5), which can reduce the loss and distortion of the original information. In addition, in the image encoding process, the present disclosure changes the absolute position encoding to the rotational position encoding (Rotary Position Embedding, RoPE), which can improve the semantic consistency of the model-generated images at different resolutions.

[0075] In step S140, the initial model performs image generation processing based on the semantic association vector and the feature representation vector to obtain a predicted generated image.

[0076] In an exemplary embodiment of the present disclosure, the predicted generated image may be an image obtained after the initial model performs image generation processing based on the semantic association vector and the feature representation vector.

[0077] During the training process of the initial model, the feature representation vector corresponding to the initial input image is used as the input of the initial model. In addition, the extracted semantic association vector can be introduced into the initial model using a specific conditional mechanism as an image generation condition for the initial model to generate the image. The initial model can perform image generation processing based on the semantic association vector and the feature representation vector to obtain a predicted generated image.

[0078] In an exemplary embodiment of the present disclosure, for step S140, an initial model is used to perform image generation processing based on a semantic association vector and a feature representation vector to obtain a predicted generated image, including: taking the feature representation vector as an input of a first initial model, the first initial model including multiple first image generation modules; obtaining a semantic category vector of the degraded sample image; and multiple first image generation modules are used to perform image generation processing based on the semantic category vector and the feature representation vector to obtain a predicted generated image.

[0079] The first initial model may be an image generation model with a large processing model based on a diffusion model of a transformer (LargeProcessing Model-Diffusion Image Transformer, LPM-DiT) as the core network architecture, indicating that the generation base model is a large processing model designed for image processing tasks. The first image generation module may be an image generation module included in the first initial model LPM-DiT, namely, a DiT Block.

[0080] Continue to refer Figure 2 For the initial input image, after encoding with the VAE encoder to obtain the corresponding feature representation vector, the feature representation vector (input image token) is used as the input of the first image generation module in the first initial model. In addition, the semantic category vector corresponding to the degraded sample image can also be obtained, and the semantic category vector is used as the control condition of the first image generation module with the category generation condition to guide the subsequent image generation process.

[0081] In an exemplary embodiment of the present disclosure, multiple first image generation modules perform image generation processing based on semantic category vectors and feature representation vectors to obtain a predicted generated image, including: using the semantic category vector as a first image generation control condition of the first hybrid expert model layer through an adaptive normalization layer; since the first gating network is based on the feature representation vector, a specified number of first target expert models are selected from multiple first expert models; the first target expert model denoises the feature representation vector based on the first image generation control condition to obtain a first predicted feature vector; and decoding the first predicted feature vector to obtain a predicted generated image.

[0082] Among them, the first hybrid expert model layer can be a mixed expert model layer (Mixed Expert Models, MoE) included in the first image generation module, and MoE is composed of a gating network (router) and multiple expert networks (Experts). The adaptive normalization layer (Adaptive Layer Normalization, AdaLN) can use the adaptive parameters in the model to perform layer normalization operations. The first gating network can be a control network for enabling one or more first expert models specified in the first hybrid expert model layer. The first expert model can be a network model implemented based on a feed-forward neural network (Feed-Forward Neural Network, FFN).

[0083] refer to Figure 4 , Figure 4 is a detailed structural diagram of a DiT Block in LPM-DiT according to an exemplary embodiment. Figure 4 The DiT Block in the paper is a Transformer structure, including a Multi-Head Self-Attention (MHA) module, a first hybrid expert model layer (MoE layer) and an adaptive normalization layer (AdaLN layer). LPM-DiT generates high-quality, realistic images by learning the gradual conversion process from random noise to the target image latent variables in the latent space, and increases the number of network layers from 28 layers of DiT to 32 layers, and the number of parameters from 0.7B (700 million) to 0.9B (900 million). The present invention adds a sparse hybrid expert model layer to the Transformer Block of each DiT; at the same time, unlike the existing solution DiT-MoE, this solution does not require a shared expert layer, so the model reasoning efficiency is higher.

[0084] Figure 4The MoE layer in improves the performance and efficiency of the model by combining multiple expert models. Specifically, in the Transformer architecture, the MoE layer contains multiple expert models implemented based on the pointwise feedforward network (FFN) structure. The following is the specific operation of the MoE model in the image generation process.

[0085] First, the feature representation vector (input image token) of the initial input image is used as the input of the first initial model. The input image token can be the feature image representation (Image token) corresponding to the image encoded in the latent space; then the input image token is input to the AdaLN layer in the first initial model, and feature extraction is performed through the multi-head self-attention (MHA) layer. Subsequently, the vector processed by the AdaLN layer and the MHA layer is input to the MoE layer.

[0086] The LPM-DiT disclosed in the present invention introduces query-key normalization (QK-Normalization) in the Self Attention part of the Transformer Block, that is, the Q and K matrices are normalized before performing the self-attention matrix operation. This operation can make the model training more stable.

[0087] Figure 4 The MoE model can include a router or a gated network. Taking the gated network as an example, the gated network can dynamically decide which expert models should be used to infer the input image token based on the characteristics of the input image token. For example, the gated network outputs a weight vector, which represents the degree of response of each expert model to the current input (the weight value of each expert). The gated network can select two expert models with the highest weight values. In other words, each image block feature (image token) can enter a specified number (such as 2) of expert models (i.e., the first target expert model) with the highest response, and the expert model can be a FFN network structure.

[0088] In addition, the hybrid expert model adopts a gating mechanism, which is part of the router and determines the contribution of each expert model to the final output. The output of each expert model is weighted averaged according to the weight provided by the gating network to obtain the final model output image feature (i.e., output image token) as the first predicted feature vector. After obtaining the first predicted feature vector, the VAE decoder is used to decode the latent variable into an image.

[0089] The hybrid expert model has the following advantages: different expert models can learn different semantic representations, making the model more flexible and powerful; during reasoning, only some expert models are activated, which can reduce the amount of calculation and improve efficiency; in the subsequent process, the capacity of the model can be expanded by adding more experts, which has good scalability. The hybrid expert model is introduced into the initial model of the present disclosure, which improves the model parameter quantity, performance upper limit and model generation effect while keeping the reasoning speed basically unchanged.

[0090] In an exemplary embodiment of the present disclosure, the initial model includes a second initial model, and the initial model performs image generation processing based on a semantic association vector and a feature representation vector to obtain a predicted generated image, and also includes: obtaining an image semantic vector of a degraded sample image; using the image semantic vector and the feature representation vector as inputs of the second initial model, the second initial model includes multiple second image generation modules; and the multiple second image generation modules perform image generation processing based on the image semantic vector and the feature representation vector to obtain a predicted generated image.

[0091] The second initial model may be an image generation model with a multimodal diffusion transformer processing model (Multimodal Diffusion Transformer-Diffusion Image Transformer, LPM-MMDiT) as the core network architecture. The second image generation module may be an image generation module included in the second initial model LPM-MMDiT, namely, an MMDiT block.

[0092] Continue to refer Figure 3 For the image semantic vector of the degraded sample image, the corresponding image generation condition mechanism can be used to take the image semantic vector as the image generation control condition of the second initial model. Figure 3 In the process, the feature representation vector is used as the input of the second image generation module (MMDiT Block) in the second initial model, and the image semantic vector is used as the image generation condition for each MMDiT Block. The second initial model can perform image generation processing based on the feature representation vector and the image semantic vector to obtain a predicted generated image. In this process, LPM-MMDiT generates high-quality and realistic images by learning the step-by-step conversion process from random noise to the target image latent variables in the latent space.

[0093] In an exemplary embodiment of the present disclosure, multiple second image generation modules perform image generation processing based on image semantic vectors and feature representation vectors to obtain predicted generated images, including: a cross-attention module of the second image generation module performs feature extraction processing on the image semantic vector and the feature representation vector to obtain fused image features; the feature representation vector is input into a second hybrid expert model layer of the second image generation module, the second hybrid expert model layer includes a second gating network and multiple second expert models; since the second gating network is based on the feature representation vector, a specified number of second target expert models are selected from multiple second expert models; the second target expert model performs denoising processing on the feature representation vector based on the fused image features to obtain a second predicted feature vector; and the second predicted feature vector is decoded to obtain a predicted generated image.

[0094] Among them, the second mixed expert model layer can be a mixed expert model layer (Mixed Expert Models, MoE) included in the second image generation module, and the second mixed expert model layer can also be composed of a gating network (or routing Router) and multiple expert networks (Experts). The second gating network can be a control network for enabling one or more second expert models specified in the second mixed expert model layer. The second expert model can be a network model implemented based on a feed-forward neural network (Feed-Forward Neural Network, FFN). The second mixed expert model layer adopts a cross attention mechanism as a conditional mechanism for image generation.

[0095] refer to Figure 5 , Figure 5 is a detailed structural diagram of the MMDiT Block in the LPM-MMDiT according to an exemplary embodiment. Figure 5 The MMDiT Block in the paper is a Transformer structure, including a multi-head cross-attention module (Multi-Head Cross-Attention), a second hybrid expert model layer (MoE layer) and an adaptive normalization layer (AdaLN layer). LPM-MMDiT generates high-quality, realistic images by learning the gradual conversion process from random noise to the target image latent variables in the latent space. The present disclosure also adds a hybrid expert model layer to the Transformer Block of each MMDiT, and the conditional mechanism is Cross-Attention.

[0096] The image semantic vector and feature representation vector input into the second image generation module can be subjected to feature extraction processing by the cross-attention module in the second image generation module to obtain fused image features; and used as the input of the subsequent second hybrid expert model layer. In the second hybrid expert model layer, the second target expert model selected by the second gating network performs denoising processing on the feature representation vector to obtain a second predicted feature vector. After obtaining the second predicted feature vector, the VAE decoder is used to decode the latent variables into a predicted generated image. The implementation process of the second hybrid expert model layer is the same as that of the first hybrid expert model layer, and this disclosure will not go into details. The second hybrid expert model layer can improve the performance and efficiency of the model by combining multiple expert models (Experts) together.

[0097] In step S150, the initial model is trained according to the predicted generated image and the target sample image to obtain an image generation model.

[0098] In an exemplary embodiment of the present disclosure, the image generation model may be a network model for processing image generation tasks. The image generation model is a graph-to-graph model, the model input is an image, the output is a generated image with better clarity and more detailed texture, and the generated image and the input image maintain semantic consistency.

[0099] For the LPM-DiT model, the degraded sample image is encoded into the latent space by the Image Encoder of CLIP (or SigLIP), and the latent space features are classified in the feature database to obtain the category C of the low-quality image. Category C is input as a condition to the LPM-DiT model. The entire reasoning process needs to be repeated T steps. At time t, the backbone model receives the noise or the state of the previous moment as input, performs the denoising process, and uses the rectified flow mechanism to perform denoising iterations. After multiple iterations (generally the number of reasoning steps can be configured as T = 25 to 50 times), the generated image result is obtained.

[0100] For the LPM-MMDiT model, the degraded sample image is encoded into the latent space by the Image Encoder of CLIP (or SigLIP) to obtain the latent space features, and the feature representation vector of the image is input as a condition into the Transformer attention layer of the LPM-MMDiT model. Similar to the LPM-DiT model, the entire reasoning process needs to be cycled for T steps. At time t (the tth iteration), the backbone model receives noise or the state of the previous moment (the previous iteration, such as the t-1th iteration) as input, performs the denoising process, and finally obtains the generated image result.

[0101] In an exemplary embodiment of the present disclosure, an initial model is trained according to a predicted generated image and a target sample image to obtain an image generation model, including: determining a plurality of noisy images under different prediction states based on the target sample image, and noise data corresponding to each noisy image; constructing a prediction speed function based on a predicted speed at which the noisy image approaches the target sample image; determining a noise image difference based on a difference between the noise data and the target sample image; constructing a model loss function based on the prediction speed function and the noise image difference, and training the initial model based on the model loss function to obtain an image generation model.

[0102] The noisy images in different prediction states may be noisy images in the t-th iteration (t-th moment) of the diffusion process. The noise data may be the noise contained in the noisy images in each prediction state. The prediction speed may be the speed at which the noisy image (such as the noisy image at the t-th moment) approaches the target sample image. The prediction speed function may be a prediction target calculation function constructed according to the prediction speed. The noise image difference may be the difference between the target sample image and the noise data.

[0103] This scheme changes the training method of the Denoising Diffusion Probabilistic Model (DDPM) to a flow matching model (Flow Matching) or a renormalized flow (Rectified-Flow), that is, the training objective is changed from predicting the variance of the noise to predicting the transport speed v of the noise to the data distribution. Specifically, the renormalized flow proposes a new diffusion noise formula, as shown in Formula 1.

[0104] x t =(1-t)x0+t∈,t∈[0,1] (Formula 1)

[0105] Among them, x0 can represent the target sample image (i.e., high-definition image), x t It can represent the noisy image at time t, that is, the noisy image under each prediction state, ∈ can represent Gaussian noise (that is, noise data); t∈ can represent the noise data at time t. The model can learn to gradually denoise from Gaussian noise to obtain x0.

[0106] In the related scheme, DDPM uses random noise in each step. Different from DDPM, R-Flow determines the target sample image (initial image x0) and the final noise image ∈, and then calculates each state x in the diffusion process. t The noisy image is determined as a linear interpolation between x0 and ∈. This linear deterministic diffusion process can make training easier, while the multi-step accumulated error in the inference phase is smaller, achieving faster sampling speed.

[0107] In the training phase, the re-rectified stream prediction target is the noisy image x at sampling time t. t Approach the speed v of the target sample x0 and construct a predicted speed function, as shown in Formula 2.

[0108]

[0109] Among them, x0 can represent the target sample image (i.e., high-definition image), x t can represent the noisy image at time t, that is, the noisy image under each prediction state, ∈ can represent Gaussian noise (that is, noise data); v θ (x t ,t) can represent the prediction speed.

[0110] During the model training process, the simplest training method is to optimize v to minimize the speed functions of the two systems, namely the predicted speed function v θ (x t ,t) and the square error between the noise image difference x0-∈. The loss function is constructed based on the predicted speed function and the noise image difference as shown in Formula 3.

[0111]

[0112] in, can represent the model loss function; x0-∈ can represent the noise image difference; v θ (x t ,t) can represent the prediction speed function; here v θ It is represented as LPM-DiT / LPM-MMDiT network. Compared with DDPM which takes noise as the prediction target, the prediction speed can more effectively guide the model to learn information close to pure noise samples.

[0113] In addition, compared to DiT, which trains a model from scratch for each resolution, this solution uses multi-level resolution training. First, the 256x256 model is trained for 1000Ksteps, then the 512x512 scale model is initialized based on the pre-training results of the 256x256 model and continued to train for 500Ksteps, and then the 1024x1024 base model is trained based on the 512x512 model for 100Ksteps.

[0114] Furthermore, the present invention introduces representation alignment (REPA) technology to accelerate model convergence, that is, for the output features of the 8th Transformer Block of LPM-DiT and LPM-MMDiT, the cosine similarity loss is calculated with the features of the input image x obtained by the DINOv2 pre-training model, so that the distance between the feature vectors of the two is as small as possible (that is, the cosine similarity is close to 1), and the feature space is aligned.

[0115] In summary, the model training method disclosed in the present invention obtains a pre-built initial model and a sample image pair, wherein the sample image pair includes a degraded sample image and a target sample image; determines the semantic association vector corresponding to the degraded sample image; determines the initial input image of the initial model and the feature representation vector of the initial input image; performs image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image; trains the initial model according to the predicted generated image and the target sample image to obtain an image generation model. On the one hand, the semantic association vector of the image is used in the image generation process, so that the model can learn the semantic features of the image, ensuring that the final generated image is semantically consistent with the input image. On the other hand, the degraded image is used as the model input, and there is no need to use text information to generate the image, so that the model can learn more image detail features, and finally output a generated image with better clarity and more detailed texture. On the other hand, the hybrid expert model is used in the image generation module, so that the image generation model has better flexibility and scalability, while improving the model calculation efficiency. On the other hand, the model is trained in a re-rectified flow manner, and the prediction speed is used as the training target to more effectively guide the model to learn information close to pure noise samples.

[0116] Figure 6 is a flowchart of an image generating method according to an exemplary embodiment. Figure 6 As shown, the image generation method can be used in a computer device. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.

[0117] Step S610, obtaining an image to be processed;

[0118] Step S620, obtaining a pre-trained image generation model, where the image generation model is trained based on a model training method;

[0119] Step S630: The image generation model performs image generation processing on the image to be processed to obtain a target generated image.

[0120] The technical solution of the present invention constructs an image generation model (Generative model) based on a diffusion model (Diffusion model). Different from the text-generated image model, the condition (Condition) of the generative model of the present invention is an image, that is, an image-generated image model. The model input of the present invention is an image, and the output is a generated image with better clarity and more detailed textures, while the generated image and the input image maintain semantic consistency.

[0121] In an exemplary embodiment of the present disclosure, an image to be processed is obtained, and the image to be processed may be a low-quality initial image to be processed for image restoration. The image to be processed is used as an input of the image generation model. Figure 7 , Figure 7 is a simplified diagram of a reasoning process according to an exemplary embodiment. Figure 7 It can be seen that compared with the related text-generated image model, which needs to use the text features of the image as input, the present invention directly uses the image as the input of the image generation model. The image generation model learns the gradual conversion process from random noise to the target image latent variables in the latent space through the model training process, and can finally generate high-quality and realistic images.

[0122] The image generation method disclosed in the present invention constructs a graph-to-graph model as an image generation model. The input of the image generation model is an image. By repairing the input image, a generated image with better clarity and more detailed texture can be output, and the generated image and the input image are semantically consistent.

[0123] Figure 8 is a block diagram of a model training device according to an exemplary embodiment. Figure 8 The model training device 800 includes: an image pair acquisition module 810, a semantic vector determination module 820, an image vector determination module 830, a predicted image generation module 840 and a model training module 850.

[0124] Specifically, the image pair acquisition module 810 is used to acquire a pre-built initial model and a sample image pair, the sample image pair including a degraded sample image and a target sample image; the semantic vector determination module 820 is used to determine the semantic association vector corresponding to the degraded sample image; the image vector determination module 830 is used to determine the initial input image of the initial model and the feature representation vector of the initial input image; the predicted image generation module 840 is used to perform image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image; the model training module 850 is used to train the initial model according to the predicted generated image and the target sample image to obtain an image generation model.

[0125] In an exemplary embodiment of the present disclosure, the semantic association vector includes a semantic category vector and an image semantic vector, and the semantic vector determination module 820 includes a semantic vector determination unit, which is used to: obtain a preconfigured image category recognition network; perform category recognition processing on the degraded sample image by the image category recognition network to obtain a semantic category vector; or perform pixel feature extraction processing and semantic feature extraction processing on the degraded sample image by the image category recognition network to obtain a pixel feature vector and a semantic feature vector; generate an image semantic vector based on the pixel feature vector and the semantic feature vector.

[0126] In an exemplary embodiment of the present disclosure, the image vector determination module 830 includes an image vector determination unit, which is used to: perform noise processing on the target sample image and use the target sample image after noise processing as the initial input image; and perform block and encoding processing on the initial input image to obtain a feature representation vector.

[0127] In an exemplary embodiment of the present disclosure, the initial model includes a first initial model, and the predicted image generation module 840 includes a first image generation unit, which is used to: use the feature representation vector as the input of the first initial model, and the first initial model includes multiple first image generation modules; obtain the semantic category vector of the degraded sample image; and perform image generation processing based on the semantic category vector and the feature representation vector by the multiple first image generation modules to obtain a predicted generated image.

[0128] In an exemplary embodiment of the present disclosure, the first image generation module includes a first hybrid expert model layer and an adaptive normalization layer, the first hybrid expert model layer includes a first gating network and multiple first expert models; the first image generation unit includes a first image generation subunit, which is used to: through the adaptive normalization layer, use the semantic category vector as the first image generation control condition of the first hybrid expert model layer; since the first gating network is based on the feature representation vector, a specified number of first target expert models are selected from multiple first expert models; the first target expert model denoises the feature representation vector based on the first image generation control condition to obtain a first predicted feature vector; and decodes the first predicted feature vector to obtain a predicted generated image.

[0129] In an exemplary embodiment of the present disclosure, the initial model includes a second initial model, and the predicted image generation module 840 includes a second image generation unit, which is used to: obtain an image semantic vector of the degraded sample image; use the image semantic vector and the feature representation vector as inputs of the second initial model, and the second initial model includes multiple second image generation modules; and perform image generation processing based on the image semantic vector and the feature representation vector by the multiple second image generation modules to obtain a predicted generated image.

[0130] In an exemplary embodiment of the present disclosure, the second image generation unit includes a second image generation subunit, which is used to: perform feature extraction processing on the image semantic vector and the feature representation vector by a cross-attention module of the second image generation module to obtain a fused image feature; input the feature representation vector to a second hybrid expert model layer of the second image generation module, the second hybrid expert model layer including a second gating network and multiple second expert models; since the second gating network is based on the feature representation vector, a specified number of second target expert models are selected from multiple second expert models; perform denoising processing on the feature representation vector based on the fused image feature by the second target expert model to obtain a second predicted feature vector; and decode the second predicted feature vector to obtain a predicted generated image.

[0131] In an exemplary embodiment of the present disclosure, the predicted generated image includes multiple noisy images under different prediction states, and the model training module 850 includes a model training unit, which is used to: determine multiple noisy images under different prediction states based on the target sample image, and noise data corresponding to each noisy image; construct a prediction speed function based on the predicted speed at which the noisy image approaches the target sample image; determine the noise image difference based on the difference between the noise data and the target sample image; construct a model loss function based on the prediction speed function and the noise image difference, and train the initial model based on the model loss function to obtain an image generation model.

[0132] Fig. 9 is a block diagram of an image generating device according to an exemplary embodiment. Fig. 9 The image generating device 900 includes: an image acquiring module 910 , a model acquiring module 920 and an image generating module 930 .

[0133] Specifically, the image acquisition module 910 is used to acquire the image to be processed; the model acquisition module 920 is used to acquire a pre-trained image generation model, where the image generation model is trained based on a model training method; and the image generation module 930 is used to perform image generation processing on the image to be processed by the image generation model to obtain a target generated image.

[0134] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0135] Refer to the following Fig.10 hereinafter describes the electronic device 1000 according to such an embodiment of the present disclosure. Fig.10 The electronic device 1000 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0136] like Fig.10As shown, the electronic device 1000 is in the form of a general computing device. The components of the electronic device 1000 may include but are not limited to: the at least one processing unit 1010, the at least one storage unit 1020, a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010), and a display unit 1040.

[0137] The storage unit stores program codes, which can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.

[0138] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1021 and / or a cache memory unit 1022 , and may further include a read-only memory unit (ROM) 1023 .

[0139] The storage unit 1020 may also include a program / utility 1024 having a set (at least one) of program modules 1025, such program modules 1025 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0140] Bus 1030 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0141] The electronic device 1000 may also communicate with one or more external devices 1070 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 1060. As shown, the network adapter 1060 communicates with other modules of the electronic device 1000 via a bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0142] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of the device to complete the above model training method and image generation method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0143] In an exemplary embodiment, a computer program product is also provided, including a computer program, which implements any one of the above-mentioned model training methods or image generation methods when executed by a processor.

[0144] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0145] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A model training method, characterized in that: include: Acquire a pre-built initial model and a sample image pair, wherein the sample image pair includes a degraded sample image and a target sample image; Determining a semantic association vector corresponding to the degraded sample image; Determining an initial input image of the initial model and a feature representation vector of the initial input image; Performing image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image; The initial model is trained according to the predicted generated image and the target sample image to obtain an image generation model.

2. The method according to claim 1, characterized in that: The semantic association vector includes a semantic category vector and an image semantic vector, and determining the semantic association vector corresponding to the degraded sample image includes: Get a pre-configured image category recognition network; The image category recognition network performs category recognition processing on the degraded sample image to obtain the semantic category vector; or The image category recognition network performs pixel feature extraction processing and semantic feature extraction processing on the degraded sample image to obtain a pixel feature vector and a semantic feature vector; The image semantic vector is generated based on the pixel feature vector and the semantic feature vector.

3. The method according to claim 1, characterized in that The determining of the initial input image of the initial model and the feature representation vector of the initial input image comprises: Performing noise processing on the target sample image, and using the target sample image after the noise processing as the initial input image; The initial input image is divided into blocks and encoded to obtain the feature representation vector.

4. The method according to claim 1, characterized in that: The initial model includes a first initial model, and the initial model is used to perform image generation processing based on the semantic association vector and the feature representation vector to obtain a predicted generated image, including: Using the feature representation vector as an input of the first initial model, wherein the first initial model includes a plurality of first image generation modules; Obtaining a semantic category vector of the degraded sample image; The plurality of first image generation modules perform image generation processing based on the semantic category vector and the feature representation vector to obtain the predicted generated image.

5. The method according to claim 4, characterized in that The first image generation module includes a first hybrid expert model layer and an adaptive normalization layer, and the first hybrid expert model layer includes a first gating network and a plurality of first expert models; the plurality of first image generation modules perform image generation processing based on the semantic category vector and the feature representation vector to obtain the predicted generated image, including: Using the semantic category vector as a first image generation control condition of the first hybrid expert model layer through the adaptive normalization layer; Since the first gating network is based on the feature representation vector, a specified number of first target expert models are selected from the plurality of first expert models; The first target expert model performs denoising processing on the feature representation vector based on the first image generation control condition to obtain a first predicted feature vector; The first predicted feature vector is decoded to obtain the predicted generated image.

6. The method according to claim 1, characterized in that The initial model includes a second initial model, and the initial model performs image generation processing based on the semantic association vector and the feature representation vector to obtain a predicted generated image, and further includes: Obtaining an image semantic vector of the degraded sample image; Using the image semantic vector and the feature representation vector as inputs of the second initial model, wherein the second initial model includes a plurality of second image generation modules; The plurality of second image generation modules perform image generation processing based on the image semantic vector and the feature representation vector to obtain the predicted generated image.

7. The method according to claim 6, characterized in that The step of performing image generation processing based on the image semantic vector and the feature representation vector by the multiple second image generation modules to obtain the predicted generated image includes: The cross attention module of the second image generation module performs feature extraction processing on the image semantic vector and the feature representation vector to obtain a fused image feature; Inputting the feature representation vector into a second hybrid expert model layer of the second image generation module, wherein the second hybrid expert model layer includes a second gating network and a plurality of second expert models; Since the second gating network is based on the feature representation vector, a specified number of second target expert models are selected from the plurality of second expert models; The second target expert model performs denoising processing on the feature representation vector based on the fused image feature to obtain a second predicted feature vector; The second predicted feature vector is decoded to obtain the predicted generated image.

8. The method according to any one of claims 1 to 7, characterized in that: The predicted generated image includes a plurality of noisy images in different predicted states, and the initial model is trained according to the predicted generated image and the target sample image to obtain an image generation model, including: Determine, based on the target sample image, the multiple noisy images in different prediction states, and the noise data corresponding to each of the noisy images; Constructing a prediction speed function based on the prediction speed at which the noisy image approaches the target sample image; Determining a noise image difference according to a difference between the noise data and the target sample image; A model loss function is constructed according to the predicted speed function and the noise image difference, and the initial model is trained based on the model loss function to obtain the image generation model.

9. An image generation method, characterized in that: include: Get the image to be processed; Obtain a pre-trained image generation model, wherein the image generation model is trained based on the model training method according to any one of claims 1 to 8; The image generation model performs image generation processing on the image to be processed to obtain a target generated image.

10. A model training device, characterized in that: include: An image pair acquisition module, used to acquire a pre-built initial model and a sample image pair, wherein the sample image pair includes a degraded sample image and a target sample image; A semantic vector determination module, used to determine the semantic association vector corresponding to the degraded sample image; An image vector determination module, used to determine an initial input image of the initial model and a feature representation vector of the initial input image; A predicted image generation module, used to perform image generation processing based on the semantic association vector and the feature representation vector by the initial model to obtain a predicted generated image; The model training module is used to train the initial model according to the predicted generated image and the target sample image to obtain an image generation model.

11. An image generating device, characterized in that: include: An image acquisition module, used for acquiring an image to be processed; A model acquisition module, used to acquire a pre-trained image generation model, wherein the image generation model is trained based on the model training method according to any one of claims 1 to 8; The image generation module is used to perform image generation processing on the image to be processed by the image generation model to obtain a target generated image.

12. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the model training method as described in any one of claims 1 to 8, or to implement the image generation method as described in claim 9.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the model training method as described in any one of claims 1 to 8, or implements the image generation method as described in claim 9.

Citation Information

Cited By

  • Industrial surface defect generation method and system based on few samples

    CN121391788A