Training method and device of image generation model and image generation method and device
By introducing the feature information of the original image into the image generation system and combining it with noise estimation and denoising processing, the problem of unsatisfactory image generation in complex backgrounds is solved, and high-quality images that meet user needs are generated in a variety of scenarios.
Patent Information
- Application Number
- CN202510916314.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-10
AI Technical Summary
Existing image generation systems are not ideal when generating defect images against complex backgrounds, and it is difficult to generate high-quality images that meet user needs.
By obtaining image-text training pairs in the training dataset, including the original image, the target image and the text description, pixel feature extraction, semantic feature extraction, noise processing and denoising processing are used to combine the feature information of the original image to generate an image that is closer to the target image.
The generated image can retain the characteristic information of the original image under complex background, meet user needs, be applicable to various scenarios, and generate high-quality target images.
Smart Images

Figure CN120766060A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of neural networks, and in particular to a method and device for training an image generation model and a method and device for generating images. BACKGROUND
[0002] With the development of neural network technology, image generation systems have been increasingly widely applied.
[0003] At present, image generation systems are applied in many industrial fields to generate images meeting certain requirements. For example, in the detection of defects of industrial products using a neural network model, diverse defect images and non-defect images are needed as a training set to train a defect detection model. However, due to the high yield rate in industrial production, defect samples are very scarce. To obtain more types and more comprehensive defect images, image generation systems are needed to generate various defect images to form a training set for training the defect detection model.
[0004] Generally, when generating defect images of a certain product, the same defect is generated in a clean background area according to a known defect type, and then the generated defect is fused with the image of the product to obtain a defect image of the product. Therefore, the above image generation system is generally only applicable to products with clean backgrounds, and for products with complex backgrounds, the effect of the generated defect image is not ideal.
[0005] In addition to product defect detection, in other application scenarios, there is also a similar requirement that a high-quality image meeting user requirements should be generated based on an initial image. SUMMARY
[0006] The present application provides a method and device for training an image generation model and a method and device for generating images, which can provide an image generation model applicable to various scenarios and generate high-quality images meeting user requirements.
[0007] To achieve the above object, the present application adopts the following technical solutions:
[0008] A method for training an image generation model, comprising:
[0009] obtaining a training data set for training an image generation model; wherein the training data set comprises a plurality of image-text training pairs, each image-text training pair comprising an original image, a target image, and a text description expressing changes of the target image relative to the original image;
[0010] training the image generation model using the training data set;
[0011] wherein the training process for each image-text training pair comprises:
[0012] pixel feature extraction and semantic feature extraction are performed on the original image respectively, and pixel features of the original image and semantic features of the original image are obtained correspondingly;
[0013] image feature extraction is performed on the target image, and noise estimation and de-noising processing are performed on the extracted image features to obtain noise features;
[0014] Based on the pixel features of the original image, the semantic features of the original image and the guidance of the text description, noise estimation and de-noising processing are performed on the noise features, and image restoration is performed on the result features obtained by the de-noising processing to obtain a generated image.
[0015] The generated image is compared with the target image, and the parameters of the pixel feature extraction, the semantic feature extraction, the noise estimation and the de-noising processing are updated based on the comparison result, so that the generated image is closer to the target image.
[0016] Preferably, the method further comprises:
[0017] Any text description and its corresponding any original image are input into the current trained image generation model, and the current generated image is obtained by inference using the image generation model. When it is determined that the current generated image meets the quality requirement, the any original image, the any text description and the current generated image are combined to form an image text training pair, which is added to the training data set for training to generate the image generation model.
[0018] Preferably, the inference using the image generation model to obtain the current generated image comprises:
[0019] Pixel feature extraction is performed on the any original image to obtain pixel features of the any original image, and semantic feature extraction is performed on the any original image to obtain semantic features of the any original image.
[0020] Randomly generated noise features are obtained, noise estimation and de-noising processing are performed on the noise features based on the pixel features of the any original image, the semantic features of the any original image and the guidance of any text description, and image restoration is performed on the result features obtained by the de-noising processing to obtain the current generated image.
[0021] Preferably, the pixel feature extraction on the original image comprises:
[0022] Image feature extraction is performed on the original image to obtain image features of the original image.
[0023] Pixel-level feature extraction is performed on the image features of the original image to obtain pixel-level features of the original image.
[0024] Preferably, the image feature extraction of the target image is completed by using the first VAE encoding unit, the image feature extraction of the original image is completed by using the second VAE encoding unit, the noise estimation and denoising processing are completed by using the denoising unit, the pixel-level feature extraction of the image feature of the original image is completed by using the reference denoising unit, and the image restoration is completed by using the first VAE decoding unit.
[0025] The network structure of the reference denoising unit and the denoising unit is the same.
[0026] Preferably, the noise estimation and denoising processing of the noised feature by using the denoising unit comprises:
[0027] The features of each layer in the pixel features of the original image output by the reference denoising unit are fused with the output features of the same level in the denoising unit, and the semantic features of the original image are fused with the output features of each level in the denoising unit.
[0028] Preferably, the reference denoising unit and the denoising unit are both UNET networks, or the reference denoising unit and the denoising unit are both transformer networks.
[0029] Preferably, the manner of obtaining the image-text training pair comprises:
[0030] For any text description, a pre-trained image understanding model is used to select one original image and one target image including the same subject in the original image data set and the target image data set corresponding to the any text description, and the selected original image, target image and the any text description are combined to form an image-text training pair.
[0031] Preferably, the processing of the image understanding model comprises:
[0032] An original image and a target image are extracted from the original image data set and the target image data set respectively, and consistency comparison is performed on the two extracted images, and if the consistency comparison result is greater than a set threshold, it is determined that the extracted original image and target image include the same subject.
[0033] An image generation method comprises:
[0034] An initial image and a text description provided by a user are input into an image generation model trained by the training method of the image generation model according to any one of claims 1 to 9; wherein the text description is used to express the change requirement for the initial image.
[0035] In the image generation model, pixel-level feature extraction is performed on the initial image to obtain pixel features of the initial image; and semantic feature extraction is performed on the initial image to obtain semantic features of the initial image.
[0036] Noise features are obtained, noise estimation and denoising processing are performed on the noise features based on the pixel features of the initial image, the semantic features of the initial image, and the guidance of the text description;
[0037] Image restoration is performed on the result features obtained through the denoising processing to obtain a generated image.
[0038] A training device of an image generation model comprises a training data set acquisition module and a training module.
[0039] The training data set acquisition module is configured to acquire a training data set for training an image generation model; wherein the training data set comprises a plurality of image-text training pairs, each of the image-text training pairs comprising an original image, a target image, and a text description used to express changes of the target image relative to the original image.
[0040] The training module is configured to train an image generation model using the training data set, and specifically comprises an original image processing sub-module, a first image feature extraction unit, a noise unit, a denoising unit, an image restoration unit, and a parameter updating unit.
[0041] The original image processing sub-module is configured to perform pixel-level feature extraction and semantic feature extraction on the original image respectively, so as to obtain pixel features of the original image and semantic features of the original image.
[0042] The first image feature extraction unit is configured to perform image feature extraction on the target image.
[0043] The noise unit is configured to perform noise processing on the image features extracted by the image feature extraction unit of the target image to obtain noise features.
[0044] The denoising unit is configured to perform noise estimation and denoising processing on the noise features based on the pixel features of the original image, the semantic features of the original image, and the guidance of the text description.
[0045] The image restoration unit is configured to perform image restoration on the result features obtained by the denoising unit to obtain a generated image.
[0046] The parameter updating unit is configured to compare the generated image with the target image, and update parameters of the pixel feature extraction unit and the denoising unit based on the comparison result, so that the generated image is closer to the target image.
[0047] Preferably, the training device further comprises a data set expansion unit configured to input any text description and its corresponding any original image into the current trained image generation model, perform inference on the image generation model to obtain a current generated image, and when it is determined that the current generated image meets the quality requirement, add the any original image, the any text description and the current generated image to the training data set to train the image generation model.
[0048] Preferably, the original image processing submodule comprises a second image feature extraction unit, a reference denoising unit and a semantic feature extraction unit.
[0049] The second image feature extraction unit is configured to perform image feature extraction on the original image to obtain image features of the original image.
[0050] The reference denoising unit is configured to perform pixel-level feature extraction on the image features of the original image to obtain pixel features of the original image.
[0051] The semantic feature extraction unit is configured to perform semantic feature extraction on the original image to obtain semantic features of the original image.
[0052] Preferably, the image feature extraction unit of the target image is a first VAE encoding unit, the image feature extraction unit of the original image is a second VAE encoding unit, and the image restoration unit is a first VAE decoding unit.
[0053] The network structure of the reference denoising unit is the same as that of the denoising unit.
[0054] An image generation device comprises an input module and an inference module of an image generation model.
[0055] The input module is configured to input an initial image and a text description provided by a user into the image generation model output by the training device of any one of the preceding image generation models.
[0056] The inference module of the image generation model is configured to perform inference on the image generation model based on the initial image and the text description to obtain a generated image.
[0057] The inference process comprises: performing pixel-level feature extraction on the initial image in the image generation model to obtain pixel features of the initial image; performing semantic feature extraction on the initial image to obtain semantic features of the initial image; obtaining randomly generated noise features, performing noise estimation and denoising processing on the noise features based on the pixel features of the initial image, the semantic features of the initial image and the guidance of the text description; and performing image restoration on the result features obtained by the denoising processing to obtain the generated image.
[0058] As can be seen from the above technical solution, in this application, a training data set for training an image generation model is obtained, and the training data set includes multiple image-text training pairs, each image-text training pair includes an original image, a target image, and a text description for expressing the changes of the target image relative to the original image; then, the training data set is used to train and generate an image generation model. The training process for each image-text training pair includes: pixel-level feature extraction and semantic feature extraction of the original image, respectively, to obtain the pixel features and semantic features of the original image, thereby providing feature information of the original image from two different aspects: semantics and pixels; image feature extraction of the target image, and noise processing of the extracted image features to obtain noisy features; then, based on the pixel features of the original image, the semantic features of the original image, and the guidance of the text description, noise estimation and denoising processing are performed on the noisy features. In this way, during the denoising process, the denoised features introduce the feature information of the original image and meet the user requirements of text description; next, the denoised features are restored to obtain the generated image of this round of training; finally, the generated image is compared with the target image, and the parameters of the pixel feature extraction unit and the denoising unit are updated based on the comparison results to make the generated image closer to the target image. The image generation model generated by the above method introduces the features of the original image during the image generation process, so that the features used to generate the target image carry the information of the original image, thereby effectively restoring the complex background of the original image, and then generating the target image that meets the text description requirements based on the original image. Therefore, this image generation model can be applied to a variety of scenarios and generate high-quality images that meet user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A schematic diagram of the basic process of the training method of the image generation model in this application;
[0060] Figure 2 Schematic diagram of the training process for each image-text training pair in this application;
[0061] Figure 3 Schematic diagram of the basic process of the image generation method in this application;
[0062] Figure 4 This is a schematic diagram of a specific process of the training method of the image generation model in Example 1 of the present application;
[0063] Figure 5 Schematic diagram of the model structure of the image generation model in Example 1 of the present application;
[0064] Figure 6 This is a schematic diagram of the process of adding data to the training data set in Example 1 of the present application;
[0065] Figure 7 The specific flowchart of the image generation method in Embodiment Two of the present application is shown in the following.
[0066] Figure 8 The network architecture diagram of the trained image generation model in Embodiment Two of the present application is shown in the following. DETAILED DESCRIPTION
[0067] In order to make the purposes, technical means and advantages of the present application more clear, the present application is further described in detail below with reference to the accompanying drawings.
[0068] The present application provides a training method of an image generation model, and simultaneously utilizes the trained image generation model to give an image generation method.
[0069] Figure 1 The basic flowchart of the training method of the image generation model in the present application is shown in the following. As shown in the following, Figure 1 the method comprises:
[0070] Step 101, obtaining a training data set for training an image generation model.
[0071] The training data set comprises a plurality of image-text training pairs, each image-text training pair comprises an original image, a target image and a text description, and the text description is used to express the changes of the target image relative to the original image. That is to say, each image-text training pair is used to express that the original image becomes the target image after the changes of the text description.
[0072] For example, the original image is an image of a defect-free industrial product, the target image is an image of a defective industrial product, and the text description is "defect". Of course, the text description can also be a specific defect type, such as "scratch", and the original image and the target image are respectively an image of the same product without scratches and an image of the same product with scratches. Or, the text description is "holding a mobile phone", and the original image and the target image are respectively an image of a person holding a mobile phone and an image of a person not holding a mobile phone.
[0073] Step 102, training the image generation model by using the training data set.
[0074] The image generation model of the present application is a neural network model, which is used to generate a new image consistent with the text description on the basis of the original image.
[0075] As described above, the training data set comprises a plurality of image-text training pairs, and each image-text training pair is used as a set of training data to train the image generation model.
[0076] The training process for each image-text training pair is shown in the following, Figure 2 and specifically comprises:
[0077] Step 102a, pixel-level feature extraction and semantic feature extraction are performed on the original image respectively, and pixel features of the original image and semantic features of the original image are obtained correspondingly;
[0078] The pixel-level feature extraction is performed on the original image to obtain the pixel features of the original image. The pixel feature extraction processing can adopt an existing feature extraction manner. For example, image feature extraction is first performed on the original image to obtain image features of the original image, and then pixel-level feature extraction is performed on the image features of the original image to obtain the pixel features of the original image. The image feature extraction processing can be implemented by using a VAE encoder, and the pixel-level feature extraction processing can be implemented by using a UNET or a transformer network structure.
[0079] The semantic feature extraction is performed on the original image to obtain the semantic features of the original image. The semantic feature extraction can adopt an existing semantic feature extraction manner. For example, an IP-Adapter network.
[0080] Step 102b, image feature extraction is performed on the target image, and noise processing is performed on the extracted image features to obtain noise features;
[0081] The image feature extraction is performed on the target image, and the image features of the target image can be obtained by using an existing image feature extraction manner. For example, a VAE encoder is used to perform image feature extraction on the target image.
[0082] The noise processing is performed on the image features of the target image, and noise can be gradually added to the image features of the target image to make them approximate pure noise.
[0083] In addition, the step 102a and the step 102b can be performed in parallel or in any order.
[0084] Step 102c, noise estimation and de-noising processing are performed on the noise features based on the pixel features of the original image, the semantic features of the original image and the guidance of the text description;
[0085] The noise estimation and de-noising processing are performed on the noise features, which can be implemented by using a UNET or a transformer network structure. In existing industrial product defect image generation models, noise estimation and de-noising processing are also included, and a UNET network is usually selected for implementation. However, in the existing model, the noise estimation and de-noising processing are only performed under the guidance of the text description, and do not include the feature information of the original image.
[0086] In the noise estimation and denoising processing in the present application, the pixel features and semantic features of the original image are introduced to guide the noise estimation and denoising from the pixel level and the semantic level, so that the resulting features after denoising processing include the pixel and semantic information of the original image, and of course, the background information of the original image.
[0087] Step 102d, image restoration is performed on the resulting features obtained by denoising processing to obtain a generated image.
[0088] The processing of the present step can be realized in the existing manner, which is usually a symmetrical operation with the image feature extraction in step 102b, that is, a process of obtaining the pixel values of the image from the feature information. For example, if the image feature extraction of step 102b is realized by a VAE encoder, a VAE decoder can be used to realize image restoration in the present step, which can also be considered as image generation, obtaining a new image, which is called generated image. Since this image is obtained based on the resulting features after denoising processing, and the resulting features include the pixel and semantic information of the original image, the generated image obtained here contains the content of the original image, that is, the changes expressed by the text description are added to the original image.
[0089] Step 102e, comparing the generated image with the target image, and updating the parameters of pixel feature extraction, semantic feature extraction, noise estimation and denoising processing based on the comparison result, so that the generated image is closer to the target image.
[0090] The training target is that the image generation model can generate the target image based on the original image and the text description, so after obtaining the generated image by the image generation model, it is compared with the target image. The target image here refers to the target image belonging to the same image-text training pair as the original image and the text description, and usually the main body of the original image and the target image is the same.
[0091] Based on the comparison result of the generated image and the target image, the loss function is determined, and the parameters of the image generation model are updated based on the loss function, so that the generated image is closer to the target image. The parameters updated here at least include the parameters of step 102a pixel feature extraction and semantic feature extraction, step 102c noise estimation and denoising processing. As for the parameters of image feature extraction in step 102b and image restoration in step 102d, they can be determined by pre-training, of course, they can also be updated in step 102e.
[0092] The above Figure 2The training process shown is for training a pair for one image sample, and a plurality of image samples in the training data set are trained according to the above process. After each training is completed, it can be judged whether the current image generation model meets the requirements. If it meets the requirements, the training is ended. If it does not meet the requirements, the next round of training is continued. Wherein, whether the current image generation model meets the requirements can be judged by using the existing method.
[0093] So far, the basic flow of the image generation model training method in this application is completed. Using the image generation model obtained by the above training, image generation can be performed. The basic flow of the image generation method given in this application is shown as Figure 3 , which specifically includes:
[0094] Step 301, input the initial image and text description provided by the user into the image generation model.
[0095] Wherein, the initial image refers to which image is expected to be changed, and the text description is used to express the change requirement of the initial image. For example, when an image of an industrial product with defects needs to be generated, the initial image can be an image of a normal industrial product, and the text description gives the defect type. The image generation model here is the image generation model generated in the foregoing Figure 1 and Figure 2 In fact, the initial image is equivalent to the original image in the foregoing training process. In order to distinguish from the training data, it is referred to as an initial image here.
[0096] Next, through the processing of steps 302-304, the inference process of the image generation model is completed to obtain the generated image.
[0097] Step 302, in the image generation model, pixel-level feature extraction is performed on the initial image to obtain the pixel feature of the initial image; semantic feature extraction is performed on the initial image to obtain the semantic feature of the initial image.
[0098] The processing of this step is the same as step 102a in Figure 2 .
[0099] Step 303, obtain randomly generated noise features, and based on the pixel feature of the initial image, the semantic feature of the initial image and the guidance of the text description, perform noise estimation and denoising processing on the noise features.
[0100] In the inference process, there is no target image, so the noise features are randomly generated and used as the input of noise estimation and denoising processing, which is equivalent to the noise feature in the foregoing training process. The specific noise estimation and denoising processing are the same as step 102c.
[0101] Step 304, perform image restoration on the result features obtained by the denoising processing to obtain the generated image.
[0102] The processing of this step is the same as step 102d in FIG. 10. Figure 2
[0103] So far, the image generation method flow in this application ends.
[0104] The above is the training method of the most basic image generation model in this application and the image generation method based on the trained image generation model. In the above method, the pixel features and semantic features of the original image or the initial image are introduced in the process of generating the image to guide the image generation, so that the generated image is generated according to the requirements of the text description on the basis of the initial image, effectively preserving the background content of the initial image, suitable for various scenes and needs and various types of initial images, and capable of generating high-quality images meeting user needs. At the same time, since the original image also participates in the training, the generation ability of the image generation model is stronger.
[0105] The specific implementation of the training method of the image generation model and the image generation method in this application will be described below through specific embodiments.
[0106] Embodiment One:
[0107] Figure 4 The specific flowchart of the training method of the image generation model in Embodiment One of this application is shown in FIG. 11. In this embodiment, in order to more appropriately introduce the feature information of the original image, the pixel feature extraction of the original image includes image feature extraction and pixel-level feature extraction in turn, wherein the image feature extraction of the original image is performed using the same structure as the image feature extraction of the target image, and the pixel-level feature extraction of the original image is performed using the same structure as the noise estimation and denoising processing of the target image. Figure 5 The model structure diagram of the image generation model in this embodiment is shown in FIG. 12. As shown in FIG. 12, the image generation model in this embodiment includes an image feature extraction module 1201, a noise estimation module 1202, a denoising processing module 1203, a pixel-level feature extraction module 1204, a text description processing module 1205, a text description feature extraction module 1206, a text description feature fusion module 1207, a text description feature processing module 1208, a target image generation module 1209, and a loss function module 1210. Figure 4 and Figure 5 The training method of this embodiment specifically includes:
[0108] Step 401, obtaining a training data set.
[0109] The image-text training pair in the training data set can be artificially organized, or it can also be automatically organized through a neural network model.
[0110] In the automatic organization of image-text training pairs, for any text description A, a pre-trained image understanding model is used to select an original image and a target image including the same subject from the original image data set and the target image data set corresponding to the text description A, and the selected original image, target image and text description A form an image-text training pair. For example, when the text description is "defects", the original image data set includes a plurality of normal industrial product images, and the target image data set includes a plurality of defective industrial product images. Here, the original image data set and the target image data set are for a certain text description. In specific implementation, the data set of images of the same category can be pre-organized, and for a certain text description, a suitable data set is found as the original image data set and the target image data set corresponding thereto.
[0111] The specific processing of the above image understanding model can include: extracting an original image X and a target image Y from the original image data set and the target image data set respectively, performing consistency comparison on the two extracted images (for example, using a CLIP model to extract image features x1 of the original image X and image features y1 of the target image Y respectively, and comparing the image features x1 and y1 by cosine distance), and if the consistency comparison result is greater than a set threshold, it is determined that the extracted original image and target image include the same subject. The identification of the original image X and the target image Y can be output, and it is indicated that it can be used as the corresponding original image and target image. In this way, this pair of original image and target image and the corresponding text description can form an image-text training pair.
[0112] After obtaining the training data set, the image generation model is trained using each image-text training pair in the training data set. The following takes the processing process of an image-text training pair as an example to illustrate a round of training process of the image generation model:
[0113] Step 402, for a certain image-text training pair, determine its text description A, original image X and target image Y.
[0114] Step 403, perform image feature extraction on the original image X to obtain the image feature EmbeddingX of the original image X.
[0115] In this embodiment, the image feature extraction is performed by the encoding module (i.e. the second VAE encoding unit in the VAE network) of the VAE network, which is mainly considered that the VAE network structure is particularly suitable for feature processing of image generation type. The parameters of the second VAE encoding unit can be determined by pre-training. Figure 5
[0116] Step 404, perform pixel-level feature extraction on the image feature EmbeddingX of the original image X to obtain the pixel feature FeatureX1 of the original image X.
[0117] In this embodiment, UNET network is used for pixel-level feature extraction, as shown in Figure 5 As mentioned earlier, in order to more accurately introduce the feature information of the original image, the pixel-level feature extraction in this step is the same as the network structure for noise estimation and denoising processing in the target image, that is, UNET network is used.
[0118] To distinguish the two UNET networks for the original image and the target image, the UNET network for the original image is called reference UNET unit, and the UNET network for the target image is called UNET unit. The network structures of the two UNET units are the same, but the parameter updates are independent of each other. Of course, in this embodiment, only the pixel-level feature extraction using the UNET network structure is taken as an example for description, and other network structures can also be used to realize pixel-level feature extraction in actual application, such as transformer network, etc., and correspondingly, the noise estimation and denoising processing of the target image can also use the transformer network with the same structure.
[0119] In this step, in order to more comprehensively represent the pixel features of the original image, the reference UNET unit outputs all the feature information at each scale to obtain FeatureX1, that is, the output features of each level in the reference UNET unit are combined to form FeatureX1.
[0120] Step 405, semantic feature extraction is performed on the original image X to obtain the semantic feature FeatureX2 of the original image X.
[0121] The processing of this step is realized by the semantic feature extraction unit in Figure 5 The specific processing can be performed in the existing manner.
[0122] Step 406, image feature extraction is performed on the target image, and noise processing is performed on the extracted image feature.
[0123] In this embodiment, the image feature extraction of the target image is also performed by using the encoding module of the VAE network (such as the first VAE encoding unit in Figure 5 ), that is, the network structure used in step 403 is the same.
[0124] After the image feature extraction of the target image Y is performed by the first VAE encoding unit, the image feature EmbeddingY of the target image Y is obtained. Then, the noise is gradually added in the image feature EmbeddingY by the noise adding unit shown in Figure 5 , such as 1000-step noise, to obtain the noise feature noiseY, which is usually approximate pure noise.
[0125] Step 407, based on the pixel features of the original image, the semantic features of the original image and the guidance of the text description, the noise estimation and denoising processing of the noise feature is carried out.
[0126] In this embodiment, the noise estimation and denoising processing is realized by the UNET network, that is, the UNET unit in Figure 5 In the existing industrial defect image generation system, the noise estimation and denoising processing part realized by the UNET network is also included, and the noise estimation and denoising processing is carried out under the guidance of a single text description. Compared with the noise estimation and denoising processing of the existing image generation system, the main difference of the present application is that the pixel features and semantic features of the original image are introduced to guide the noise estimation and denoising processing.
[0127] Specifically, in this embodiment, in order to more appropriately introduce the pixel features FeatureX1 and semantic features FeatureX2 of the original image, the features of each layer in FeatureX1 output by the UNET unit are fused with the output features of the same level in the UNET unit, and the semantic features FeatureX2 are respectively fused with the output features of each level in the UNET unit. Through the above processing, the pixel features and semantic features of the original image are introduced into the features of the same level to realize the feature fusion of the corresponding level, and the noise estimation and denoising processing is completed based on this. The resulting feature after denoising can effectively preserve the content information of the original image.
[0128] Step 408, image restoration is performed on the resulting feature obtained by the denoising processing to obtain a generated image Y'.
[0129] In this embodiment, the image restoration processing is realized by the decoding module of the VAE network, that is, the first VAE decoding unit in Figure 5 The specific processing is the same as the decoding module of the existing VAE network, which restores the resulting feature into a pixel-level image to obtain a generated image.
[0130] Step 409, the generated image Y' is compared with the target image Y, and the parameters of the reference UNET unit, the semantic feature extraction unit and the UNET unit are updated based on the comparison result, so that the generated image is closer to the target image.
[0131] In this embodiment, the value of the loss function is calculated by comparing the generated image Y' and the target image Y, and the parameters of the reference UNET unit, the semantic feature extraction unit and the UNET unit are updated based on the value of the loss function. Of course, the parameters of the first VAE encoding unit, the first VAE decoding unit and the second VAE encoding unit can also be updated. The loss function can be selected as needed, such as MSE loss function, etc.
[0132] So far, the training of the image generation model is completed, and the generation of the image generation model is completed through several rounds of training.
[0133] The above Figure 5 In the image generation model architecture of the present embodiment, the network structure involved in the processing of the target image in the lower half (specifically including the first VAE encoding unit, the noise unit, the UNET unit, and the first VAE decoding unit) is the existing industrial defect image generation model. In the present application, by further introducing the processing of the original image in the upper half (i.e. Figure 5 the reference image understanding part in the reference image, specifically including the processing of the second VAE encoding unit, the image semantic feature extraction unit, and the reference UNET unit), the feature information of the original image is introduced into the noise estimation and denoising processing of the existing image generation model, so that the newly generated image not only meets the requirements of the text description, but also retains the content of the original image, and can meet the application of more scenes and adapt to various types of original images including complex backgrounds.
[0134] In addition, in the training of the image generation model, the larger the amount of data in the training data set and the richer the types, the better the accuracy of the image generation model obtained by training. Considering that in some scenarios, the number of target images in the training data set may be limited, for example, the number and types of real defect images may be limited, preferably, the processing shown in Figure 6 can be further included to increase the amount of data in the training data set:
[0135] Step 601, input a certain text description and its corresponding certain original image into the current image generation model obtained by training;
[0136] The current image generation model obtained by training can not be the image generation model obtained after training is completed, but the latest image generation model in the training process, which is referred to as the current image generation model hereinafter.
[0137] For a certain text description A1, determine its corresponding certain original image X1, and input the text description A1 and the original image X1 into the current image generation model.
[0138] Step 602, using the current image generation model to obtain the current generated image.
[0139] The specific inference process includes:
[0140] 1) performing pixel feature extraction on the original image X1 to obtain the pixel feature of the original image X1; performing semantic feature extraction on the original image X1 to obtain the semantic feature of the original image X1;
[0141] Specifically, in the embodiment, the second VAE encoding unit is used to perform image feature extraction on the original image X1, and the reference UNET is used to perform pixel-level feature extraction on the extracted image features to obtain pixel features of the original image X1; the semantic feature extraction unit is used to perform semantic feature extraction on the original image X1 to obtain semantic features of the original image X1.
[0142] 2) Obtain randomly generated noise features, and perform noise estimation and denoising processing on the noise features based on the pixel features of the original image X1, the semantic features of the original image X1, and the guidance of the text description A1;
[0143] Specifically, in the embodiment, the randomly generated noise features, the pixel features of the original image X1, the semantic features of the original image X1, and the text description A1 are input into the UNET unit for processing, wherein the features of each layer in the pixel features output by the reference UNET unit are fused with the output features of the same level in the UNET unit, and the semantic features are respectively fused with the output features of each level in the UNET unit, and noise estimation and denoising processing are performed based on the fused features.
[0144] 3) Perform image restoration on the result features obtained by the denoising processing to obtain the current generated image Y1.
[0145] Specifically, in the embodiment, the first VAE decoding unit is used to process the result features, and output the current generated image.
[0146] In the above inference process, the parameters of each processing unit are the latest parameters of the current image generation model.
[0147] Step 603, when it is determined that the current generated image meets the quality requirement, the original image X1, the text description A1, and the current generated image Y1 are combined to form an image-text training pair and added to a training data set for training an image generation model.
[0148] Whether the current generated image meets the quality requirement, that is, whether the effect of the current generated image meets the requirement. When it meets the requirement, the original image X1, the text description A1, and the current generated image Y1 can be combined to form an image-text training pair, that is, the current generated image Y1 is taken as a target image corresponding to the original image X1 and the text description A1, and the image-text training pair is organized and added to a training data set, so that a new target image meeting the requirement can be further generated based on the original image X1 and the text description A1, the types and quantities of target images are expanded, and the types and quantities of data in the training data set are further expanded. The expanded training data set can be used to continue training the image generation model in the future.
[0149] The above is a specific implementation of the training method of the image generation model in the present application.
[0150] Embodiment Two:
[0151] Figure 7 The specific flowchart of the image generation method in Embodiment Two is shown in the following. Figure 8 The network architecture of the trained image generation model can be obtained by the training method shown in the above. Figure 4 As shown in the above, compared with the training flow model framework of the image generation model, the trained image generation model no longer includes the first VAE encoding unit and the noise unit. As shown in the above, the specific flow of the image generation method in the embodiment includes: Figure 8 Figure 5 Figure 7 Figure 8
[0152] Step 701, input the initial image and the text description provided by the user into the trained image generation model.
[0153] The text description is used to express the change requirement for the initial image.
[0154] Step 702, perform image feature extraction on the initial image by using the second VAE encoding unit to obtain the image feature of the initial image.
[0155] Step 703, perform pixel-level feature extraction on the image feature of the initial image by using the reference UNET unit to obtain the pixel feature of the initial image.
[0156] Step 704, perform semantic feature extraction on the initial image by using the image semantic feature extraction unit to obtain the semantic feature of the initial image.
[0157] Step 705, obtain the randomly generated noise feature.
[0158] Step 706, based on the pixel feature of the initial image, the semantic feature of the initial image and the guidance of the text description, perform noise estimation and de-noising processing on the noise feature by using the UNET network to obtain the result feature.
[0159] Step 707, perform image restoration on the result feature by using the first VAE decoding unit to obtain the generated image.
[0160] Up to now, the image generation method flow in Embodiment Two is ended. Since the feature information of the original image is introduced into the image generation model, the complete target image conforming to the text description can be directly generated based on the original image, and it is not necessary to merge the generated image with the original image to obtain the target image, which is suitable for various application scenarios and initial image types including complex background.
[0161] The above is the specific implementation of the image generation method in the present application.
[0162] The application further provides a training device of an image generation model and an image generation device, which can be respectively used to implement the training method and the image generation method of the image generation model.
[0163] The basic structure of the training device of the image generation model provided in the application comprises a training data set acquisition module and a training module.
[0164] The training data set acquisition module is used to acquire a training data set of the image generation model; wherein the training data set comprises a plurality of image-text training pairs, and each image-text training pair comprises an original image, a target image and a text description used to express the change of the target image relative to the original image.
[0165] The training module is used to train the image generation model by using the training data set, and specifically comprises an original image processing submodule, a first image feature extraction unit, a noise unit, a denoising unit, an image restoration unit and a parameter updating unit.
[0166] The original image processing submodule is used to perform pixel-level feature extraction and semantic feature extraction on the original image respectively, and correspondingly obtain the pixel feature of the original image and the semantic feature of the original image.
[0167] The first image feature extraction unit is used to perform image feature extraction on the target image.
[0168] The noise unit is used to perform noise processing on the image feature extracted by the image feature extraction unit of the target image, and obtain a noise feature.
[0169] The denoising unit is used to perform noise estimation and denoising processing on the noise feature based on the pixel feature of the original image, the semantic feature of the original image and the guidance of the text description.
[0170] The image restoration unit is used to perform image restoration on the result feature obtained by the denoising unit, and obtain a generated image.
[0171] The parameter updating unit is used to compare the generated image with the target image, and update the parameters of the pixel feature extraction unit and the denoising unit based on the comparison result, so that the generated image is closer to the target image.
[0172] Optionally, the training device further comprises a data set expansion unit, which is used to input any text description and any original image corresponding thereto into the image generation model obtained by the current training, perform inference by using the image generation model to obtain a current generated image, and when it is determined that the current generated image meets the quality requirement, combine the any original image, the any text description and the current generated image to form an image-text training pair, and add the image-text training pair to the training data set for training the image generation model.
[0173] Optionally, the original image processing submodule can comprise: a second image feature extraction unit, a reference denoising unit and a semantic feature extraction unit.
[0174] The second image feature extraction unit is configured to perform image feature extraction on the original image to obtain image features of the original image.
[0175] The reference denoising unit is configured to perform pixel-level feature extraction on the image features of the original image to obtain pixel features of the original image.
[0176] The semantic feature extraction unit is configured to perform semantic feature extraction on the original image to obtain semantic features of the original image.
[0177] Optionally, the first image feature extraction unit can be a first VAE encoding unit, the second image feature extraction unit can be a second VAE encoding unit, and the image restoration unit can be a first VAE decoding unit.
[0178] The network structures of the reference denoising unit and the denoising unit are the same.
[0179] The basic structure of the image generation device provided in the present application comprises: an input module and an inference module of an image generation model.
[0180] The input module is configured to input an initial image and a text description provided by a user into the image generation model output by the training device of the image generation model.
[0181] The inference module of the image generation model is configured to use the image generation model to perform inference based on the initial image and the text description to obtain a generated image.
[0182] The inference process comprises: performing pixel-level feature extraction on the initial image in the image generation model to obtain pixel features of the initial image; performing semantic feature extraction on the initial image to obtain semantic features of the initial image; obtaining randomly generated noise features, performing noise estimation and denoising processing on the noise features based on the pixel features of the initial image, the semantic features of the initial image and the guidance of the text description; and performing image restoration on the result features obtained by the denoising processing to obtain the generated image.
[0183] In the image generation device, the inference module of the image generation model has a specific structure comprising the network structure of the image generation model, i.e., an original image processing submodule, a denoising unit and an image restoration unit.
[0184] The original image processing submodule is configured to perform pixel-level feature extraction and semantic feature extraction on the initial image respectively to obtain pixel features of the initial image and semantic features of the initial image.
[0185] The denoising unit is configured to perform noise estimation and denoising processing on the noised feature based on pixel features of the initial image, semantic features of the initial image, and guidance of the text description.
[0186] The image restoration unit is configured to perform image restoration on the result feature obtained by the denoising unit to obtain a generated image.
[0187] Optionally, the original image processing sub-module can include an image feature extraction unit of the original image, a reference denoising unit, and a semantic feature extraction unit.
[0188] The second image feature extraction unit is configured to perform image feature extraction on the initial image to obtain image features of the initial image.
[0189] The reference denoising unit is configured to perform pixel-level feature extraction on the image features of the initial image to obtain pixel features of the initial image.
[0190] The semantic feature extraction unit is configured to perform semantic feature extraction on the initial image to obtain semantic features of the initial image.
[0191] Optionally, the first image feature extraction unit can be a first VAE encoding unit, the second image feature extraction unit can be a second VAE encoding unit, and the image restoration unit can be a first VAE decoding unit.
[0192] The network structures of the reference denoising unit and the denoising unit are the same.
[0193] Optionally, the reference denoising unit and the denoising unit can both be UNET networks or both be transformer networks.
[0194] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for an image generation model, characterized in that: include: Obtaining a training dataset for training an image generation model; wherein the training dataset includes a plurality of image-text training pairs, each of the image-text training pairs including an original image, a target image, and a text description for expressing a change in the target image relative to the original image; Using the training data set to train and generate an image generation model; The training process for each image-text training pair includes: Performing pixel feature extraction and semantic feature extraction on the original image respectively, and obtaining pixel features of the original image and semantic features of the original image respectively; Extracting image features from the target image and performing noise processing on the extracted image features to obtain noise features; Based on the pixel features of the original image, the semantic features of the original image, and the guidance of the text description, the noise estimation and denoising processing are performed on the noisy features, and the image restoration is performed on the result features obtained by the denoising processing to obtain a generated image; The generated image is compared with the target image, and based on the comparison result, the parameters of the pixel feature extraction, the semantic feature extraction, the noise estimation and the denoising process are updated to make the generated image closer to the target image.
2. The method according to claim 1, characterized in that The method further comprises: Any text description and any corresponding original image are input into the currently trained image generation model, and the image generation model is used for reasoning to obtain the current generated image. When it is determined that the current generated image meets the quality requirements, the any original image, the any text description and the current generated image are combined into an image-text training pair and added to the training data set for training and generating the image generation model.
3. The method according to claim 2, characterized in that The method of using the image generation model to perform reasoning to obtain a current generated graph includes: Performing pixel feature extraction on any of the original images to obtain pixel features of the any of the original images; performing semantic feature extraction on any of the original images to obtain semantic features of the any of the original images; Obtain randomly generated noise features, perform noise estimation and denoising on the noise features based on the pixel features of any original image, the semantic features of any original image, and the guidance of any text description, and perform image restoration on the resulting features obtained by the denoising process to obtain the currently generated image.
4. The method according to claim 1, wherein The extracting pixel features from the original image includes: Performing image feature extraction on the original image to obtain image features of the original image; Perform pixel-level feature extraction on the image features of the original image to obtain pixel-level features of the original image.
5. The method according to claim 4, characterized in that The first VAE encoding unit is used to complete the image feature extraction of the target image, the second VAE encoding unit is used to complete the image feature extraction of the original image, the denoising unit is used to complete the noise estimation and denoising processing, the reference denoising unit is used to complete the pixel-level feature extraction of the image features of the original image, and the first VAE decoding unit is used to complete the image restoration; The reference denoising unit and the denoising unit have the same network structure.
6. The method according to claim 5, characterized in that The performing noise estimation and denoising processing on the noisy features by using a denoising unit includes: The features of each layer in the pixel features of the original image output by the reference denoising unit are fused with the output features of the same level in the denoising unit, and the semantic features of the original image are fused with the output features of each level in the denoising unit.
7. The method according to claim 5, characterized in that The reference denoising unit and the denoising unit are both UNET networks, or the reference denoising unit and the denoising unit are both transformer networks.
8. The method according to claim 1, characterized in that The method of obtaining the image-text training pair includes: For any text description, a pre-trained image understanding model is used to select an original image and a target image including the same subject from the original image dataset and the target image dataset corresponding to the any text description, and the selected original image, target image and the any text description are combined into an image-text training pair.
9. The method according to claim 8, characterized in that The processing of the image understanding model includes: An original image and a target image are extracted from the original image dataset and the target image dataset respectively, and a consistency comparison is performed on the two extracted images. If the result of the consistency comparison is greater than a set threshold, it is determined that the extracted original image and the target image include the same subject.
10. An image generation method, characterized in that: include: Inputting an initial image and a text description provided by a user into an image generation model trained by the image generation model training method according to any one of claims 1 to 9; wherein the text description is used to express a change requirement for the initial image; In the image generation model, pixel-level feature extraction is performed on the initial image to obtain pixel features of the initial image; semantic feature extraction is performed on the initial image to obtain semantic features of the initial image; Obtaining randomly generated noise features, and performing noise estimation and denoising on the noise features based on pixel features of the initial image, semantic features of the initial image, and guidance from the text description; The resulting features obtained by denoising are restored to obtain a generated image.
11. A training device for an image generation model, characterized in that: include: Training data set acquisition module and training module; The training data set acquisition module is used to acquire a training data set for training the image generation model; wherein the training data set includes a plurality of image-text training pairs, each of which includes an original image, a target image, and a text description for expressing the change of the target image relative to the original image; The training module is used to train and generate an image generation model using the training data set, and specifically includes an original image processing submodule, a first image feature extraction unit, a noise generation unit, a denoising unit, an image restoration unit, and a parameter updating unit; The original image processing submodule is used to perform pixel-level feature extraction and semantic feature extraction on the original image, respectively, to obtain pixel features of the original image and semantic features of the original image; The first image feature extraction unit is used to extract image features from the target image; The noise generation unit performs noise generation on the image features extracted by the image feature extraction unit of the target image to obtain noised features; The denoising unit is configured to perform noise estimation and denoising on the noisy features based on pixel features of the original image, semantic features of the original image, and guidance from the text description; The image restoration unit is used to perform image restoration on the result features obtained by the denoising unit to obtain a generated image; The parameter updating unit is used to compare the generated image with the target image, and update the parameters of the pixel feature extraction unit and the denoising unit based on the comparison result to make the generated image closer to the target image.
12. The training device according to claim 11, characterized in that The training device further includes a data set expansion unit, which is used to input any text description and any corresponding original image into the image generation model obtained by current training, use the image generation model to perform inference to obtain the current generated image, and when it is determined that the current generated image meets the quality requirements, the any original image, the any text description and the current generated image are combined into an image-text training pair and added to the training data set for training and generating the image generation model.
13. The training device according to claim 11, characterized in that The original image processing submodule includes: a second image feature extraction unit, a reference denoising unit and a semantic feature extraction unit; The second image feature extraction unit is used to extract image features from the original image to obtain image features of the original image; The reference denoising unit is used to perform pixel-level feature extraction on the image features of the original image to obtain pixel features of the original image; The semantic feature extraction unit is used to extract semantic features from the original image to obtain the semantic features of the original image.
14. The training device according to claim 13, characterized in that The target image feature extraction unit is a first VAE encoding unit, the original image feature extraction unit is a second VAE encoding unit, and the image restoration unit is a first VAE decoding unit; The reference denoising unit and the denoising unit have the same network structure.
15. An image generating device, characterized in that: include: Input module and inference module of image generation model; The input module is used to input the initial image and text description provided by the user into the image generation model output by the image generation model training device according to any one of claims 11 to 14; The inference module of the image generation model is used to use the image generation model to perform inference based on the initial image and text description to obtain a generated graph; The reasoning process includes: in the image generation model, performing pixel-level feature extraction on the initial image to obtain the pixel features of the initial image; performing semantic feature extraction on the initial image to obtain the semantic features of the initial image; obtaining randomly generated noise features, and performing noise estimation and denoising on the noise features based on the pixel features of the initial image, the semantic features of the initial image and the guidance of the text description; and performing image restoration on the result features obtained by the denoising process to obtain the generated graph.
Citation Information
Cited By
Image generation method and device and electronic equipment
CN121437671A