A training method of an image generation model, an image generation method, and related devices

By utilizing descriptive information from low-quality training data and employing an optimizer to improve image quality during the training process of the image generation model, the problem of image generation under low-quality training data is solved, and the training of a high-quality image generation model is achieved.

CN119810588BActive Publication Date: 2026-01-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411858783.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2026-01-27
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

In image generation scenarios, how can we train a high-quality image generation model with only low-quality training data?

Method used

By taking the descriptive information of low-quality training data as input, an image generation model is trained, and an optimizer is used to refine the trained images to improve their quality. The optimized images are then used as the next round of training data to continue training the model.

Benefits of technology

Under low-quality training data conditions, the image quality of the training data is gradually improved to train an image generation model that can generate high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810588B_ABST
    Figure CN119810588B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method of an image generation model, comprising: for the Nth training process, obtaining first training data of the Nth training process; taking description information of the first training image as input of the image generation model and taking the first training image as a label to train the image generation model; obtaining second description information; taking the second description information as input of the image generation model after the Nth training process to obtain a first image output by the image generation model; using an optimizer to finely process the first image to obtain a second image output by the optimizer; and for the (N+1)th training process, taking at least part of the second image and description information of the at least part of the second image as first training data of the (N+1)th training process to train the image generation model. The method can improve the image quality of the training data while training the image generation model, and can train the image generation model to generate high-quality images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a training method for an image generation model, an image generation method, a training device for an image generation model, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] In image generation scenarios, users (such as business entities with image generation needs) can use image generation models to generate desired images. Therefore, training image generation models is particularly important.

[0003] The quality of images generated by image generation models depends on the quality of the training data; that is, training the model with high-quality training data will result in an image generation model capable of generating high-quality images. However, it is usually difficult to obtain a large amount of high-quality training data. How to train a high-quality image generation model using low-quality training data has become an urgent problem to be solved. Summary of the Invention

[0004] This disclosure provides a method for training an image generation model. This method enables the training of an image generation model capable of generating high-quality images even with only low-quality training data. This disclosure also provides an image generation method corresponding to the above method, a training apparatus for the image generation model, an image generation apparatus, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] In a first aspect, this disclosure provides a method for training an image generation model, the method comprising:

[0006] For the Nth round of training, the first training data of the Nth round of training is obtained; wherein, the first training data includes a first training image and descriptive information of the first training image, and N is an integer greater than 0;

[0007] The image generation model is trained by using the description information of the first training image as input and the first training image as a label.

[0008] Obtain the second description information;

[0009] The second description information is used as the input to the image generation model after the Nth round of training to obtain the first image output by the image generation model;

[0010] The first image is refined using an optimizer to obtain a second image output by the optimizer; wherein the image quality of the second image is better than that of the first image.

[0011] For the N+1th round of training, at least a portion of the second image and at least a portion of the description information of the second image are used as the first training data for the N+1th round of training to train the image generation model.

[0012] Secondly, this disclosure provides an image generation method, the method comprising:

[0013] Obtain first description information; wherein, the first description information is used to describe the desired image;

[0014] The first description information is input into the image generation model to obtain the target image output by the image generation model; wherein the image generation model is trained by the method described in any one of claims 1 to 8.

[0015] Thirdly, this disclosure provides a training apparatus for an image generation model, the apparatus comprising:

[0016] The acquisition module is used to acquire the first training data of the Nth training round for the Nth training round; wherein the first training data includes a first training image and descriptive information of the first training image, and N is an integer greater than 0;

[0017] The training module is used to train the image generation model by taking the description information of the first training image as input and the first training image as a label.

[0018] The acquisition module is further configured to acquire second description information;

[0019] The generation module is used to take the second description information as input to the image generation model after the Nth round of training to obtain the first image output by the image generation model;

[0020] An optimization module is used to refine the first image using an optimizer to obtain a second image output by the optimizer; wherein the image quality of the second image is better than that of the first image.

[0021] The training module is further configured to, for the N+1th round of training, use at least a portion of the second image and at least a portion of the description information of the second image as the first training data for the N+1th round of training, and train the image generation model.

[0022] Fourthly, this disclosure provides an image generation apparatus, the apparatus comprising:

[0023] An acquisition module is used to acquire first description information; wherein, the first description information is used to describe the desired image;

[0024] A generation module is used to input the first description information into an image generation model to obtain a target image output by the image generation model; wherein the image generation model is trained by a training method for an image generation model as described in the first aspect or any implementation thereof.

[0025] Fifthly, this disclosure provides an electronic device including a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory to cause the electronic device to perform a training method for an image generation model as described in the first aspect or any implementation thereof, or to perform an image generation method as described in the second aspect.

[0026] In a sixth aspect, this disclosure provides a computer-readable storage medium storing instructions that instruct an electronic device to perform a training method for an image generation model as described in the first aspect or any implementation thereof, or to perform an image generation method as described in the second aspect.

[0027] In a seventh aspect, this disclosure provides a computer program product containing instructions that, when run on an electronic device, causes the electronic device to execute the training method of the image generation model described in the first aspect or any implementation thereof, or to execute the image generation method described in the second aspect.

[0028] Based on the implementation methods provided in the above aspects, this disclosure can be further combined to provide more implementation methods.

[0029] As can be seen from the above technical solutions, this disclosure has the following advantages:

[0030] This disclosure provides a training method for an image generation model. For the Nth training round, the method acquires first training data for the Nth training round, where the first training data includes a first training image and its descriptive information, where N is a positive integer. The descriptive information of the first training image is used as input to the image generation model, and the first training image is used as a label to train the image generation model. Next, second descriptive information is acquired and used as input to the image generation model after the Nth training round to obtain a first image output by the image generation model. An optimizer is then used to refine the first image to obtain a second image output by the optimizer, wherein the image quality of the second image is better than that of the first image. For the N+1th training round, at least a portion of the second image and its descriptive information are used as the first training data for the N+1th training round to train the image generation model.

[0031] In this method, when dealing with low-quality training data, the image generation model is first trained using this low-quality training data (i.e., the first training data). Then, an optimizer is used to optimize the first image output by the image generation model, improving its quality to obtain a higher-quality second image. In the next training round, the image generation model is trained again using this higher-quality second image. In this way, even with only low-quality training data, the image generation model is trained while simultaneously improving the image quality of the training data, resulting in an image generation model capable of generating high-quality images. Attached Figure Description

[0032] To more clearly illustrate the technical methods of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below.

[0033] Figure 1 A schematic flowchart illustrating a training method for an image generation model provided in an embodiment of this disclosure;

[0034] Figure 2 This is a schematic diagram of the structure of an image generation model provided in an embodiment of the present disclosure;

[0035] Figure 3 This is a schematic flowchart of an image generation method provided in an embodiment of the present disclosure;

[0036] Figure 4 This is a schematic diagram of the structure of an image generation model provided in an embodiment of the present disclosure;

[0037] Figure 5 A schematic diagram of the structure of a training device for an image generation model provided in an embodiment of this disclosure;

[0038] Figure 6 This is a schematic diagram of the structure of an image generation apparatus provided in an embodiment of the present disclosure;

[0039] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0040] The terms "first" and "second" used in the embodiments of this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0041] First, some technical terms and application scenarios involved in the embodiments of this disclosure will be introduced.

[0042] Image generation can be understood as the process of automatically generating new images. Typically, the industry utilizes image generation models for image generation; for example, an image generation model could be a diffusion model based on deep learning.

[0043] The quality of images generated by image generation models depends on the quality of the training data. In other words, training the model with high-quality training data will result in an image generation model capable of generating high-quality images. High-quality training data can be obtained through methods such as design by designers, purchase from high-quality image websites, or data cleaning. However, these methods are costly, making it difficult to acquire large amounts of high-quality training data in scenarios like model training.

[0044] In view of this, this disclosure provides a training method for an image generation model. For the Nth training round, the method acquires first training data for the Nth training round, wherein the first training data includes a first training image and descriptive information of the first training image, where N is an integer greater than 0. The descriptive information of the first training image is used as input to the image generation model, and the first training image is used as a label to train the image generation model. Next, second descriptive information is acquired and used as input to the image generation model after the Nth training round to obtain a first image output by the image generation model. An optimizer is used to refine the first image to obtain a second image output by the optimizer, wherein the image quality of the second image is better than that of the first image. For the N+1th training round, at least a portion of the second image and at least a portion of the descriptive information of the second image are used as the first training data for the N+1th training round to train the image generation model.

[0045] In this method, when dealing with low-quality training data, the image generation model is first trained using this low-quality training data (i.e., the first training data). Then, an optimizer is used to optimize the first image output by the image generation model, improving its quality to obtain a higher-quality second image. In the next training round, the image generation model is trained again using this higher-quality second image. In this way, even with only low-quality training data, the image generation model is trained while simultaneously improving the image quality of the training data, resulting in an image generation model capable of generating high-quality images.

[0046] To facilitate understanding of the technical solutions provided in the embodiments of this disclosure, a description will be given below in conjunction with the accompanying drawings. See also... Figure 1 The diagram illustrates a training method for an image generation model, which specifically includes:

[0047] S101: For the Nth round of training, obtain the first training data for the Nth round of training.

[0048] In this embodiment of the disclosure, the training of the image generation model may include multiple training processes, and the model parameters of the image generation model may be updated once in each training process.

[0049] The Nth round of training can be understood as any training process in the training of the image generation model, where N can be an integer greater than 0. The first training data in the Nth round of training can be understood as the training data used to train the image generation model in the Nth round of training.

[0050] Specifically, the first training data may include a first training image and descriptive information of the first training image. The descriptive information of the first training image can be used to describe the image content of the first training image. In this embodiment of the disclosure, the descriptive information of the first training image can be described in natural language, representing the natural language content; that is, the descriptive information of the first training image can describe the image content of the first training image in the form of natural language.

[0051] Thus, by using the first training image and its descriptive information as the first training data, the image generation model can generate images based on text.

[0052] S102: Use the description information of the first training image as input to the image generation model, and use the first training image as a label to train the image generation model.

[0053] In this embodiment of the disclosure, the training of the image generation model can be supervised training, and the label can be understood as the correct output result in supervised training.

[0054] In practice, the image generation model is input with the descriptive information of the first training image. The image generation model generates an output image that matches the descriptive information of the first training image. Then, based on the output image of the image generation model and the first training image as the label, the loss function value is calculated. This loss function value can characterize the degree of difference between the output image of the image generation model and the first training image. The model parameters of the image generation model are adjusted to minimize the loss function value, thus completing the Nth round of training for the image generation model.

[0055] S103: Obtain the second description information.

[0056] After completing the Nth round of training for the image generation model, the image generation model trained in the Nth round is used to generate images. In other words, the second descriptive information can be used as input to the image generation model after the Nth round of training.

[0057] Similar to the descriptive information of the first training image, the second descriptive information can be described in natural language, representing the natural language content. That is, the second descriptive information can describe the image content of the desired image in the form of natural language.

[0058] This disclosure does not limit the source of the second description information. For example, the second description information can be extracted from an existing image library, or it can be written by a user (e.g., an annotator) based on the desired image.

[0059] S104: Use the second description information as input to the image generation model after the Nth round of training to obtain the first image output by the image generation model.

[0060] like Figure 2 As shown, the input to the image generation model can include an initial noisy image and second descriptive information. The initial noisy image can be understood as an image composed of initial noise (e.g., random numbers). Typically, the initial noisy image has the same size as the first image. The image generation model can use the initial noisy image as a reference and generate an image based on it.

[0061] It should be noted that the image generation model can also include more modules. For example, the image generation model can also include a variational autoencoder module, which includes an encoder and a decoder. The encoder is used to downsample the initial noisy image (e.g., compressing the initial 512×512 noisy image into a 16×16 image), and the decoder is used to upsample the image output by the image generation model (e.g., restoring the 512×512 first image from the 16×16 image output by the image generation model).

[0062] Since the image generation model has undergone N rounds of training, after inputting the second descriptive information into the image generation model, the image generation model can input a first image that satisfies the second descriptive information.

[0063] As mentioned earlier, since it is difficult to obtain a large amount of high-quality first training data during the training process, the image quality of the first image generated by the image generation model after training the image generation model in the Nth round of training using the first training data may be poor. For example, the first image may have defects such as incomplete figures or erasure marks, resulting in poor image quality.

[0064] S105: Use the optimizer to refine the first image and obtain the second image output by the optimizer.

[0065] Continue as Figure 2As shown in this embodiment, for the first image generated by the image generation model after each round of training, an optimizer is used to optimize the first image to obtain a second image. The image quality of the second image can be better than that of the first image. For example, the details of the second image are clearer, the objects in the second image are more complete, and the tone and style of the second image are more in line with the image generation requirements of the business scenario.

[0066] S106: For the N+1th round of training, at least a portion of the second image and the description information of at least a portion of the second image are used as the first training data for the N+1th round of training to train the image generation model.

[0067] Continue as Figure 2 As shown in this embodiment, an optimizer is used to optimize the first image output by the image generation model after the Nth training round. The optimized second image is then used as the first training data for the N+1th training round to continue training the image generation model. In this way, the quality of the first training data in each round is improved compared to the first training data in the previous round, and the training effect of the image generation model gradually improves.

[0068] In this method, when dealing with low-quality training data, the image generation model is first trained using this low-quality training data (i.e., the first training data). Then, an optimizer is used to optimize the first image output by the image generation model, improving its quality to obtain a higher-quality second image. In the next training round, the image generation model is trained again using this higher-quality second image. In this way, even with only low-quality training data, the image generation model is trained while simultaneously improving the image quality of the training data, resulting in an image generation model capable of generating high-quality images.

[0069] The training process of the image generation model provided in the embodiments of this disclosure has been described above. The specific contents involved in the training process will be introduced below.

[0070] In this embodiment of the disclosure, the first training data can be selected based on actual image generation requirements. Specifically, the first training data can be related to at least one of the following: target service, target content, or target style.

[0071] The target business can be understood as a specific business scenario with image generation requirements. For example, in business A, it is necessary to use an image generation model to generate template images for advertising. In this case, the target business is business A. The first training data may include a first training image that meets the requirements of the template image and descriptive information of the first training image.

[0072] The target content can be understood as the specific object in the image. For example, if a user needs to use an image generation model to generate an image related to a cartoon character B, then the cartoon character B is the target content. The first training data may include a first training image containing the cartoon character B and descriptive information of the first training image.

[0073] The target style can be understood as a specific image style. For example, if a user needs to use an image generation model to generate an image in the style of an oil painting, the target style is the oil painting style. The first training data can include the first training image in the style of an oil painting and the descriptive information of the first training image.

[0074] Thus, by selecting the first training data based on the actual image generation needs, the image generation model can learn to meet specific business, content, or style requirements, thereby improving the matching degree between the image generation model and the actual image generation needs. Furthermore, when image generation needs involve the aforementioned target business, content, or style, it is often difficult to obtain a large amount of high-quality training data. Applying the training method of the image generation model provided in this disclosure can yield an image generation model with good inference performance.

[0075] Similarly, the desired image described by the second descriptive information can also be related to at least one of the target business, target content, and target style. That is, after training the image generation model to generate images under a specific business, specific content, or specific style using the first training data, the second descriptive information under the same business, same content, or same style can be obtained. On the one hand, this verifies the training effect of the image generation model after the Nth round of training, and on the other hand, it improves the quality of training data in subsequent training processes.

[0076] Since both the first training data and the second descriptive information are related to the same business, content, or style, the second image output by the optimizer can also be related to the same business, content, or style, thereby improving the image quality of images related to the same business, content, or style.

[0077] Model training is typically configured with termination conditions, and the model training process ends when these conditions are met. In this embodiment of the disclosure, after the Mth training round, if the image quality of the first image output by the image generation model meets a first preset condition, the training of the image generation model ends. Here, M is an integer greater than 0.

[0078] In other words, the termination condition in this embodiment can be "the image quality of the first image meets the first set condition". Different methods of determining image quality can correspond to different first set conditions. In some embodiments, a score representing the image quality of the first image is determined by a multimodal scoring model. In this case, the first set condition can be a score threshold. When the score representing the image quality of the first image is greater than the score threshold, it is determined that the image quality of the first image meets the first set condition. In other embodiments, a rating representing the image quality of the first image is determined by manual annotation. In this case, the first set condition can be a set level. When the rating representing the image quality of the first image is greater than the set level, it is determined that the image quality of the first image meets the first set condition.

[0079] Thus, when the image generation model can generate a first image of high quality without relying on the optimizer, the training of the image generation model is terminated, ensuring that the image generation model has good reasoning ability.

[0080] The following describes in detail the process by which the optimizer optimizes the first image. Specifically, the first image is encoded to obtain its latent space code. The second description information and the latent space code of the first image are used as input to the optimizer to obtain the latent space code of the second image output by the optimizer. The latent space code of the second image is then decoded to obtain the second image.

[0081] The process of obtaining the latent space encoding of the first image can be understood as encoding the image data of the first image into a low-dimensional latent space. Through latent space encoding, the image generation model can improve the quality and resolution of the first image while maintaining computational efficiency.

[0082] In this embodiment, the core principle of the optimizer is to optimize the image by performing denoising processing. The denoising process can be divided into multiple steps, in which the noise in the first image is gradually reduced and the image quality is improved, resulting in a more refined second image. Thus, for the first image output by the image generation model, the optimizer performs denoising processing on the first image, resulting in a second image with better image quality output by the optimizer.

[0083] This disclosure supports image optimization using an optimizer in different ways. In some embodiments, the second description information and the latent space code of the first image are used as inputs to the optimizer, causing the optimizer to perform a first denoising process on the latent space code of the first image to obtain the latent space code of the second image output by the optimizer.

[0084] In other embodiments, the latent space code of the first image is denoised, and the second description information and the latent space code of the first image after denoising are used as inputs to the optimizer, so that the optimizer performs a second denoising process on the latent space code of the first image after denoising, and obtains the latent space code of the second image output by the optimizer.

[0085] In other words, the optimizer can directly denoise the first image output by the image generation model, or it can first add noise to the first image output by the image generation model, and then use the optimizer to denoise the denoised first image. Thus, in the method of directly denoising the first image output by the image generation model, the optimizer reduces noise in the first image, making the second image, compared to the first image, have a blurred background, a smoother background, and a stronger focus on the subject. In the method of first adding noise to the first image output by the image generation model, and then denoising, the addition of noise adds more content to the first image, making the second image, compared to the first image, have more image details and richer image information, with more textures and lines. Simultaneously, since the optimizer's input also includes second descriptive information, the second image, compared to the first image, is a better match to the second descriptive information and better meets the user's image generation needs.

[0086] In some possible implementations, the degree of noise reduction in the first and second denoising processes can differ. For example, the first denoising process can be used to perform the final n% of noise reduction, where n can be a real number greater than 0 and less than 100, and the second denoising process can be used to perform 100% noise reduction.

[0087] After refining the first image using an optimizer to obtain a second image with better image quality, at least a portion of the second images can be used as the first training images for the next round of training (i.e., the N+1th round of training). Specifically, second target images that meet the second set conditions are selected from the second images. For the N+1th round of training, the descriptive information of the second target images is used as the input to the image generation model, and the second target images are used as labels to train the image generation model.

[0088] The second setting condition can be understood as a selection criterion for choosing the second target image from the second image. That is, the second image that meets the second setting condition is determined as the second target image, and the image generation model is trained for the N+1th round using the description information of the second target image.

[0089] The embodiments disclosed herein do not limit the second setting conditions. For example, the second setting conditions can be configured from different dimensions (such as the clarity of details, the completeness of objects, the degree of matching with the requirements of image generation, etc.). When the second image meets the second setting conditions in all dimensions, the second image is determined as the second target image.

[0090] Thus, the optimized second image is filtered, and the high-quality second target image is determined as the first training data in the N+1th round of training, thereby improving the quality of the training data in the N+1th round of training.

[0091] Furthermore, considering that the optimizer can process the first image in different ways, such as directly denoising the first image output by the image generation model, or adding noise to the first image output by the image generation model first and then denoising, different processing methods can correspond to different image styles, such as a style that blurs the background or a style that enriches the details. Therefore, in this embodiment of the present disclosure, the description information of the second target image may include an image style identifier. The image style identifier can be used to identify the image style of the second target image. The image style of the second target image is determined based on the processing method of the optimizer on the first image.

[0092] In other words, by adding image style labels to the first training data in the (N+1)th round of training, the image generation model can learn different image styles. Thus, after training the image generation model, during actual inference, if the model needs to generate images of a certain style, the corresponding image style label can be input, enabling the model to generate images of different styles.

[0093] Based on the image generation model training method provided above, this disclosure also provides an image generation method. See [link to previous document]. Figure 3 The diagram shows a flowchart of an image generation method, which specifically includes:

[0094] S301: Obtain the first description information.

[0095] The first descriptive information can be used to describe the desired image; in other words, the first descriptive information can characterize the user's image generation needs.

[0096] This disclosure does not limit the source of the first description information. For example, users can extract the first description information from existing advertising materials, or in a template generation scenario, users can extract the first description information from video materials.

[0097] S302: Input the first description information into the image generation model to obtain the target image output by the image generation model.

[0098] The image generation model can be trained using the training method for image generation models provided above. Specifically, such as... Figure 4 As shown, unlike the training process, during the inference process, there is no need to add an optimizer. By inputting the first descriptive information into the image generation model, the image generation model can generate a target image that matches the first descriptive information based on the original noisy image, thus achieving image generation.

[0099] In this method, as the quality of the training data gradually improves during the training process of the image generation model, the image generation model has good image generation capabilities and can generate high-quality images that meet the image generation requirements (i.e., the first descriptive information).

[0100] The above text combined Figures 1 to 4 The training method and image generation method of the image generation model provided in the embodiments of this disclosure have been described in detail. The apparatus and device provided in the embodiments of this disclosure will be described below with reference to the accompanying drawings.

[0101] See Figure 5 The schematic diagram of the training device for the image generation model shown indicates that the device 50 includes:

[0102] The acquisition module 501 is used to acquire the first training data of the Nth training round for the Nth training round; wherein the first training data includes a first training image and descriptive information of the first training image, and N is an integer greater than 0;

[0103] The training module 502 is used to train the image generation model by taking the description information of the first training image as input and the first training image as a label.

[0104] The acquisition module 501 is further configured to acquire second description information;

[0105] The generation module 503 is used to take the second description information as the input of the image generation model after the Nth round of training to obtain the first image output by the image generation model;

[0106] Optimization module 504 is used to refine the first image using an optimizer to obtain a second image output by the optimizer; wherein the image quality of the second image is better than that of the first image;

[0107] The training module 502 is further configured to, for the N+1th round of training, use at least a portion of the second image and at least a portion of the description information of the second image as the first training data for the N+1th round of training, and train the image generation model.

[0108] In some possible implementations, the training module 502 is further configured to:

[0109] After the Mth round of training, if the image quality of the first image output by the image generation model meets the first set condition, the training of the image generation model ends, where M is an integer greater than 0.

[0110] In some possible implementations, the first training data is related to at least one of the following: target business, target content, or target style.

[0111] In some possible implementations, the optimization module 504 is specifically used for:

[0112] The first image is encoded to obtain the latent space code of the first image;

[0113] The second description information and the latent space code of the first image are used as inputs to the optimizer to obtain the latent space code of the second image output by the optimizer.

[0114] The latent space encoding of the second image is decoded to obtain the second image.

[0115] In some possible implementations, the optimization module 504 is specifically used for:

[0116] The second description information and the latent space code of the first image are used as inputs to the optimizer, so that the optimizer performs a first denoising process on the latent space code of the first image to obtain the latent space code of the second image output by the optimizer.

[0117] In some possible implementations, the optimization module 504 is specifically used for:

[0118] The latent space encoding of the first image is subjected to noise processing;

[0119] The second description information and the latent space code of the first image after noise processing are used as inputs to the optimizer, so that the optimizer performs a second denoising process on the latent space code of the first image after noise processing to obtain the latent space code of the second image output by the optimizer.

[0120] In some possible implementations, the training module 502 is specifically used for:

[0121] Select a second target image from the second image whose image quality meets the second set conditions;

[0122] For the N+1th round of training, the description information of the second target image is used as the input of the image generation model, and the second target image is used as the label to train the image generation model.

[0123] In some possible implementations, the description information of the second target image includes an image style identifier, which is used to identify the image style of the second target image, and the image style of the second target image is determined based on the processing method of the first image by the optimizer.

[0124] The image generation model training apparatus 50 according to the embodiments of this disclosure can correspond to performing the methods described in the embodiments of this disclosure, and the above and other operations and / or functions of each module / unit of the image generation model training apparatus 50 are respectively for implementing Figure 1 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.

[0125] See Figure 6 The schematic diagram of the image generation apparatus shown shows that the apparatus 60 includes:

[0126] The acquisition module 601 is used to acquire first description information; wherein, the first description information is used to describe the desired image;

[0127] The generation module 602 is used to input the first description information into the image generation model to obtain the target image output by the image generation model; wherein the image generation model is trained by the aforementioned image generation model training method.

[0128] The image generation apparatus 60 according to the embodiments of this disclosure can correspond to performing the methods described in the embodiments of this disclosure, and the above and other operations and / or functions of each module / unit of the image generation apparatus 60 are respectively for implementing Figure 3 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.

[0129] This disclosure also provides an electronic device. This electronic device is specifically used to implement, as described above. Figure 5 The training device 50 of the image generation model in the illustrated embodiment functions as follows, or is used to achieve, for example... Figure 6 The image generation device 60 in the illustrated embodiment has the following functions.

[0130] Figure 7 A structural schematic diagram of an electronic device 700 is provided, such as... Figure 7As shown, the electronic device 700 includes a bus 701, a processor 702, a communication interface 703, and a memory 704. The processor 702, the memory 704, and the communication interface 703 communicate with each other via the bus 701.

[0131] The 701 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0132] The processor 702 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0133] Communication interface 703 is used for external communication. For example, communication interface 703 can be used to communicate with a terminal.

[0134] Memory 404 may include volatile memory, such as random access memory (RAM). Memory 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0135] The memory 704 stores executable code, and the processor 702 executes the executable code to perform the training method or image generation method of the aforementioned image generation model.

[0136] Specifically, in achieving Figure 5 or Figure 6 In the case of the illustrated embodiment, and Figure 5 The training device 50 for the image generation model described in the embodiments or Figure 6 When the modules or units of the image generation apparatus 60 described in the embodiment are implemented by software, the following steps are performed: Figure 5 or Figure 6 The software or program code required for the functions of each module / unit can be partially or wholly stored in the memory 704. The processor 702 executes the program code corresponding to each unit stored in the memory 704 to execute the training method or image generation method of the aforementioned image generation model.

[0137] This disclosure also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the training method for the image generation model applied to the training apparatus 50 of the image generation model, or to execute the image generation method applied to the image generation apparatus 60.

[0138] This disclosure also provides a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this disclosure are generated.

[0139] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0140] When the computer program product is executed by a computer, the computer executes any method of the aforementioned image generation model training method or any method of the aforementioned image generation method. The computer program product can be a software installation package; when it is necessary to use any method of the aforementioned image generation model training method or any method of the aforementioned image generation method, the computer program product can be downloaded and executed on the computer.

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0142] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units / modules do not necessarily limit the specific unit itself.

[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0144] In the context of embodiments of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0145] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.

[0146] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0147] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0148] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0149] The above description of the disclosed embodiments enables those skilled in the art to make or use this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A training method for an image generation model, characterized in that, The method includes: For the Nth round of training, the first training data of the Nth round of training is obtained; wherein, the first training data includes a first training image and descriptive information of the first training image, and N is an integer greater than 0; The image generation model is trained by using the description information of the first training image as input and the first training image as a label. Obtain second descriptive information, which describes the image content of the desired image in natural language. The second description information is used as the input to the image generation model after the Nth round of training to obtain the first image output by the image generation model; The first image is refined using an optimizer to obtain a second image output by the optimizer; wherein the image quality of the second image is better than that of the first image. For the N+1th training round, at least a portion of the second image and at least a portion of the description information of the second image are used as the first training data for the N+1th training round to train the image generation model; After the Mth training round, if the image quality of the first image output by the image generation model meets the first set condition, the training of the image generation model ends, where M is an integer greater than 0. The step of using at least a portion of the second image and at least a portion of the description information of the second image as the first training data in the N+1th round of training to train the image generation model includes: Select a second target image from the second image whose image quality meets the second set conditions; For the N+1th round of training, the description information of the second target image is used as the input of the image generation model, and the second target image is used as the label to train the image generation model.

2. The method according to claim 1, characterized in that, The first training data is related to at least one of the following: target business, target content, or target style.

3. The method according to claim 1, characterized in that, The step of refining the first image using an optimizer to obtain the second image output by the optimizer includes: The first image is encoded to obtain the latent space code of the first image; The second description information and the latent space code of the first image are used as inputs to the optimizer to obtain the latent space code of the second image output by the optimizer. The latent space encoding of the second image is decoded to obtain the second image.

4. The method according to claim 3, characterized in that, The step of using the second description information and the latent space encoding of the first image as input to the optimizer to obtain the latent space encoding of the second image output by the optimizer includes: The second description information and the latent space code of the first image are used as inputs to the optimizer, so that the optimizer performs a first denoising process on the latent space code of the first image to obtain the latent space code of the second image output by the optimizer.

5. The method according to claim 3, characterized in that, The step of using the second description information and the latent space encoding of the first image as input to the optimizer to obtain the latent space encoding of the second image output by the optimizer includes: The latent space encoding of the first image is subjected to noise processing; The second description information and the latent space code of the first image after noise processing are used as inputs to the optimizer, so that the optimizer performs a second denoising process on the latent space code of the first image after noise processing to obtain the latent space code of the second image output by the optimizer.

6. The method according to claim 1, characterized in that, The description information of the second target image includes an image style identifier, which is used to identify the image style of the second target image. The image style of the second target image is determined based on the processing method of the first image by the optimizer.

7. An image generation method, characterized in that, The method includes: Obtain first description information; wherein, the first description information is used to describe the desired image; The first description information is input into the image generation model to obtain the target image output by the image generation model; wherein the image generation model is trained by the method described in any one of claims 1 to 6.

8. A training device for an image generation model, characterized in that, The device includes: The acquisition module is used to acquire the first training data of the Nth training round for the Nth training round; wherein the first training data includes a first training image and descriptive information of the first training image, and N is an integer greater than 0; The training module is used to train the image generation model by taking the description information of the first training image as input and the first training image as a label. The acquisition module is further configured to acquire second descriptive information, wherein the second descriptive information describes the image content of the desired image in the form of natural language; The generation module is used to take the second description information as input to the image generation model after the Nth round of training to obtain the first image output by the image generation model; An optimization module is used to refine the first image using an optimizer to obtain a second image output by the optimizer; wherein the image quality of the second image is better than that of the first image. The training module is further configured to, for the N+1th round of training, use at least a portion of the second image and at least a portion of the description information of the second image as the first training data for the N+1th round of training, and train the image generation model. The training module is also used for: After the Mth training round, if the image quality of the first image output by the image generation model meets the first set condition, the training of the image generation model ends, where M is an integer greater than 0. The training module is specifically used for: Select a second target image from the second image whose image quality meets the second set conditions; For the N+1th round of training, the description information of the second target image is used as the input of the image generation model, and the second target image is used as the label to train the image generation model.

9. An image generation apparatus, characterized in that, The device includes: An acquisition module is used to acquire first description information; wherein, the first description information is used to describe the desired image; A generation module is used to input the first description information into an image generation model to obtain a target image output by the image generation model; wherein the image generation model is trained by the method described in any one of claims 1 to 6.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory; The processor is configured to execute instructions stored in the memory, causing the electronic device to perform the method as claimed in any one of claims 1 to 6 or 7.

11. A computer-readable storage medium, characterized in that, Includes instructions that instruct an electronic device to perform the method as claimed in any one of claims 1 to 6 or 7.

12. A computer program product, characterized in that, The computer program product includes computer-readable instructions for implementing the method of any one of claims 1 to 6 or claim 7.

Citation Information

Patent Citations

  • Text-to-image generation method and device, storage medium and terminal

    CN115700519A

  • Video resolution improving system and method based on pre-training video generation model

    CN119048356A