Pentograph model training method, image generation method and corresponding device
By introducing the differential optimization goals of positive and negative samples in the training of literary and biographical graph model, and using the diffusion model to train the literary and biographical graph model, the problem of insufficient user preference satisfaction in the prior art is solved, and more efficient user preference image generation is achieved.
Patent Information
- Application Number
- CN202510180360.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
The existing literary and biographical graphics technology based on diffusion model still needs to be improved in satisfying user preferences.
By obtaining training data including positive and negative samples, the diffusion model is used to train the literary graph model, and the model parameters are optimized to minimize the difference from the positive samples and maximize the difference from the negative samples, thereby improving the user preference compliance of image generation.
It realizes the generation of images that are more in line with user preferences, reduces the ability to generate images that are not in line with user preferences, and does not require additional training of auxiliary models, simplifies the training process and reduces costs.
Smart Images

Figure CN120125934A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a training method for a text-to-image generation model, an image generation method, and corresponding devices. Background Art
[0002] Text-to-Image Generation is a technology for generating images from text. The core of this technology lies in combining deep learning and natural language processing technologies, and using a deep learning model to analyze the input text description and generate corresponding images. Among them, the diffusion model has been widely used in the text-to-image technology in the field of deep learning. The principle of the diffusion model is to gradually add noise to the data and then gradually remove the noise to achieve the generation and restoration of the data. In practical applications, by introducing additional guidance information such as text descriptions, the diffusion model is guided to generate images that match the text descriptions.
[0003] However, the existing text-to-image technology based on the diffusion model still needs to be improved in meeting user preferences. Summary of the Invention
[0004] In view of this, the present application provides a training method for a text-to-image generation model, an image generation method, and corresponding devices, so that the generated images are more in line with user preferences.
[0005] The present application provides the following solutions:
[0006] According to a first aspect, a training method for a text-to-image generation model is provided. The method includes: obtaining training data including a plurality of sample pairs, where a sample pair includes a text description sample, a positive sample corresponding to the text description sample, and a negative sample, the positive sample is an image preferred by a user, and the negative sample is an image not preferred by the user; training a text-to-image generation model implemented based on a diffusion model using the training data, where the training includes: using the text description sample and a noise image obtained by adding noise to the positive sample as input data, and using the text-to-image generation model to perform denoising processing on the noise image in the input data to obtain a first image; and using the text description sample and a noise image obtained by adding noise to the negative sample as input data, and using the text-to-image generation model to perform denoising processing on the noise image in the input data to obtain a second image; optimizing model parameters of the text-to-image generation model based on a training objective, where the training objective includes: minimizing the difference between the first image and the positive sample corresponding to the text description sample, and maximizing the difference between the second image and the negative sample corresponding to the text description sample.
[0007] According to a second aspect, there is provided an image generation method, the method comprising: obtaining a text description; inputting the text description and a noise image into a text-to-image model to obtain an image corresponding to the text description, wherein the text-to-image model is pre-trained by the method described in the first aspect above.
[0008] According to a third aspect, there is provided an image generation apparatus, the apparatus comprising:
[0009] a description acquisition unit configured to acquire a text description;
[0010] an image generation unit configured to input the text description and a noise image into a text-to-image model to obtain an image corresponding to the text description, wherein the text-to-image model is pre-trained by the method described in the first aspect above.
[0011] According to a fourth aspect, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of the first aspect or the second aspect above.
[0012] According to the specific embodiments provided in the present application, the following technical effects are disclosed in the present application:
[0013] The present application trains a text-to-image model based on a diffusion model, uses the images preferred by the user and the images not preferred by the user as positive samples and negative samples respectively, and uses the text description samples and the noise-added positive / negative samples as the input data of the text-to-image model during the training process to train the learning ability of the text-to-image model for user preferences. During the process of optimizing the parameters of the text-to-image model based on the training objective, the training objective is used to improve the generation ability of the text-to-image model for the images preferred by the user and reduce the generation ability of the text-to-image model for the images not preferred by the user, so that the trained text-to-image model can generate images more in line with user preferences. Moreover, this training method does not require additional training of other auxiliary models, has a simpler implementation and lower training costs. Description of the Drawings
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0015] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;
[0016] Figure 2 is a flowchart of the training method of the text-to-image model provided by the embodiments of the present application;
[0017] Figure 3 Schematic diagrams of the forward process and the reverse process in a traditional diffusion model;
[0018] Figure 4 Schematic principle diagram of the text-to-image model provided by an embodiment of the present application;
[0019] Figure 5 Flowchart of the image generation method provided by an embodiment of the present application;
[0020] Figure 6 Schematic diagram of the training device of the text-to-image model provided by an embodiment of the present application;
[0021] Figure 7 Schematic diagram of the image generation device provided by an embodiment of the present application;
[0022] Figure 8 Schematic block diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the present application.
[0024] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should be understood that the term " / and / " used herein is only a description of the associated relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0026] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0027] Currently, there are already some methods aimed at meeting users' preference correction for text-to-image technology, such as through an additionally trained reward model. During the process of training a text-to-image model, the images generated by the text-to-image model are used as the input to the reward model, which can output a reward value for the input images. This reward value reflects the degree of users' preference for the images. Among them, maximizing the reward value is used as the training objective of the model, so that the trained text-to-image model better conforms to users' preferences. However, this method requires training an additional reward model, introducing new training overhead and debugging costs, and the implementation is complex.
[0028] In view of this, the present application provides a new idea. To facilitate the understanding of the present application, first, the system architecture on which the present application is based will be described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown, as Figure 1 shown in, this system architecture may include: a user device, and an image generation device and a model training device provided on the server side.
[0029] The user device can be any device with text or voice input and image display functions, and may include but are not limited to: intelligent mobile terminals, smart home devices, wearable devices, PCs (Personal Computers), etc. Among them, intelligent mobile devices may include, for example, mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), Internet cars, etc. Smart home devices may include smart TVs, smart refrigerators, etc. Wearable devices may include, for example, smart watches, smart glasses, virtual reality devices, augmented reality devices, mixed reality devices (i.e., devices that can support virtual reality and augmented reality), etc.
[0030] The server side can be set as an independent server, or can be set in a server group, or can also be set in a cloud server. A cloud server, also known as a cloud computing server or a cloud host, is a host product in the cloud computing service system, which is used to solve the defects of large management difficulty and weak service scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0031] As one of the feasible ways, the user can input a text description through the user device or input a voice description through the user device. The user device sends the text description or voice description to the server side. If the server side receives a voice description, it can convert the voice description into a text description. The image generation device on the server side uses the text description to perform denoising processing on the input noisy image to generate an image, and returns the generated image to the user device through the network, and the user device presents the image to the user. Among them, when the image generation device generates an image, it involves the use of a text-to-image model, and the text-to-image model can be pre-trained by the model training device in the manner of the embodiments of the present application. For specific references, see the relevant records in the subsequent embodiments.
[0032] In addition to the above method, the model training device and / or the image generation device can also be set in a computer terminal with strong computing power.
[0033] It should be understood that Figure 1 the user device, model training device, image generation device, and text-to-image model in
[0034] Figure 2 is only illustrative for the method flow chart of training the text-to-image model provided by the embodiments of the present application. This method can be executed by Figure 1 the server side in the system shown in Figure 2 As shown in
[0035] Step 202: Obtain training data including multiple sample pairs. The sample pair includes a text description sample, a positive sample corresponding to the text description sample, and a negative sample. The positive sample is an image preferred by the user, and the negative sample is an image not preferred by the user.
[0036] Step 204: Use the training data to train a text-to-image model implemented based on a diffusion model. The training includes: inputting the text description sample and the noisy image obtained by adding noise to the positive sample into the text-to-image model, and obtaining the first image obtained by the text-to-image model after denoising the input noisy image; and inputting the text description sample and the noisy image obtained by adding noise to the negative sample into the text-to-image model, and obtaining the second image obtained by the text-to-image model after denoising the input noisy image; optimizing the model parameters of the text-to-image model based on the training objective, and the training objective includes: minimizing the difference between the first image and the positive sample corresponding to the text description sample, and maximizing the difference between the second image and the negative sample corresponding to the text description sample.
[0037] As can be seen from the above process, the present application trains an image generation model from text based on a diffusion model, using the images preferred by the user and the images not preferred by the user as positive samples and negative samples respectively, and using the text description samples and the noisy positive / negative samples as the input data for the image generation model from text during the training process to train the learning ability of the image generation model from text for user preferences. During the process of optimizing the parameters of the image generation model from text based on the training objective, the generation ability of the image generation model from text for the images preferred by the user is improved and the generation ability of the image generation model from text for the images not preferred by the user is reduced through the training objective, so that the trained image generation model from text can generate images more in line with user preferences. Moreover, this training method does not require additional training of other auxiliary models, with a simpler implementation and lower training cost.
[0038] The following will describe in detail each step in the above process and the further effects that can be produced in conjunction with embodiments.
[0039] First, the above step 202, i.e., "obtaining training data including multiple sample pairs", will be described in detail in conjunction with embodiments.
[0040] In the embodiments of the present application, the training data is in the form of sample pairs, and each sample pair includes a text description sample, a positive sample and a negative sample corresponding to the text description sample. The positive sample is an image preferred by the user, and the negative sample is an image not preferred by the user.
[0041] Among them, the text description sample in the embodiments of the present application can be understood as a text sample used to describe the characteristics of an image. Optionally, the characteristics of the image may include the content, layout, color, texture, shape, etc. of the image. For example, the text description sample can be "Please generate an image containing a cat", or "Please generate an image in which a black cat is sitting on a white stool", and so on.
[0042] As one possible implementation, user preferences can be distinguished from machine preferences. That is, user preferences can be embodied as human preferences, i.e., images that humans prefer more. When generating positive and negative samples, the text description sample can be input into an existing image generation model from text, and the generated images of the image generation model from text are labeled. If the image belongs to the images preferred by humans, it is labeled as a positive sample. If the generated image of the image generation model from text has obvious traces of machine generation, for example, it is clearly an image generated by a machine (i.e., a model), it is labeled as a negative sample. Such positive and negative samples can enable the trained image generation model from text to reduce the characteristics of machine generation and have more characteristics of human preferences.
[0043] As another achievable manner, user preferences can also be embodied as personalized preferences of specific users, specific types of users, users in specific regions, etc. For example, it can be the preference of a specific user for image features, or the preference of specific types of users such as children and white-collar workers for image features, or the preference of users in specific regions such as Central Asia and Europe for image features, and so on. Such positive and negative samples can enable the trained text-to-image model to better understand and reflect the personalized preferences of users.
[0044] In the embodiments of the present application, training data of multiple sample pairs for training a text-to-image model is constructed using text description samples and corresponding positive and negative samples of the text description samples. During the training process, the positive samples are used to guide the text-to-image model to learn the key features in the images to generate images that meet user preferences; while introducing negative samples enables the text-to-image model to more clearly distinguish between images that meet user preferences (positive samples) and images that do not meet user preferences (negative samples), which helps the text-to-image model avoid generating images that do not meet user preferences (i.e., reduce the output probability of negative samples). By combining the use of positive and negative samples during the training process, a more accurate text-to-image model that better meets user preferences can be trained.
[0045] The following describes in detail the above step 204, that is, "training a text-to-image model implemented based on a diffusion model using training data" in conjunction with embodiments.
[0046] In the embodiments of the present application, the training process of the text-to-image model includes: using the text description sample and the noisy image obtained by adding noise to the positive sample as input data, obtaining the first image obtained by the text-to-image model after denoising the noisy image in the input data, and using the text description sample and the noisy image obtained by adding noise to the negative sample as input data, obtaining the second image obtained by the text-to-image model after denoising the noisy image in the input data; optimizing the model parameters of the text-to-image model based on the training objective, where the training objective includes: minimizing the difference between the first image and the positive sample corresponding to the text description sample, and maximizing the difference between the second image and the negative sample corresponding to the text description sample.
[0047] Here, the above noise addition process can refer to the process of adding Gaussian noise to the positive sample and the negative sample respectively to obtain the noisy image corresponding to the positive sample and the noisy image corresponding to the negative sample.
[0048] For the text-to-image model, after inputting the text description sample and the noisy image into the text-to-image model, the text-to-image model can, under the guidance of the text description sample, perform denoising processing on the noisy image to obtain the generated image. For the sake of convenient description and distinction, the image generated for the noisy image corresponding to the positive sample is called the first image, and the image generated for the noisy image corresponding to the negative sample is called the second image.
[0049] When training a text graph model, it is essentially a process of optimizing the model parameters of the initial text graph model. The initial text graph model can be any existing text graph model. When optimizing the model parameters, a loss function can be constructed based on the training objective, and the values of the model parameters can be updated using methods such as gradient descent using the values of the loss function in each round until a preset training end condition is reached. The training end condition can be, for example, that the value of the loss function is less than a preset threshold, the number of training iterations reaches a preset round number threshold, and so on.
[0050] As one of the feasible ways, we can design the loss function L1 for the above training objective, such as Figure 4 As shown in , the L1 aims to shorten the distance between the first image and the positive sample and to increase the distance between the second image and the negative sample. In essence, for the same text description, the probability of the text graph model generating positive samples (i.e., images preferred by the user) is increased and the probability of the text graph model generating negative samples (i.e., images not preferred by the user) is reduced.
[0051] Since the Wensheng graph model in the embodiment of the present application is implemented based on the diffusion model, the principle of the diffusion model is briefly introduced below: The training process of the diffusion model usually includes a forward process and a reverse process. The idea of the forward process comes from the Markov chain, which means that from time step 0 to time step T, the image X is gradually 0 Add Gaussian noise, making it increasingly blurry and random, until you get the noisy image X T The reverse process is to process the noise image X from time step T to time step 0. T Denoising is performed to generate the final image. Ideally, the generated image should be consistent with the image before the forward process, that is, X 0 Thus, image reconstruction is achieved. Wherein, T in the figure is a preset positive integer, t is any time step between time step 0 and T, and t is less than or equal to T.
[0052] like Figure 3 As shown, X 0 is an image. In the forward process, from time point 0 to time point 1 is the first time step, from time point 1 to time point 2 is the second time step, ..., from time point t-1 to time point t is the tth time step, ..., from time point T-1 to time point T is the Tth time step, X T is the noise image output after T time steps of noise processing, from X 0 To X T is a Markov chain. Among them, q(X t |X t -1,X 0) refers to the Gaussian noise added to X at the t-th time step during the forward process. During the reverse process, from T to T - 1 is the 1st time step, and so on, with the last time step being from 1 to 0. pθ(X t-1 |X t-1 |X t ) refers to the noise predicted at the t-th time step during the reverse process, and θ are the model parameters of the diffusion model. Figure 3 What is shown is an ideal situation. However, in an actual model, there are differences between the predicted noise and the added noise. Therefore, the images generated during the reverse process also differ from the images at each time step during the forward process. In the embodiments of this application, the images predicted at each time step, such as the t-th time step, are denoted as X^ t .
[0053] The principle of the text-to-image model in the embodiments of this application will be specifically described below for the images preferred by the user and the images not preferred by the user:
[0054] For the images preferred by the user, during the forward process, Gaussian noise is added to the images preferred by the user during the noise addition process from time point 0 to time point T, generating noise images corresponding to the images preferred by the user. Moreover, as the time step increases, the proportion of noise in the images preferred by the user becomes higher and higher. During the reverse process, denoising is performed from time point T to time point 0 to obtain the denoising result corresponding to the starting time step, which is called the first image.
[0055] For the images not preferred by the user, during the forward process, Gaussian noise is added to the images not preferred by the user during the noise addition process from time point 0 to time point T, generating noise images corresponding to the images not preferred by the user. Moreover, as the time step increases, the proportion of noise in the images not preferred by the user becomes higher and higher. During the reverse process, denoising is performed from time point T to time point 0 to obtain the denoising result corresponding to the starting time step, which is called the second image.
[0056] If in the traditional way, the text-to-image model is used to predict the noise and perform denoising processing step by step during the reverse process, the efficiency is very low. Therefore, in the embodiments of this application, a better training method is provided. The above reverse process is divided into two processes. By randomly selecting a time step t between time steps T and 0, the reverse process is divided into two stages: the denoising process from time point T to time point t and the denoising process from t to 0, including:
[0057] Stage 1: For the denoising process from time point T to time point t, estimate the denoising result corresponding to the t-th time step; then input the denoising result corresponding to the t-th time step into the text-to-image model to predict the denoising result corresponding to the t - 1-th time step.
[0058] Stage 2: For the denoising process from time point t to time point 0, use the denoising result corresponding to the (t - 1)-th time step to estimate the denoising result corresponding to the starting time step.
[0059] Here, the starting time step can refer to the time step at which Gaussian noise is first added to the input image in the forward process (e.g., from time point 0 to time point 1 is the starting time step).
[0060] As one possible implementation, when estimating the denoising result corresponding to the t-th time step in the above Stage 1, instead of using the text-to-image model to make predictions step by step in time, a correction coefficient array is introduced. Thus, using the noise added in the forward noise addition process and the correction coefficient array, the denoising result corresponding to the t-th time step can be obtained.
[0061] Among them, the correction coefficient array includes T elements, and each element corresponds to the value of the correction coefficient corresponding to each time step among the above T time steps. The correction coefficient corresponding to each time step reflects the difference between the noise prediction result corresponding to this time step and the noise added at this time step in the forward process. With the correction coefficient, in theory, the noise prediction result of the t-th time step can be directly estimated using the noise added at the t-th time step, and then the denoising result corresponding to the t-th time step can be estimated. This method does not need to start from the T-th time step and use the text-to-image model to make predictions step by step in time until the t-th time step, significantly improving the efficiency.
[0062] The following details the denoising process from time point T to time point t:
[0063] Starting from the T-th time step, for each time step i, the following steps are sequentially executed until the t-th time step is completed: Use the value of the correction coefficient corresponding to the i-th time step in the correction coefficient array and the noise-added result corresponding to the i-th time step sampled from the noise addition process to estimate the noise corresponding to the i-th time step.
[0064] Then, use the noises corresponding to the T-th time step to the t-th time step obtained by estimation to perform denoising processing on the noisy image in the data, and obtain the denoising result corresponding to the t-th time step.
[0065] In Figure 3 , in the forward process, each of the T time steps corresponds to a noise-added result. After recording this noise-added result, during the training process, the noise-added results corresponding to each time step from the T-th time step to the t-th time step can be sampled from the T noise-added results.
[0066] Assume that the correction coefficient corresponding to the i-th time step is r i , and the noise q(X i |Xi-1 , X 0 ), then the noise corresponding to the i-th time step in the reverse denoising process can be estimated as pθ(X i |X i-1 ), for example, the following formula can be used:
[0067]
[0068] Starting from the T-th time step, the noises corresponding to the T-th time step to the t-th time step can be obtained respectively in the above manner. Then, by using the noise images in the input data to remove the noises corresponding to the T-th time step to the t-th time step respectively, the denoising result of the t-th time step can be obtained. Essentially, it uses the noises sampled in the forward noise addition process and the bias correction coefficient array to perform a simple "calculation" as shown in formula (1) to obtain the denoising result of the t-th time step. Compared with using the text-to-image model to perform step-by-step inference, the computational complexity is greatly reduced and the computational efficiency is improved.
[0069] After obtaining the denoising result of the t-th time step, inputting the denoising result of the t-th time step into the text-to-image model for denoising processing can infer the denoising result of the t - 1-th time step.
[0070] It can be seen that from the T-th time step to the t - 1-th time step, it is equivalent to only using the text-to-image model for inference and denoising processing once (as described before, the denoising results from the T-th time step to the t-th time step are obtained by a simple "calculation" as shown in formula (1)), which greatly improves the efficiency.
[0071] Before training the text-to-image model in the embodiments of this application, the method may further include initializing the bias correction coefficient array. For example, assigning initial values to the bias correction coefficients in the bias correction coefficient array.
[0072] Continuing from the above, in the stage from time point T to time point t, after obtaining the denoising result corresponding to the t - 1-th time step, the method may further include: using the denoising result corresponding to the t - 1-th time step to update the value of the bias correction coefficient corresponding to the t - 1-th time step in the bias correction coefficient array. That is to say, the bias correction coefficient array is also gradually updated during the model training process.
[0073] In one example, the continuous update of the bias correction coefficient value can be achieved by using the exponential moving average method, so as to continuously correct the estimation of the denoising process from T to t. The process of updating the bias correction coefficient value is described in detail below:
[0074] Updating the correction coefficient value corresponding to the (t - 1)-th time step using the denoising result corresponding to the (t - 1)-th time step includes: determining the correction coefficient error value corresponding to the (t - 1)-th time step using the denoising result corresponding to the (t - 1)-th time step and the noisy result corresponding to the (t - 1)-th time step sampled from the noise addition process; performing weighted averaging on the correction coefficient value r t-1 corresponding to the (t - 1)-th time step in the correction coefficient array and the correction coefficient error value; using the value obtained from the weighted averaging to update the correction coefficient value corresponding to the (t - 1)-th time step in the correction coefficient array. For example, the following formula can be used:
[0075]
[0076] where α is a hyperparameter, and empirical values or experimental values can be used, is the updated correction coefficient value corresponding to the (t - 1)-th time step, μ θ (X t , t) is the distribution mean of the denoising result corresponding to the (t - 1)-th time step, and μ(X t , X 0 ) is the distribution mean of the noisy result corresponding to the (t - 1)-th time step sampled.
[0077] As one possible implementation, when estimating the denoising result corresponding to the initial time step in the above stage 2, instead of using the text-to-image model to make predictions step by step in time, the denoising result corresponding to the (t - 1)-th time step can be used for implicit denoising diffusion processing to obtain the denoising result corresponding to the starting time step.
[0078] Here, the implicit denoising diffusion processing can refer to using DDIM (Denoising Diffusion Implicit Model) for implicit denoising diffusion processing. Among them, the core idea of DDIM is to regard the time point t and the time point 0 as two adjacent vertices of the sub-chain of the original denoising Markov chain through non-Markov. Different from the traditional diffusion model, DDIM allows skipping some intermediate steps in the reverse process, so that it is possible to determine the denoising result corresponding to the starting time step through one-step reasoning (that is, directly performing one-step reasoning from t to 0), avoiding an overly complex reverse propagation gradient graph, and ensuring the feasibility of training the text-to-image model while guaranteeing the training efficiency. The process of estimating X 0 using DDIM can be shown by the following formula:
[0079]
[0080] where α t-1 and α tThey are the weight coefficients corresponding to the t-th time step and the (t-1)-th time step respectively. In DDIM, the contribution of different time steps is controlled by the weight coefficients corresponding to each time step, which are pre-set hyperparameters. is the noise predicted at the t-th time step, σ t is the variance of the noise distribution predicted at the t-th time step.
[0081] It should be noted that during the process of stage 2 from time point t to time point 0, in addition to using DDIM for one-step inference, a denoising process step by step in time can also be adopted, or DDIM can be used for multi-step inference, but the step size used for inference is greater than the step size of the above time steps. For example, 10 time steps are used as one step size for DDIM inference, and so on.
[0082] Through experiments, it is found that during the training process of the text-to-image model, the value of the correction coefficient is larger when approaching time point 0, which can push the loss function into the gradient saturation region, thereby suppressing the optimization near time point 0. In DDIM, the time steps close to time point T usually correspond to a high noise level, and the noise estimation may not be accurate enough at these time steps. DDIM suppresses the optimization near time point T by introducing weight coefficients. For example, the value of the weight coefficient is larger near time point T, which causes the error between the predicted noise and the actually added noise to be amplified more as it gets closer to time point T, making it easier for the optimization for preference to enter the saturation region of the loss function, thereby suppressing the optimization near time point T and reducing the influence of the noise estimation error at these time steps on the noise-added result corresponding to the initial time step, thus improving the quality of the generated image. Combining these two methods weakens the optimization at both ends (i.e., near time point 0 and time point T) during the reverse process, emphasizes more the optimization of the time steps near the middle part, realizes more refined supervision. Through experiments, this method is significantly better than the traditional supervision method, and the efficiency is also significantly improved.
[0083] Furthermore, when training the text-to-image model, it is often based on the existing text-to-image model and the method provided in the embodiments of the present application is used for further training and optimization, that is, the existing text-to-image model is used as the initial model for optimization to make it more in line with user preferences. The existing text-to-image model usually already has relatively excellent capabilities in text understanding and image generation accuracy. Therefore, when further training and optimizing the text-to-image model in the embodiments of the present application, it is also necessary to ensure that the difference between the text-to-image model and the initial model is as small as possible.
[0084] In view of this, the initial text-to-image model can be used as a reference model. The text description sample and the noisy image obtained by adding noise to the positive sample are input into the reference model at the same time to obtain the third image obtained after the reference model denoises the input noisy image. Also, the text description sample and the noisy image obtained by adding noise to the negative sample are input into the reference model at the same time to obtain the fourth image obtained after the reference model denoises the input noisy image. The training objective at this time can further include minimizing the difference between the first image and the third image, and minimizing the difference between the second image and the fourth image. As Figure 4 shown, the loss function L2 can be further designed according to this training objective so that L2 represents the difference between the first image and the third image, and the difference between the second image and the fourth image.
[0085] As a more preferred implementation, when training the text-to-image model, the loss function L used can be jointly determined by the above L1 and L2. For example, L is obtained by weighted summation of L1 and L2. Then, using the loss function L, the model parameters of the text-to-image model are updated by means such as gradient descent until the preset training end condition is reached.
[0086] In addition to the above preferred implementation, other implementation methods can also be used. For example, the loss function L is only determined by the above L1, or the loss function L is determined by the above L1 combined with other loss functions, or the loss function is determined by L1, L2 combined with other loss functions, and so on.
[0087] Figure 5 is a schematic diagram of the image generation method of the embodiment of the present application. This method can be executed by Figure 1 the image generation set on the server side as shown in Figure 5 shown. This method can include the following steps:
[0088] Step 501: Obtain a text description.
[0089] In the embodiment of the present application, after the text-to-image model is trained, the text-to-image service can be provided for users. The text-to-image model can accurately understand the context and intention of the text description input by the user and generate an image based on the text description.
[0090] In different application scenarios, the source of the obtained text description can be different. For example, the text input by the user through the text box provided by the text-to-image service, or the text obtained after speech recognition of the speech input by voice.
[0091] For another example, the text description can be extracted from existing text data on a page or in a database. For instance, in an online food service, if it is desired to provide users with pictures corresponding to the dishes on the menu, then the text description can be extracted from the dish names, consumers' reviews, and merchants' descriptions. Using a text-to-image model, an image can be generated based on the text description, and this image can be displayed on the web page as the picture corresponding to the dish.
[0092] Step 504: Input the text description and the noise image into the text-to-image model to obtain the image corresponding to the text description.
[0093] The above-mentioned noise image can be a randomly generated noise image. According to the default size of the service or the size set by the user, etc., a noise image of the corresponding size can be randomly generated. Inputting the noise image and the text description into the text-to-image model pre-trained using the Figure 2 method shown can obtain the image output by the text-to-image model, and this image is more in line with the user's preferences.
[0094] As can be seen from the above embodiments, the technical solution provided by this application can also have the following advantages:
[0095] 1) In the embodiment of this application, the denoising process performed during the training process of the text-to-image model is divided into two stages: the stage from time point T to time point t and the stage from time point t to time point 0. And the text-to-image model only needs to perform one-step inference (that is, use the denoising result corresponding to the t-th time step that has been estimated to obtain the denoising result corresponding to the t - 1-th time step) to estimate the denoising result corresponding to the starting time step. Compared with the denoising process of using the text-to-image model step by step from time point T to time point 0, the speed of the denoising process is accelerated, and the training efficiency of the text-to-image model is improved.
[0096] 2) In the embodiment of this application, for the stage from time point T to time point t, an array of bias correction coefficients for correcting the error of the denoising result is introduced. Using the array of bias correction coefficients and the noise addition result sampled for the forward noise addition process, the noise corresponding to the t-th time step can be estimated, and then the denoising result corresponding to the t-th time step can be obtained, without the need to use the text-to-image model for step-by-step denoising processing, significantly improving the training efficiency.
[0097] 3) In the embodiment of this application, using the denoising result corresponding to the t - 1-th time step obtained by the text-to-image model, the value of the bias correction coefficient corresponding to the t - 1-th time step in the array of bias correction coefficients is dynamically updated, so that the bias correction coefficient can also be continuously optimized during the training process of the text-to-image model, and further continuously improve the estimation accuracy of the denoising result for the randomly sampled t-th time step.
[0098] 4) In the embodiment of the present application, for the stage from time point t to time point 0, implicit denoising diffusion processing is performed using the denoising result corresponding to the (t - 1)-th time step to obtain the denoising result corresponding to the starting time step. Compared with the traditional denoising diffusion processing, the number of time steps used for denoising prediction is reduced, thereby significantly improving the training efficiency of the text-to-image model.
[0099] 5) In the process of training the text-to-image model in the embodiment of the present application, the difference between the image generation ability of the text-to-image model during the training process and the image generation ability of the reference model (i.e., the initial text-to-image model) is further considered, so as to ensure that the existing image generation ability is maintained during the optimization of the user preferences for the text-to-image model, thereby meeting the user preferences while ensuring the accuracy of image generation.
[0100] Of course, it is not necessary for any product implementing the present application to achieve all the above-mentioned advantages simultaneously.
[0101] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0102] According to an embodiment of another aspect, a training device for a text-to-image model is provided. Figure 6 A schematic block diagram of the training device for the text-to-image model according to one embodiment is shown. The device is provided in Figure 1 the server side in the architecture shown. As Figure 6 shown, the device 600 includes: a data acquisition unit 601 and a model training unit 602. The main functions of each component unit are as follows:
[0103] The data acquisition unit 601 is configured to acquire training data including a plurality of sample pairs. The sample pair includes a text description sample, a positive sample corresponding to the text description sample, and a negative sample. The positive sample is an image preferred by the user, and the negative sample is an image not preferred by the user.
[0104] The model training unit 602 is configured to train a text-to-image model implemented based on a diffusion model using training data, where the training includes: using a text description sample and a noisy image obtained by adding noise to a positive sample as input data, and using the text-to-image model to denoise the noisy image in the input data to obtain a first image; and using the text description sample and a noisy image obtained by adding noise to a negative sample as input data, and using the text-to-image model to denoise the noisy image in the input data to obtain a second image; optimizing the model parameters of the text-to-image model based on a training objective, where the training objective includes: minimizing the difference between the first image and the positive sample corresponding to the text description sample, and maximizing the difference between the second image and the negative sample corresponding to the text description sample.
[0105] As one possible implementation, the above noise addition includes noise addition processing for T time steps. When the model training unit 602 uses the text-to-image model to denoise the noisy image in the input data, it is specifically configured to:
[0106] Randomly select the t-th time step among the T time steps, and estimate the denoising result corresponding to the t-th time step;
[0107] Input the denoising result corresponding to the t-th time step into the text-to-image model to obtain the denoising result corresponding to the (t - 1)-th time step;
[0108] Use the denoising result corresponding to the (t - 1)-th time step to estimate the denoising result corresponding to the starting time step;
[0109] Wherein, T and t are preset positive integers, and t is less than or equal to T.
[0110] As one possible implementation, when the model training unit 602 estimates the denoising result corresponding to the t-th time step, it is specifically configured to:
[0111] Starting from the T-th time step, for each time step i in sequence, perform the following steps until the t-th time step is completed: use the correction coefficient value corresponding to the i-th time step in the correction coefficient array and the noise addition result corresponding to the i-th time step sampled from the noise addition processing to estimate the noise corresponding to the i-th time step, and the correction coefficient array includes the correction coefficient values corresponding to each time step among the T time steps;
[0112] Use the estimated noise corresponding to the T-th time step to denoise the noisy image in the input data to obtain the denoising result corresponding to the t-th time step.
[0113] Furthermore, the model training unit 602 can also be configured to: initialize the correction coefficient array; after obtaining the denoising result corresponding to the (t-1)-th time step, use the denoising result corresponding to the (t-1)-th time step to update the value of the correction coefficient corresponding to the (t-1)-th time step in the correction coefficient array.
[0114] As one possible implementation, when the model training unit 602 updates the value of the correction coefficient corresponding to the (t-1)-th time step in the correction coefficient array, it can be specifically configured to:
[0115] Use the denoising result corresponding to the (t-1)-th time step and the noise addition result corresponding to the (t-1)-th time step sampled from the noise addition process to determine the correction coefficient error value corresponding to the (t-1)-th time step;
[0116] Perform a weighted average on the value of the correction coefficient corresponding to the (t-1)-th time step in the correction coefficient array and the correction coefficient error value;
[0117] Use the value obtained by the weighted average to update the value of the correction coefficient corresponding to the (t-1)-th time step in the correction coefficient array.
[0118] As one possible implementation, when the model training unit 602 estimates the denoising result corresponding to the starting time step by using the denoising result corresponding to the (t-1)-th time step, it can be specifically configured to:
[0119] Perform implicit denoising diffusion processing on the denoising result corresponding to the (t-1)-th time step to obtain the denoising result corresponding to the starting time step.
[0120] Furthermore, the above training is obtained by optimizing the model parameters of the initial text-to-image model. The model training unit 602 can also be configured to: use the initial text-to-image model as a reference model, input the text description sample and the noise image obtained by adding noise to the positive sample into the reference model, obtain the third image obtained by the reference model after denoising the input noise image, and, input the text description sample and the noise image obtained by adding noise to the negative sample into the reference model, obtain the fourth image obtained by the reference model after denoising the input noise image; correspondingly, the training objective also includes: minimizing the difference between the first image and the third image, and, minimizing the difference between the second image and the fourth image.
[0121] According to an embodiment of another aspect, an image generation device is provided. Figure 7 A schematic block diagram of the image generation device according to an embodiment is shown. As Figure 7 shown, the device 700 includes: a description acquisition unit 701 and an image generation unit 702. The main functions of each component unit are as follows:
[0122] A description acquisition unit 701, configured to acquire a text description.
[0123] An image generation unit 702, configured to input the text description and a noise image into a text-to-image model to obtain an image corresponding to the text description, where the text-to-image model is pre-trained by using the Figure 6 device shown.
[0124] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0125] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0126] In addition, the embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any one of the foregoing method embodiments.
[0127] And an electronic device, including:
[0128] One or more processors; and
[0129] A memory associated with the one or more processors, where the memory is used to store program instructions. When the program instructions are read and executed by the one or more processors, they execute the steps of the method described in any one of the foregoing method embodiments.
[0130] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the method described in any one of the foregoing method embodiments.
[0131] Among them, Figure 8 An exemplary architecture of an electronic device is shown, which may specifically include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The above-mentioned processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, and the memory 820 can be communicatively connected through a communication bus 830.
[0132] Among them, the processor 810 can be implemented in ways such as a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the present application.
[0133] The memory 820 can be implemented in forms such as ROM (Read Only Memory), RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 820 can store an operating system 821 for controlling the operation of the electronic device 800, and a basic input / output system (BIOS) 822 for controlling the low-level operations of the electronic device 800. In addition, a web browser 823, a data storage management system 824, and a training device 600 / image generation device 700 of the text-to-image model, etc. can also be stored. The above-mentioned training device 600 / image generation device 700 of the text-to-image model can be the application program that specifically implements the operations of the foregoing steps in the embodiments of the present application. In short, when implementing the technical solutions provided by the present application through software or firmware, the relevant program codes are stored in the memory 820 and are called and executed by the processor 810.
[0134] The input / output interface 813 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0135] The network interface 814 is used to connect to a communication module (not shown in the figure) to enable communication and interaction between this device and other devices. The communication module can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0136] The bus 830 includes a path for transmitting information between various components of the device (such as the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, and memory 820).
[0137] It should be noted that although the above device only shows the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, memory 820, bus 830, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of this application, and does not necessarily include all the components shown in the figure.
[0138] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer program product. The computer program product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0139] The above has introduced the technical solution provided by this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A training method for a text graph model, characterized in that: The method comprises: Acquire training data including a plurality of sample pairs, wherein the sample pairs include text description samples, positive samples and negative samples corresponding to the text description samples, wherein the positive samples are images preferred by the user, and the negative samples are images not preferred by the user; The training data is used to train a Vincent graph model based on a diffusion model, wherein the training includes: The text description sample and the noise image obtained by adding noise to the positive sample are used as input data, and the noise image in the input data is denoised using the Wensheng graph model to obtain a first image; and the text description sample and the noise image obtained by adding noise to the negative sample are used as input data, and the noise image in the input data is denoised using the Wensheng graph model to obtain a second image; and the model parameters of the Wensheng graph model are optimized based on a training objective, and the training objective includes: minimizing the difference between the first image and the positive sample corresponding to the text description sample, and maximizing the difference between the second image and the negative sample corresponding to the text description sample.
2. The method according to claim 1, characterized in that The adding of noise includes a denoising process of T time steps, and the denoising process of the noise image in the input data using the Wensheng graph model includes: Randomly select a t-th time step from the T time steps, and estimate a denoising result corresponding to the t-th time step; Inputting the denoising result corresponding to the t-th time step into the Wensheng graph model to obtain the denoising result corresponding to the t-1-th time step; Using the denoising result corresponding to the t-1th time step, estimating the denoising result corresponding to the starting time step; Wherein, T and t are preset positive integers, and t is less than or equal to T.
3. The method according to claim 2, characterized in that The estimating the denoising result corresponding to the t-th time step includes: Starting from the Tth time step, for each time step i, the following steps are performed in sequence until the tth time step: using the correction coefficient value corresponding to the i-th time step in the correction coefficient array and the noise addition result corresponding to the i-th time step obtained by sampling from the noise addition process, the noise corresponding to the i-th time step is estimated, and the correction coefficient array includes the correction coefficient values corresponding to each time step in the T time steps; The noise image in the input data is denoised using the estimated noises corresponding to the T-th time step to the t-th time step to obtain a denoising result corresponding to the t-th time step.
4. The method according to claim 3, characterized in that Before training the Vincent graph model, the method further includes: initializing the correction coefficient array; After obtaining the denoising result corresponding to the t-1th time step, the method further includes: using the denoising result corresponding to the t-1th time step to update the correction coefficient value corresponding to the t-1th time step in the correction coefficient array.
5. The method according to claim 4, characterized in that The updating of the correction coefficient value corresponding to the t-1th time step in the correction coefficient array by using the denoising result corresponding to the t-1th time step includes: Determine the correction coefficient error value corresponding to the t-1th time step by using the denoising result corresponding to the t-1th time step and the denoising result corresponding to the t-1th time step sampled from the denoising process; Taking a weighted average of the correction coefficient value corresponding to the t-1th time step in the correction coefficient array and the correction coefficient error value; The value obtained by the weighted average is used to update the correction coefficient value corresponding to the t-1th time step in the correction coefficient array.
6. The method according to claim 2, characterized in that The estimating the denoising result corresponding to the starting time step by using the denoising result corresponding to the t-1th time step includes: The denoising result corresponding to the t-1th time step is used to perform implicit denoising diffusion processing to obtain the denoising result corresponding to the starting time step.
7. The method according to any one of claims 1 to 6, characterized in that The training is obtained by optimizing the model parameters of the initial Wensheng graph model, and the initial Wensheng graph model is implemented based on the diffusion model; The method further includes: using the initial text image model as a reference model, inputting the text description sample and the noise image obtained by adding noise to the positive sample into the reference model, obtaining a third image obtained by denoising the input noise image by the reference model, and inputting the text description sample and the noise image obtained by adding noise to the negative sample into the reference model, obtaining a fourth image obtained by denoising the input noise image by the reference model; The training objective also includes minimizing the difference between the first image and the third image, and minimizing the difference between the second image and the fourth image.
8. An image generation method, characterized in that: The method comprises: Get text description; The text description and the noise image are input into a text-generated graph model to obtain an image corresponding to the text description, wherein the text-generated graph model is pre-trained using the method according to any one of claims 1 to 7.
9. An image generating device, characterized in that: The device comprises: A description acquisition unit, configured to acquire a text description; The image generation unit is configured to input the text description and the noise image into a text-generated graph model to obtain an image corresponding to the text description, wherein the text-generated graph model is pre-trained using the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.