Fine-tuning method, device, electronic device and storage medium of text-to-image model
Through a two-stage training method, combined with the training of text encoder and diffusion generation model, the problems of limited learning ability and overfitting in traditional fine-tuning methods are solved, achieving more efficient and accurate image generation.
Patent Information
- Application Number
- CN202411382118.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The traditional fine-tuning method of literary and biographical graphics model only fine-tune the image diffusion model and freezes the text encoder, resulting in limited model learning ability and serious overfitting behavior, which affects the efficiency and accuracy of image generation.
A two-stage training method is proposed. The first stage trains the text encoder and the diffusion generation model at the same time. The second stage freezes the text encoder and only trains the diffusion generation model. By adjusting the keyword weight of the text encoder, and optimizing the model parameters through the loss function during the training process.
It effectively enhances the learning ability of the image generation model, reduces overfitting behavior, improves the efficiency and accuracy of image generation, and avoids the "burning effect" and "catastrophic forgetting" phenomena.
Smart Images

Figure CN119312842B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for fine-tuning a text-to-image model. Background Art
[0002] In the field of AI (artificial intelligence) image generation technology, a diffusion model is usually used to generate images from text, that is, the text-to-image process, or also known as the AI painting process. The specific process of text-to-image includes: the text input by the user is output by the text encoder model and encoded into a text embedding vector. At the same time, the text embedding vector is passed into the image diffusion model as a conditioning mechanism, and an image with the concept described by the text is output in the image diffusion model.
[0003] In related technologies, when new knowledge and concepts need to be added to a pre-trained text-to-image diffusion model, people will achieve this by fine-tuning the text-to-image model network. The traditional model fine-tuning method only fine-tunes the network structure of the image diffusion model, while freezing (i.e., not fine-tuning) the text encoder, which greatly limits the learning ability of the text-to-image model. Summary of the Invention
[0004] To solve or partially solve the problems existing in the related technologies, this application provides a method, device, electronic device, and storage medium for fine-tuning a text-to-image model, which enhances the learning ability of the image generation model, reduces the overfitting behavior of training the image generation model during the image generation process, and at the same time improves the efficiency and accuracy of image generation.
[0005] The first aspect of this application provides a method for fine-tuning a text-to-image model, including:
[0006] Obtain a preset image training set, where the preset image training set includes a first image and a first text corresponding to the first image;
[0007] Train a preset text-to-image model according to the preset image training set to obtain a target text-to-image model, and the target text-to-image model is used to output the first image according to the first text;
[0008] Among them, the preset text-to-image model includes a text encoder and a diffusion generation model. The preset text-to-image model is trained according to the preset image training set to obtain the target text-to-image model, including: in the first training stage, the preset image training set is trained based on the text encoder and the diffusion generation model to obtain the first target model corresponding to the text encoder and the second candidate model corresponding to the diffusion generation model; in the second training stage, the second candidate model is trained based on the preset image training set to obtain the second target model; the target text-to-image model is obtained according to the first target model and the second target model;
[0009] The user image text is obtained according to the user input instruction, and the target image is obtained based on the user image text and the target text-to-image model.
[0010] Optionally, obtaining the preset image training set includes:
[0011] Obtain image keywords, which are used to describe the generation of the first image;
[0012] Add the image keywords to the first text for annotation, and the first text is at least annotated with one image keyword.
[0013] Optionally, obtaining the image keywords includes:
[0014] Expand the first text based on the preset semantic rules to obtain multiple first expanded texts;
[0015] Semantically supplement the first expanded text based on the preset painting style and / or artist name to obtain the second expanded text;
[0016] Retrieve the associated text of the second expanded text to obtain the third expanded text associated with the second expanded text;
[0017] Based on the preset classification algorithm, classify the third expanded text and the first image to obtain the image keywords corresponding to each first image.
[0018] Optionally, adding the image keywords to the first text for annotation, where the first text is at least annotated with one image keyword, includes:
[0019] Arrange the image keywords according to the preset sequence numbers to obtain the sequence numbers corresponding to the image keywords;
[0020] Obtain the image keywords associated with the first text, and use the sequence numbers of the image keywords associated with the first text as the annotation of the first text;
[0021] Among them, each image keyword in the annotation of the first text carries an initial weight.
[0022] Optionally, in the first training stage, based on the text encoder and the diffusion generation model, train a preset image training set to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model, including:
[0023] Input the preset image training set into the first preset model, and output a second text and semantic information corresponding to the second text, where the second text is an annotated text obtained by updating the keyword weights of the first text;
[0024] Input the second text and the semantic information corresponding to the second text into the diffusion generation model to obtain a first type of initial image;
[0025] Calculate the first loss function of the preset text-to-image model according to the first type of initial image and the first image;
[0026] Repeat training the preset text-to-image model until the first loss function converges to obtain the preset text-to-image model completed in the first stage. The preset text-to-image model completed in the first stage includes the first target model and the second candidate model.
[0027] Optionally, in the second training stage, train the second candidate model based on the preset image training set to obtain a second target model, including:
[0028] In the second training stage, obtain a third text, where the third text is the text obtained by updating the annotation weights of the first text in the first training stage;
[0029] According to the third text and the first target model, obtain semantic information corresponding to the third text;
[0030] Use the semantic information corresponding to the third text and the first image corresponding to the third text as the training set for the second stage;
[0031] Train the second candidate model according to the training set for the second stage to obtain a second target model. The second target model is used to output the first image corresponding to the third text according to the semantic information of the third text.
[0032] Optionally, train the second candidate model according to the training set for the second stage to obtain a second target model, including:
[0033] Input the semantic information corresponding to the third text into the second candidate model to obtain a second type of initial image;
[0034] Calculate the second loss function of the preset text-to-image model according to the second type of initial image and the first image;
[0035] Retrain the second candidate model until the second loss function converges to obtain a preset text-to-image model completed in the first stage. The preset text-to-image model completed in the second stage includes a first target model and a second target model.
[0036] Optionally, obtaining user image text according to a user input instruction, and obtaining a target image according to the user image text and the target text-to-image model, includes:
[0037] Search for image keywords corresponding to the user instruction according to the user input instruction to obtain user image text;
[0038] Obtain a first target semantics according to the image text and the first target model, and generate a target image according to the first target semantics and the second target model.
[0039] In a second aspect, the present application provides an image generation device, including:
[0040] An acquisition unit, configured to acquire a preset image training set, where the preset image training set includes a first image and a first text corresponding to the first image;
[0041] A training unit, configured to train a preset text-to-image model according to the preset image training set to obtain a target text-to-image model, where the target text-to-image model is used to output a first image according to the first text;
[0042] Wherein, the preset text-to-image model includes a text encoder and a diffusion generation model, and the training unit includes a first training unit, a second training unit, and a third training unit. The training unit trains the preset text-to-image model according to the preset image training set to obtain a target text-to-image model, including: the first training unit is configured to, in a first training stage, train the preset image training set based on the text encoder and the diffusion generation model to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model; the second training unit is configured to, in a second training stage, train the second candidate model based on the preset image training set to obtain a second target model; the third training unit is configured to obtain a target text-to-image model according to the first target model and the second target model;
[0043] An image generation unit, configured to obtain user image text according to a user input instruction, and obtain a target image based on the user image text and the target text-to-image model.
[0044] Optionally, the image generation unit includes: a first image subunit and a second image subunit; the first image subunit is configured to search for image keywords corresponding to the user instruction according to the user input instruction to obtain user image text; the second image subunit is configured to obtain a first target semantics according to the image text and the first target model, and generate a target image according to the first target semantics and the second target model.
[0045] A third aspect of the present application provides an electronic device, including:
[0046] a processor; and
[0047] a memory storing executable code thereon, which when executed by the processor, causes the processor to execute the method as described above.
[0048] A fourth aspect of the present application provides a computer-readable storage medium storing executable code thereon, which when executed by a processor of an electronic device, causes the processor to execute the method as described above.
[0049] The technical solution provided by the present application may include the following beneficial effects:
[0050] In a first aspect, the present application provides a method for training a preset text-to-image model based on a preset image training set to obtain a target text-to-image model. The preset image training set includes a first image and a first text corresponding to the first image. When using the preset image training set, each first text carries the weight of an image keyword. By adjusting the keyword weight of the first text during the training process, it is convenient for users to identify key features, improving the training efficiency. At the same time, the weight feature can suppress data noise generated during the training process, enhancing the robustness of model training.
[0051] In a second aspect, the present application includes two training stages when training the preset text-to-image model. In the first training stage, the text encoder and the diffusion generation model are trained simultaneously to accelerate the training efficiency. In the second training stage, the text encoder is frozen and the diffusion generation model is trained, effectively avoiding overfitting of the text encoder, which may lead to the "burning effect" of the target text-to-image model, that is, the generated image has too high contrast and brightness, or "catastrophic forgetting", that is, the model forgets the original knowledge during the training process. The two-stage training method of the present application not only accelerates the convergence speed of training, but also improves the final generation quality of the model, and effectively avoids the negative effects of text encoder overfitting.
[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.
[0054] Figure 1 is a schematic flowchart of the fine-tuning method of the text-to-image model shown in the embodiments of the present application;
[0055] Figure 2 It is another process schematic diagram of the fine-tuning method of the text-to-image model shown in the embodiments of the present application;
[0056] Figure 3 It is a process schematic diagram of training the target text-to-image model shown in the embodiments of the present application;
[0057] Figure 4 It is the target image generated by the target text-to-image model according to the user input instruction shown in the embodiments of the present application;
[0058] Figure 5 It is another target image generated by the target text-to-image model according to the user input instruction shown in the embodiments of the present application;
[0059] Figure 6 It is a structural schematic diagram of the text-to-image processing device shown in the embodiments of the present application;
[0060] Figure 7 It is a structural schematic diagram of the electronic device shown in the embodiments of the present application. Detailed Embodiments
[0061] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0062] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0063] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0064] In the field of AI (artificial intelligence) image generation technology, diffusion models are usually used to generate images from text, that is, the text-to-image process, or the so-called AI painting process. The specific process of text-to-image includes: the text input by the user is encoded into a text embedding vector through a text encoder model (such as the CLIP model or the T5 Encoder model), and at the same time, the text embedding vector is passed into the image diffusion model (such as the UNet architecture) as a conditioning mechanism, and an image with the concepts described by the text is output in the image diffusion model.
[0065] In related technologies, when new knowledge and concepts need to be added to a pre-trained text-to-image model, it is usually achieved by fine-tuning the network of the text-to-image model. During the fine-tuning process, the text-to-image model is fine-tuned through a specific data set containing new knowledge and concepts, so that the text-to-image model can adapt to a specific task domain. The traditional model fine-tuning method only fine-tunes the network structure of the image diffusion model while freezing (i.e., not fine-tuning) the text encoder, which greatly limits the learning ability of the text-to-image model.
[0066] In view of the above problems, the embodiments of the present application provide a fine-tuning method for a text-to-image model, which enhances the learning ability of the image generation model, reduces the overfitting behavior of training the image generation model during the image generation process, and at the same time improves the efficiency and accuracy of image generation.
[0067] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0068] Figure 1 It is a schematic flowchart of the fine-tuning method for the text-to-image model shown in the embodiments of the present application.
[0069] See Figure 1 , the fine-tuning method for the text-to-image model of the present application includes:
[0070] Step S101, obtain a preset image training set, where the preset image training set includes a first image and a first text corresponding to the first image.
[0071] Among them, the first text and the first image form a text-image pair of training samples, and the first text is the text description information corresponding to the first image.
[0072] In one embodiment, obtaining a preset image training set includes: obtaining image keywords, where the image keywords are used to describe the generation of a first image; adding the image keywords to a first text for annotation, and the first text is at least annotated with one image keyword.
[0073] Among them, first, the first image and image keywords are collected. Each image keyword is used to determine the relevant elements for generating the first image, and the keyword corresponds to the text description information of the relevant elements. For example. The image keywords include, but are not limited to, the name of a well-known artist, the painting style of the image, the painting elements appearing in the image, the color of the painting elements, the theme expressed by the image, etc. For example, for the first text "a wolf in ink painting style", it includes at least image keywords such as "wolf", "ink painting style", and "a". It should be noted that the image keywords can be either Chinese or English and can be switched at any time according to the user's choice. Preprocess the first image and image keywords, match the first image and image keywords, and after the matching is completed, annotate the image keywords with weights in sequence to generate the first text, where the weights are used to represent the importance of the image keywords in the text-to-image generation process.
[0074] In this embodiment, by using the form of annotating image keywords with weights, the data content corresponding to the first text is made more accurate, which helps to reduce the training error of the model, improve the training accuracy, and enable it to learn better and have stronger generalization ability during the training process.
[0075] To further improve the accuracy of image annotation in the preset image training set, as Figure 2 shown, the present application provides a fine-tuning method for a text-to-image model for obtaining image keywords, including:
[0076] Step S201, expanding the first text based on a preset semantic rule to obtain multiple first expanded texts.
[0077] In this embodiment, step S201 includes: using a preset language text model to perform semantic expansion on the first text. For example, encoding the first text and encoding each of the multiple images are both implemented using the Contrastive Language-Image Pre-Training (CLIP) model. The CLIP model (Contrastive Language-Image Pre-Training, CLIP contrastive language-image pre-training model), by training on a large number of text-image pairs, learns the mapping relationship between images and texts, so as to realize semantic encoding of texts and images respectively in the same semantic space and obtain the corresponding vectors. Thus, semantic encoding of texts and images respectively in the same semantic space and obtaining the corresponding vectors to generate multiple first expanded texts.
[0078] Step S202: Semantically supplement the first extended text based on a preset painting style and / or artist name to obtain a second extended text.
[0079] In this embodiment, step S202 includes: determining the correspondence between the text and the painting type and / or artist name by means of pre-traversal and screening. Specifically, a set of painting types and / or a set of artist names can be obtained, and then each item in the set is used to expand the text respectively to construct extended texts such as [text, artist 1], [text, artist 2], [text, artist 3]. Corresponding images are generated based on each extended text respectively, and images with higher quality are screened through the similarity between the text vector and the image vector, so as to determine the correspondence between the text and the corresponding artist and / or painting type set and form a template, and apply the constructed correspondence between the text and the painting type and / or artist name, that is, the template, to the expansion of the first text.
[0080] Step S203: Retrieve the associated text of the second extended text to obtain a third extended text associated with the second extended text.
[0081] In this embodiment, the OOV (Out Of Vocabulary) method is used for retrieval. When performing natural language processing or text processing, there is usually a dictionary. This dictionary can be pre-loaded, or user-defined, or extracted from the current data set. Suppose there is another data set later, and there are some words in this data set that are not in the existing dictionary. These words are called out-of-dictionary words. Based on the above OOV method, traverse the second extended text in the preset database dictionary to further expand the second extended text, so as to make the semantic description of the first image more abundant.
[0082] Step S204: Classify the third extended text and the first image based on a preset classification algorithm to obtain image keywords corresponding to each first image.
[0083] In this embodiment, the preset classification algorithm can be manual classification, or the image keywords can be assigned to the corresponding first image according to a preset classifier. For example, use a preset classifier to classify and score the third extended text and the first image, and select the third extended text that meets the scoring requirements as the image keywords corresponding to the first image. Among them, the preset classifier can be the Naive Bayes algorithm.
[0084] In steps S201 - S204, various rules are used to expand the first text for generating images in multiple dimensions, so that the expression of the first text is more complete and rich, and further the images generated based on the expanded text have richer content.
[0085] In one embodiment, image keywords are added to the first text for annotation. The first text is annotated with at least one image keyword, including: arranging the image keywords in accordance with a preset sequence number to obtain the corresponding sequence numbers of the image keywords; obtaining the image keywords associated with the first text, and using the sequence numbers of the image keywords associated with the first text as the annotation of the first text; wherein, each image keyword in the annotation of the first text carries an initial weight.
[0086] By designing and training the initial weights for the image keywords, the embodiment can assign different weights to the keywords, enabling the model to better identify and distinguish important features during subsequent training, thereby improving the accuracy of classification and prediction. At the same time, during the training process, assigning higher weights to important keywords can make the model more focused on key features, thus optimizing the allocation of computing resources and facilitating the model to process large-scale data. Further, the weight design can help the model maintain stable performance when facing noise and unbalanced data. By adjusting the weights, the model can better handle outliers and noise in the data.
[0087] In this embodiment, some special words are used to mark new concept data. The new concept data uses one word to mark each special concept, namely "image keyword", and the image keyword is added to the data annotation. Specifically, the image keywords in the data are numbered in accordance with a preset sequence number, such as numbered 1, 2, …, k, and the i-th concept where 1 ≤ i ≤ k is artificially named and marked as c i . For each concept where 1 ≤ i ≤ k, the data is annotated with a text description containing c i . Initially, an initial weight is assigned to each image keyword, such as 10%, and the weight is used to identify the importance of the image keyword. If the user wants to weaken the content of a certain segment in the image keyword, the user can assign a lower initial weight to the image keyword.
[0088] The initial weight can, according to the user's instruction, assign an attention weight to each image keyword, or directly modify the output of the text encoder attention according to the user's instruction. Specifically, according to the attention weight of each image keyword, each dimension of the initial text embedding vector output by the text encoder can be weighted to obtain a weighted target text embedding vector, so as to achieve the modification of the attention output of the text encoder to the standard weight.
[0089] Step S102, training a preset text-to-image model according to a preset image training set to obtain a target text-to-image model, where the target text-to-image model is used to output a first image according to the first text.
[0090] Among them, the preset text-to-image model includes a text encoder and a diffusion generation model. The text encoder includes a text encoder, which can be used to convert text into a vector for representation. Therefore, the image keywords are input into the text encoder so that the text encoder can encode the text prompt into the initial text embedding vector of the diffusion generation model. Among them, the types of text encoders can include, but are not limited to: CLIP (Contrastive Language-Image Pre-training) model, T5 Encoder (Transfer Text-to-Text Transformer) model. The diffusion generation model can be used to generate a corresponding image from the initial text embedding vector generated by the text encoder. The diffusion generation model includes, but is not limited to, the U-Net network. U-Net is an encoder-decoder network based on convolutional neural networks and skip connections, generally used to generate an image of the same size as the input image. In order to make the preset text-to-image model better adapt to concept learning in different fields, it is necessary to perform fine-tuning training on the preset text-to-image model.
[0091] Step S103, where the preset text-to-image model includes a text encoder and a diffusion generation model. Training the preset text-to-image model according to the preset image training set to obtain the target text-to-image model includes: in the first training stage, training the preset image training set based on the text encoder and the diffusion generation model to obtain the first target model corresponding to the text encoder and the second candidate model corresponding to the diffusion generation model; in the second training stage, training the second candidate model based on the preset image training set to obtain the second target model; obtaining the target text-to-image model according to the first target model and the second target model.
[0092] In step S103, in the first training stage, the text encoder and the diffusion generation model are unfrozen, and at the same time, the text encoder and the diffusion generation model are trained to update the training parameters of the text encoder and the diffusion generation model, improving the model's learning and expression ability for image keywords. In the second training stage, the training parameters of the text encoder are frozen, and only the training parameters of the diffusion generation model are adjusted. By training the text encoder and the diffusion generation model respectively in the first training stage and the second training stage, not only the convergence speed of the training is accelerated, but also the final generation quality of the model is improved, and at the same time, overfitting of the text encoder is effectively avoided.
[0093] Figure 3 This is a schematic flowchart of a method for training a preset text-to-image model according to the present application. In combination with Figure 3 Step S103 is described.
[0094] In one embodiment, in the first training stage 301, based on a text encoder and a diffusion generation model, a preset image training set is trained to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model, including: inputting the preset image training set into a first preset model to output a second text and semantic information corresponding to the second text, where the second text is an annotated text obtained by updating keyword weights of a first text; inputting the second text and the semantic information corresponding to the second text into the diffusion generation model to obtain a first type of initial image; calculating a first loss function of a preset text-to-image model according to the first type of initial image and a first image; repeatedly training the preset text-to-image model until the first loss function converges to obtain the preset text-to-image model completed in the first stage of training, and the preset text-to-image model completed in the first stage of training includes the first target model and the second candidate model.
[0095] In this embodiment, first, a suitable text encoder and a diffusion generation model are selected. For example, the CLIP model is selected as the text encoder, and the architecture and sub-modules of the text encoder are determined. The U-Vit model is selected as the diffusion generation model, and the number of layers and sub-modules of the diffusion generation model are determined. The text encoder encodes the text description into an embedding vector, and during the training process, the learning rate, batch size, and other hyperparameters of the text encoder are adjusted to adjust the output embedding vector. The diffusion generation model receives the embedding vector and outputs an image according to the embedding vector. During the training process, the model channel number and other structural parameters of the diffusion generation model are adjusted. In the first training process 301, the text encoder outputs a second text and semantic information (embedding vector) corresponding to the second text. The second text and the semantic information corresponding to the second text are input into the diffusion generation model to obtain a first type of initial image. Among them, the diffusion generation model uses the semantic information corresponding to the second text as noise, and gradually adds noise to the clean noise to generate a white noise image. The white noise image is the first type of initial image. During the training process, by comparing the white noise image and the first image, the parameters of the diffusion generation model are optimized so that the parameters of the diffusion generation model can adapt to the increase in noise, and the rate and shape of noise addition are adjusted.
[0096] In this embodiment, a contrast loss function is used to train the model. The contrast loss function can adopt, for example, the InfoNCE loss function. In the first training process, positive samples (from the same image) and negative samples (from different images) are randomly selected. By calculating the similarity scores between the positive samples and the negative samples, the similarity of the positive samples is maximized and the similarity of the negative samples is minimized. According to the loss function, the training gradient of the preset training model is calculated, and the weights of the model are updated through backpropagation until the contrast loss function converges, and the first stage of training is completed.
[0097] In this embodiment, by training the text encoder and the diffusion generation model simultaneously in the first stage, when the model learns unknown new concepts (which can also be said to be image keywords), the encoding method of the text encoder for the image keyword can be optimized, so that when the concepts trained by the diffusion generation model are generated, higher attention weights can be obtained, and this image keyword can be learned better.
[0098] In one embodiment, in the second training stage 302, training the second candidate model based on a preset image training set to obtain a second target model includes: in the second training stage, obtaining a third text, where the third text is the text with updated annotation weights of the first text in the first training stage; obtaining semantic information corresponding to the third text according to the third text and the first target model; using the semantic information corresponding to the third text and the first image corresponding to the third text as the training set for the second stage; training the second candidate model according to the training set for the second stage to obtain a second target model, and the second target model is used to output the first image corresponding to the third text according to the semantic information of the third text.
[0099] Among them, the third text is the text after the weight of the image keyword is adjusted after the first text is input into the first target model during the first training process. Inputting the third text into the first target model to obtain the semantic information output by the first target model, that is, the text embedding vector. When inputting the third text and its corresponding first image into the second candidate model, the loss function is calculated based on the image output according to the third text and the first image.
[0100] In one embodiment, training the second candidate model according to the training set for the second stage to obtain a second target model includes: inputting the semantic information corresponding to the third text into the second candidate model to obtain a second type of initial image; calculating a second loss function of a preset text-to-image generation model based on the second type of initial image and the first image; repeating the training of the second candidate model until the second loss function converges to obtain a preset text-to-image generation model completed in the first stage. The preset text-to-image generation model completed in the second stage includes the first target model and the second target model.
[0101] In the second training stage of this embodiment: refreeze the text encoder and fine-tune the text encoder until the loss function converges. After the first training stage, the text encoder of the text encoder model has basically mastered the image keywords in the training data, but the model network usually does not fully converge. Therefore, the fine-tuning in the second stage will continue to train the diffusion generation model network based on the fine-tuned text encoder until the overall model training converges. Obtain the first target model and the second target model corresponding to the target text-to-image generation model in the training completion stage 303 as shown in Figure 3 It should be noted that during the process of training the diffusion generation model, the calculated loss function is the overall loss function of the text encoder and the diffusion generation model.
[0102] In this embodiment, by freezing the text encoder in the second stage and training the diffusion generation model, it is possible to prevent overfitting of the preset training model caused by excessive fine-tuning of the text encoder during training, which may lead to the "burning effect" of the model, that is, the generated images have too high contrast and brightness, or "catastrophic forgetting", that is, the model forgets the original knowledge during training.
[0103] Step S104: Obtain the user image text according to the user input instruction, and obtain the target image based on the user image text and the target text-to-image model.
[0104] In step S104, the target text-to-image model extracts image keywords according to the user input instruction and obtains the target image according to the image keywords.
[0105] In one embodiment, obtaining the user image text according to the user input instruction and obtaining the target image according to the user image text and the target text-to-image model includes: finding the image keywords corresponding to the user instruction according to the user input instruction to obtain the user image text; obtaining the first target semantics according to the image text and the first target model, and generating the target image according to the first target semantics and the second target model.
[0106] In this embodiment, obtaining the image keywords includes: inputting the user input instruction into a tokenizer, and dividing each segment in the user input instruction into at least one token by the tokenizer. The tokenizer can be used to split a patch into tokens for representation. Search for adjacent image keywords according to the decomposed tokens. Expand the image keywords to obtain the image text, input the image text into the first target model to obtain the first target semantics, that is, the first text embedding vector, input the first text embedding vector into the second target model, and gradually add noise to the blank image according to the text embedding vector to generate the target image.
[0107] Specifically, as Figure 4 shown, input the user input instruction "a wolf in Chinese ink style", and obtain the image keywords "Chinese ink style" and "wolf". After expansion, limit the quantity to obtain the keyword "one", search for the corresponding artist of the style, such as "Zhang Daqian", and limit the image to obtain the keyword "head portrait", so as to generate the user image text "one", "Chinese ink style", "Zhang Daqian style", "wolf", "head portrait", and obtain the schematic diagram of the user input instruction as Figure 4 shown.
[0108] Specifically, as Figure 5The input user enters the instruction "A cat wearing a gold medal is playing tennis", and obtains the image keywords "a", "gold medal", "cat", "playing tennis". After expansion, search for "cat playing posture" and the background "tennis training ground" to generate "A cat is playing tennis on the tennis court". Based on "playing tennis", generate the action of hitting the tennis ball, "swatting". Search for "gold medal wearing posture", "cat expression", etc. according to the gold medal-related elements, so as to generate the user image text keywords such as "a", "wearing a gold medal", "cat", "on the court", "swatting the tennis ball", "with a serious expression", etc., and input them into the target model to obtain as Figure 5 the schematic diagram of the user input instruction shown.
[0109] The technical solution provided by this application may include the following beneficial effects:
[0110] In the first aspect, this application provides a method for training a preset text-to-image model based on a preset image training set to obtain a target text-to-image model. The preset image training set includes a first image and a first text corresponding to the first image. When using the preset image training set, each first text carries the weight of the image keyword. By adjusting the keyword weight of the first text during the training process, it is convenient for users to identify key features, improves the training efficiency, and at the same time, the weight feature can suppress the data noise generated during the training process and enhance the robustness of the model training.
[0111] In the second aspect, this application includes two training stages when training the preset text-to-image model. In the first training stage, the text encoder and the diffusion generation model are trained simultaneously to accelerate the training efficiency. In the second training stage, the text encoder is frozen and the diffusion generation model is trained, thus effectively avoiding the overfitting of the text encoder, which in turn causes the "burning effect" of the target text-to-image model, that is, the generated image has too high contrast and brightness, or "catastrophic forgetting", that is, the model forgets the original knowledge during the training process. The two-stage training method of this application not only accelerates the convergence speed of the training, but also improves the final generation quality of the model, and at the same time effectively avoids the negative effects of the overfitting of the text encoder.
[0112] As Figure 6 shown, in the second aspect, this application provides an image generation device, including:
[0113] An acquisition unit 610, configured to acquire a preset image training set, where the preset image training set includes a first image and a first text corresponding to the first image;
[0114] A training unit 620 is configured to train a preset text-to-image model according to a preset image training set to obtain a target text-to-image model, which is used to output a first image according to a first text. The preset text-to-image model includes a text encoder and a diffusion generation model. The training unit 620 includes a first training unit 621, a second training unit 622, and a third training unit 623. Training the preset text-to-image model according to the preset image training set to obtain the target text-to-image model includes: The first training unit 621 is configured to, in a first training stage, train the preset image training set based on the text encoder and the diffusion generation model to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model. The second training unit 622 is configured to, in a second training stage, train the second candidate model based on the preset image training set to obtain a second target model. The third training unit 623 is configured to obtain the target text-to-image model according to the first target model and the second target model.
[0115] An image generation unit 630 is configured to obtain a user image text according to a user input instruction, and obtain a target image based on the user image text and the target text-to-image model.
[0116] In one embodiment, the image generation unit 630 includes: a first image subunit 631 and a second image subunit 632. The first image subunit is configured to find an image keyword corresponding to the user instruction according to the user input instruction to obtain the user image text. The second image subunit is configured to obtain a first target semantics according to the image text and the first target model, and generate a target image according to the first target semantics and the second target model.
[0117] In one embodiment, obtaining the preset image training set includes: obtaining an image keyword, where the image keyword is used to describe generating the first image; adding the image keyword to the first text for annotation, and the first text is at least annotated with one image keyword.
[0118] In one embodiment, obtaining the image keyword includes: expanding the first text based on a preset semantic rule to obtain a plurality of first expanded texts; performing semantic supplementation on the first expanded texts based on a preset painting style and / or artist name to obtain second expanded texts; retrieving associated texts of the second expanded texts to obtain third expanded texts associated with the second expanded texts; classifying the third expanded texts and the first images based on a preset classification algorithm to obtain an image keyword corresponding to each first image.
[0119] In one embodiment, image keywords are added to the first text for annotation. The first text is at least annotated with one image keyword, including: arranging the image keywords in accordance with a preset sequence number to obtain the sequence numbers corresponding to the image keywords; obtaining the image keywords associated with the first text, and using the sequence numbers of the image keywords associated with the first text as the annotation of the first text; wherein each image keyword in the annotation of the first text carries an initial weight.
[0120] In one embodiment, in the first training stage, based on a text encoder and a diffusion generation model, a preset image training set is trained to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model, including: inputting the preset image training set into a first preset model to output a second text and the semantic information corresponding to the second text, wherein the second text is the annotated text obtained by updating the keyword weights of the first text; inputting the second text and the semantic information corresponding to the second text into the diffusion generation model to obtain a first type of initial image; calculating a first loss function of a preset text-to-image model according to the first type of initial image and the first image; repeatedly training the preset text-to-image model until the first loss function converges to obtain the preset text-to-image model completed in the first stage of training. The preset text-to-image model completed in the first stage of training includes the first target model and the second candidate model.
[0121] In one embodiment, in the second training stage, the second candidate model is trained based on the preset image training set to obtain a second target model, including: in the second training stage, obtaining a third text, where the third text is the text of the first text with updated annotation weights in the first training stage; obtaining the semantic information corresponding to the third text according to the third text and the first target model; using the semantic information corresponding to the third text and the first image corresponding to the third text as the training set for the second stage; training the second candidate model according to the training set for the second stage to obtain the second target model, and the second target model is used to output the first image corresponding to the third text according to the semantic information of the third text.
[0122] In one embodiment, training the second candidate model according to the training set for the second stage to obtain the second target model, including: inputting the semantic information corresponding to the third text into the second candidate model to obtain a second type of initial image; calculating a second loss function of the preset text-to-image model according to the second type of initial image and the first image; repeatedly training the second candidate model until the second loss function converges to obtain the preset text-to-image model completed in the first stage of training. The preset text-to-image model completed in the second stage of training includes the first target model and the second target model.
[0123] In one embodiment, user image text is obtained according to a user input instruction, and a target image is obtained according to the user image text and a target text-to-image model, including: searching for an image keyword corresponding to the user instruction according to the user input instruction to obtain user image text; obtaining a first target semantics according to the image text and a first target model, and generating a target image according to the first target semantics and a second target model.
[0124] The technical solution provided by this application may include the following beneficial effects:
[0125] In a first aspect, this application provides a method for training a preset text-to-image model with a preset image training set to obtain a target text-to-image model. The preset image training set includes a first image and a first text corresponding to the first image. When using the preset image training set, each first text carries the weight of an image keyword. By adjusting the keyword weight of the first text during the training process, it is convenient for users to identify key features, improving the training efficiency. At the same time, the weight feature can suppress the data noise generated during the training process, enhancing the robustness of model training.
[0126] In a second aspect, when training the preset text-to-image model in this application, it includes two training stages. In the first training stage, the text encoder and the diffusion generation model are trained simultaneously to accelerate the training efficiency. In the second training stage, the text encoder is frozen and the diffusion generation model is trained, effectively avoiding overfitting of the text encoder, which may lead to the "burn-in effect" of the target text-to-image model, that is, the generated image has too high contrast and brightness, or "catastrophic forgetting", that is, the model forgets the original knowledge during the training process. The two-stage training method of this application not only accelerates the convergence speed of training, but also improves the final generation quality of the model, and effectively avoids the negative effects of text encoder overfitting.
[0127] Corresponding to the foregoing method embodiments for implementing application functions, this application also provides an image generation device, an electronic device, a computer-readable storage medium, and corresponding embodiments.
[0128] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the method embodiments related thereto, and will not be elaborated herein again.
[0129] Figure 7 It is a schematic structural diagram of an electronic device shown in an embodiment of this application.
[0130] See Figure 7 , the electronic device 700 includes a memory 710 and a processor 720.
[0131] The processor 720 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0132] The memory 710 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM may store static data or instructions required by the processor 720 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all of the instructions and data required by the processor during operation. In addition, the memory 710 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 710 may include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, ultra-density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.
[0133] An executable code is stored on the memory 710, and when the executable code is processed by the processor 720, it may cause the processor 720 to execute some or all of the methods described above.
[0134] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for performing some or all of the steps in the above method of the present application.
[0135] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device, the processor is caused to execute some or all of the steps of the above method according to the present application.
[0136] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.
Claims
1. A fine-tuning method for a text graph model, characterized in that: include: Acquire a preset image training set, where the preset image training set includes a first image and a first text corresponding to the first image; Training a preset text graph model according to the preset image training set to obtain a target text graph model, wherein the target text graph model is used to output the first image according to the first text; The preset text graph model includes a text encoder and a diffusion generation model, and the training of the preset text graph model according to the preset image training set to obtain a target text graph model includes: in a first training stage, training the preset image training set based on the text encoder and the diffusion generation model to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model; in a second training stage, training the second candidate model based on the preset image training set to obtain a second target model; and obtaining the target text graph model according to the first target model and the second target model. The user image text is acquired according to the user input instruction, and the target image is obtained based on the user image text and the target text-image model.
2. The method according to claim 1, characterized in that The step of obtaining a preset image training set includes: Obtaining image keywords, where the image keywords are used to describe generating a first image; The image keyword is added to the first text for annotation, and the first text is annotated with at least one image keyword.
3. The method according to claim 2, characterized in that The obtaining of image keywords comprises: Expanding the first text based on a preset semantic rule to obtain a plurality of first expanded texts; Based on a preset painting style and / or artist name, semantically supplement the first extended text to obtain a second extended text; Retrieving the associated text of the second extended text to obtain the third extended text associated with the second extended text; Based on a preset classification algorithm, the third extended text and the first image are classified to obtain image keywords corresponding to each first image.
4. The method according to claim 2, characterized in that: The step of adding the image keyword to the first text for annotation, wherein the first text is annotated with at least one image keyword, includes: Arrange the image keywords according to preset sequence numbers to obtain sequence numbers corresponding to the image keywords; Acquire image keywords associated with the first text, and use sequence numbers of the image keywords associated with the first text as labels for the first text; Each image keyword in the annotation of the first text carries an initial weight.
5. The method according to any one of claims 1 to 4, characterized in that: In the first training stage, the preset image training set is trained based on the text encoder and the diffusion generation model to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model, including: Inputting the preset image training set into the first preset model, outputting a second text and semantic information corresponding to the second text, wherein the second text is annotated text obtained after updating the keyword weights of the first text; Inputting the second text and the semantic information corresponding to the second text into the diffusion generation model to obtain a first type of initial image; Calculating a first loss function of the preset Wensheng graph model according to the first type of initial images and the first image; Repeat the training of the preset text graph model until the first loss function converges to obtain the preset text graph model trained in the first stage, wherein the preset text graph model trained in the first stage includes the first target model and the second candidate model.
6. The method according to any one of claims 1 to 4, characterized in that: The second candidate model is trained based on the preset image training set in the second training stage to obtain the second target model, including: In the second training stage, a third text is obtained, where the third text is the text whose annotation weight is updated for the first text in the first training stage; Obtaining semantic information corresponding to the third text according to the third text and the first target model; Using the semantic information corresponding to the third text and the first image corresponding to the third text as a training set for the second stage; The second candidate model is trained according to the training set of the second stage to obtain the second target model, and the second target model is used to output the first image corresponding to the third text according to the semantic information of the third text.
7. The method according to claim 6, characterized in that The step of training the second candidate model according to the training set of the second stage to obtain the second target model includes: Inputting semantic information corresponding to the third text into the second candidate model to obtain a second type of initial image; Calculating a second loss function of the preset text graph model according to the second type of initial image and the first image; Repeat the training of the second candidate model until the second loss function converges to obtain a preset text graph model trained in the second stage, wherein the preset text graph model trained in the second stage includes the first target model and the second target model.
8. The method according to claim 1, characterized in that: The step of acquiring the user image text according to the user input instruction and obtaining the target image based on the user image text and the target text-image model includes: According to the user input instruction, search for image keywords corresponding to the user instruction to obtain user image text; A first target semantics is obtained according to the user image text and the first target model, and a target image is generated according to the first target semantics and the second target model.
9. An image generating device, characterized in that: include: An acquisition unit, configured to acquire a preset image training set, wherein the preset image training set includes a first image and a first text corresponding to the first image; A training unit, configured to train a preset text-graph model according to the preset image training set to obtain a target text-graph model, wherein the target text-graph model is used to output the first image according to the first text; Wherein, the preset text graph model includes a text encoder and a diffusion generation model, the training unit includes a first training unit, a second training unit and a third training unit, and the training unit trains the preset text graph model according to the preset image training set to obtain a target text graph model, including: the first training unit is used to train the preset image training set based on the text encoder and the diffusion generation model in a first training stage to obtain a first target model corresponding to the text encoder and a second candidate model corresponding to the diffusion generation model; the second training unit is used to train the second candidate model based on the preset image training set in a second training stage to obtain a second target model; the third training unit is used to obtain the target text graph model according to the first target model and the second target model; The image generation unit is used to obtain the user image text according to the user input instruction, and obtain the target image based on the user image text and the target text-image model.
10. The device according to claim 9, characterized in that The image generation unit includes: a first image sub-unit and a second image sub-unit; the first image sub-unit is used to search for image keywords corresponding to user instructions according to the user input instructions to obtain user image text; the second image sub-unit is used to obtain first target semantics according to the user image text and the first target model, and generate a target image according to the first target semantics and the second target model.
11. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 8.
12. A computer-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
General feature map acquisition method and related equipment
CN117649578A
Method and system for generating facet defect image sample
CN118247604A