Image generation model training method and device and image generation method
By combining text features and appearance features in the training of image generation models, the problem of low matching between images generated by the image generation model and text prompt words is solved, and the accuracy of image generation is improved.
Patent Information
- Application Number
- CN202510615157.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-23
AI Technical Summary
The images generated by existing image generation models have a low degree of match with the text prompt words, resulting in low accuracy of the generated images.
The image generation model is trained by extracting the features of text prompt words and the appearance features of instance images. The training is combined with text features and appearance features, and the model parameters are adjusted to improve the matching degree of the image generation model.
The matching degree between the generated image and the text prompt words in appearance is improved, which improves the accuracy of image generation.
Smart Images

Figure CN120689694A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to an image generation model training method, device, image generation method, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Text-to-image generation (TAGE) is a cutting-edge technology in the field of artificial intelligence. It allows users to generate corresponding images or paintings based on simple text descriptions. For example, when a user enters a text prompt word into the image generation model, the model extracts features from the text prompt word, obtains the corresponding text features, and then generates an image based on these text features.
[0003] However, in existing text-to-image generation technologies, the images generated by the image generation model may have a low degree of match with the text prompt words, that is, the accuracy of the generated images is low.
[0004] Therefore, it is necessary to provide a new image generation method to solve the above technical problems. Summary of the Invention
[0005] The embodiments of the present application provide an image generation model training method, device, image generation method, electronic device, computer-readable storage medium and computer program product, which can solve the problem of low appearance matching between the generated image and the text prompt word.
[0006] In a first aspect, an embodiment of the present application provides an image generation model training method, comprising:
[0007] Acquire a training set, wherein the training set includes a plurality of sample pairs, wherein the sample pairs include text prompt words and instance images corresponding to the text prompt words;
[0008] For the sample pairs in the training set, extracting features of text prompt words in the sample pairs to obtain text features, and extracting appearance features of instance images in the sample pairs to obtain appearance features;
[0009] The image generation model to be trained is trained according to the text features and the appearance features corresponding to the text features to obtain a trained image generation model.
[0010] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0011] In an embodiment of the present application, after obtaining a training set, for each sample pair in the training set, features of the text prompt words in the sample pair are extracted to obtain text features, and appearance features of the instance images in the sample pair are extracted to obtain appearance features. The image generation model to be trained is then trained based on the text features and the appearance features corresponding to the text features, thereby obtaining a trained image generation model. Because the image generation model to be trained is trained based on the text features and the appearance features corresponding to the text features, the trained image generation model not only considers the text features of the text prompt words but also the appearance features of the text prompt words when generating images, thereby facilitating improved appearance matching between the generated image and the text prompt words, i.e., improving the accuracy of the generated image.
[0012] In a second aspect, an embodiment of the present application provides an image generation method, comprising:
[0013] Receive text messages;
[0014] extracting features of the text information;
[0015] The trained image generation model as described in the first aspect is used to generate images corresponding to the extracted features.
[0016] In a third aspect, an embodiment of the present application provides an image generation model training device, comprising:
[0017] A training set acquisition module, configured to acquire a training set, wherein the training set includes a plurality of sample pairs, each of which includes a text prompt word and an example image corresponding to the text prompt word;
[0018] a feature extraction module configured to extract, for each sample pair in the training set, features of text prompt words in the sample pair to obtain text features, and to extract appearance features of instance images in the sample pair to obtain appearance features;
[0019] The model training module is used to train the image generation model to be trained based on the text features and the appearance features corresponding to the text features to obtain a trained image generation model.
[0020] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect or the second aspect when executing the computer program.
[0021] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the method described in the first aspect or the second aspect.
[0022] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method described in the first or second aspect above.
[0023] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art.
[0025] Figure 1 This is a flowchart of an image generation model training method provided by an embodiment of the present application;
[0026] Figure 2 This is a flowchart of another image generation model training method provided in one embodiment of the present application;
[0027] Figure 3 This is a flowchart of an image generation method provided by an embodiment of the present application;
[0028] Figure 4 This is a structural diagram of an image generation model training device provided by an embodiment of the present application;
[0029] Figure 5 is a structural schematic diagram of an image generating device provided by another embodiment of the present application;
[0030] Figure 6 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0032] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0033] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0034] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0035] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.
[0036] With the development of artificial intelligence technology, image generation models can generate corresponding images based on the text prompts entered by users.
[0037] However, existing image generation models that generate images based on text often fail to reflect the appearance information of the text prompt. For example, if the text prompt is to generate a person wearing red clothes, existing image generation models may generate a person who is not wearing red clothes based on this text prompt. This is because existing image generation models only extract text features from the text prompt and train the image generation model based on these text features. Since the training process relies on text features rather than appearance features, the images generated by the trained image generation model are prone to losing the appearance features of the text prompt during application.
[0038] In order to improve the appearance matching between the image generated by the image generation model and the text prompt word, that is, to improve the accuracy of the generated image, an embodiment of the present application provides an image generation model training method.
[0039] In this training method, since the image generation model to be trained is trained based on the text features and appearance features of the text prompt words, it is beneficial to improve the appearance matching degree between the image generated by the trained image generation model and the text prompt words.
[0040] The image generation model training method provided in the embodiments of the present application is described below with reference to the accompanying drawings.
[0041] Figure 1 A flow chart of an image generation model training method provided in an embodiment of the present application is shown. The training method can be applied to electronic devices, such as servers, and is described in detail as follows:
[0042] S11, obtaining a training set, wherein the training set includes a plurality of sample pairs, and the sample pairs include text prompt words and instance images corresponding to the text prompt words.
[0043] In the embodiments of the present application, the sample pairs in the training set typically include positive sample pairs and negative sample pairs. In the positive sample pairs, the example image matches the text prompt word, while in the negative sample pairs, the example image does not match the text prompt word. The text prompt word can contain specific scene, style, details, emotion, and other information to help the subsequent image generation model better understand the user's intent and thus generate high-quality images.
[0044] For example, if text prompt word 1 in a sample pair is "Generate a person wearing red clothes", the example in example image 1 is "child wearing red clothes", the example in example image 2 is "child wearing red clothes on the beach", and the example in example image 3 is "child wearing red clothes in a shopping mall", then text prompt word 1 can form a positive sample pair with example image 1, the same text prompt word 1 can also form a positive sample pair with example image 2, and the same text prompt word 1 can also form a positive sample pair with example image 3.
[0045] For example, if the text prompt word 2 in a sample pair is "generate a girl wearing red clothes" and the instance in the instance image 4 is "a girl wearing yellow clothes", then the text prompt word 2 can form a negative sample pair with the instance image 4.
[0046] In the embodiment of the present application, in a positive sample pair, the number of instances included in the example image corresponding to the text prompt word is the same as the number of instances included in the text prompt word. For example, when the text prompt word includes two (or more) instances, the number of instances included in the example image corresponding to the text prompt word is also two (or more).
[0047] In an embodiment of the present application, to improve the generalization capability of the trained image generation model, multiple text prompts can be generated. For each text prompt, multiple instance images corresponding to the text prompt are generated. Based on these text prompts and the corresponding instance images, multiple sample pairs are constructed. That is, in the multiple sample pairs constructed in the training set, the same text prompt corresponds to different instance images. By constructing these sample pairs, the probability that the trained image generation model will generate different images when the user inputs the same text prompt is increased, thereby improving the generalization capability of the trained image generation model.
[0048] S12, for the sample pairs in the training set, extracting features of the text prompt words in the sample pairs to obtain text features, and extracting appearance features of the instance images in the sample pairs to obtain appearance features.
[0049] In the embodiment of the present application, text feature refers to the information in the form of a numerical value or vector that can represent the content and semantics of the text extracted from the text prompt word, and the text feature includes one or more of the frequency of word occurrence, topic information, and the dependency relationship between words. For the text prompt word in a sample pair, the text feature of the text prompt word is extracted by a text feature extraction method. Wherein, the text feature extraction method includes a statistical method (such as a bag of words model, word frequency and inverse document frequency (Term Frequency-Inverse Document Frequency, TF-IDF)), a word embedding-based method (such as Word2Vec, bidirectional encoder representation (Bidirectional Encoder Representations from Transformers, BERT)), a deep learning-based method (such as a long short-term memory network (Long Short-Term Memory, LSTM)), etc. In actual situations, other methods can also be selected to extract the text features of text prompt words, which will not be listed here one by one.
[0050] In the embodiments of the present application, the appearance features of an instance image refer to the attributes of the instance image that can be directly observed visually, and the appearance features include one or more of color features, texture features, shape features, and spatial relationship features. For an instance image in a sample pair, the appearance features of the instance image are extracted by an appearance feature extraction method. Among them, the appearance feature extraction method includes: using a color histogram to represent the color features of the instance image, using a directional gradient histogram to represent the texture features of the instance image, using an edge detection method to identify the outline or boundary of the instance image, etc. In actual situations, other methods can also be selected to extract the appearance features of the instance image, which will not be listed here one by one.
[0051] S13, training the image generation model to be trained according to the text features and the appearance features corresponding to the text features to obtain a trained image generation model.
[0052] Among them, the image generation model can be a U-type network: Convolutional Networks for Biomedical Image Segmentation (U-Net), and the image generation model can also be a Generative Adversarial Network (GAN) or other models, which are not limited here.
[0053] In the embodiments of the present application, the image generation model to be trained includes an initial image generation model that has not been trained, and also includes an intermediate image generation model that has been trained but still needs to be trained. After the image generation model to be trained extracts image features from text features, it generates a corresponding image based on the extracted image features.
[0054] In an embodiment of the present application, when the image generation model to be trained is trained, the parameters of the image generation model to be trained are actually adjusted. For example, when the image generation model to be trained is a U-Net model, if the U-Net model includes a cross-attention layer, then when adjusting the parameters of the image generation model to be trained, the parameters of the cross-attention layer are adjusted. Among them, the cross-attention layer calculates the similarity between different image features extracted by the image generation model to be trained and dynamically adjusts the weights of the image features, thereby enhancing the interaction between the image features. When the image generation model to be trained includes a cross-attention layer, it is beneficial to improve the correlation between the image generated by the image generation model to be trained and the text prompt word.
[0055] In an embodiment of the present application, after obtaining a training set, for each sample pair in the training set, features of the text prompt words in the sample pair are extracted to obtain text features, and appearance features of the instance images in the sample pair are extracted to obtain appearance features. The image generation model to be trained is then trained based on the text features and the appearance features corresponding to the text features, thereby obtaining a trained image generation model. Because the image generation model to be trained is trained based on the text features and the appearance features corresponding to the text features, the trained image generation model not only considers the text features of the text prompt words but also the appearance features of the text prompt words when generating images, thereby facilitating improved appearance matching between the generated image and the text prompt words, i.e., improving the accuracy of the generated image.
[0056] In an embodiment of the present application, when training the image generation model to be trained, text features, appearance features, or the text features and the appearance features can be combined to achieve the desired effect. For example, text features and appearance features can be used as inputs to the image generation model to be trained, respectively, and the image generation model to be trained is trained based on the processing results of the image generation model to be trained. For another example, appearance features can be used as inputs to the image generation model to be trained, and text features and the appearance features can be used as inputs to the image generation model to be trained, and the parameters of the image generation model to be trained are adjusted based on the processing results of the two features in the image generation model to be trained.
[0057] In some embodiments, when the text features and the appearance features are used as inputs to the image generation model to be trained, the parameters of the image generation model to be trained can be adjusted by comparing the appearance of the image output by the image generation model to be trained with the appearance of the example image. In this case, the above S13, training the image generation model to be trained based on the text features and the appearance features corresponding to the text features, includes:
[0058] Extracting the appearance features of the instance images in the sample pair to obtain a first appearance feature to be evaluated;
[0059] Generate an image to be evaluated based on the image generation model to be trained, the text features, and the appearance features corresponding to the text features;
[0060] Extracting the appearance features of the image to be evaluated to obtain a second appearance feature to be evaluated;
[0061] An appearance consistency loss is calculated based on the first appearance feature to be evaluated and the second appearance feature to be evaluated, and the image generation model to be trained is trained based on the calculated appearance consistency loss result.
[0062] The above sample pair is a positive sample pair, that is, the instance image in the sample pair matches the text prompt word, that is, the instance image can accurately reflect the appearance characteristics of the text prompt word.
[0063] In the embodiment of the present application, considering that the example image can accurately reflect the appearance features of the text prompt word, the appearance features of the example image can be compared with the appearance features of the image generated by the image generation model to be trained to determine whether the appearance of the image generated by the image generation model to be trained meets the requirements. Specifically, the appearance features of the example image are extracted to obtain a first appearance feature to be evaluated. Furthermore, the text features and the appearance features corresponding to the text features are used as inputs to the image generation model to be trained to obtain an image to be evaluated output by the image generation model, and the appearance features of the image to be evaluated are extracted to obtain a second appearance feature to be evaluated. Finally, the difference between the first appearance feature to be evaluated and the second appearance feature to be evaluated is calculated to calculate the appearance consistency loss. The difference can be represented by the distance between the first appearance feature to be evaluated and the second appearance feature to be evaluated, for example, by the cosine distance between the first appearance feature to be evaluated and the second appearance feature to be evaluated. Of course, it can also be represented by the Euclidean distance between the first appearance feature to be evaluated and the second appearance feature to be evaluated, which is not limited here.
[0064] Optionally, when the difference between the first appearance feature to be evaluated and the second appearance feature to be evaluated (or appearance consistency loss) is represented by the cosine distance between the two appearance features to be evaluated, the difference can be calculated using the following formula:
[0065] L apper =1-Cos(φ(I instance )-φ(G ta ))).
[0066] Among them, L apper represents the appearance consistency loss between the first appearance feature to be evaluated and the second appearance feature to be evaluated, Cos(,) represents the cosine similarity, φ(·) is the encoder for extracting appearance features, and I instance is the instance image, G ta It is an image generated by using appearance features and text features as control conditions of the image generation model to be trained. That is, the encoder with the same structure is used to extract I instance Appearance features and extraction G ta The appearance features of the two extracted appearance features are then calculated, and the cosine similarity of the two extracted appearance features is finally expressed as the appearance consistency loss according to the obtained cosine distance.
[0067] When the appearance consistency loss is large, it indicates that I instance The appearance characteristics and G taThe similarity of the appearance features of the two images is low. At this time, it is necessary to adjust the parameters of the image generation model to be trained. When adjusting the parameters, the backpropagation algorithm (BP) can be used. Among them, the core idea of the BP algorithm is to calculate the gradient of the function of the appearance consistency loss relative to the network parameters through the chain rule, thereby updating the weights and biases of the network in the image generation model to be trained. Of course, in actual situations, other algorithms (such as genetic algorithms) can also be used to adjust the parameters of the image generation model to be trained, which is not limited here.
[0068] In some embodiments, considering the complex relationship between appearance features and text features, that is, the need for the model to automatically learn feature interactions, before training the pre-trained image generation model, the appearance features and text features can be spliced, and then the pre-trained image generation model can be trained using the spliced features. Before splicing, the dimensions of the text features need to be the same as the dimensions of the appearance features. That is, in the above S12, the appearance features of the instance images in the above sample pairs are extracted to obtain the appearance features, including:
[0069] The appearance features of the instance images in the sample pairs are extracted, and the extracted appearance features are dimensionally processed according to the dimensions of the text features to obtain the appearance features having the same dimensions as the text features.
[0070] Specifically, based on the dimension of the text feature, if the dimension of the appearance feature of the extracted instance image is equal to the dimension of the text feature, there is no need to perform dimensionality processing on the appearance feature; if the dimension of the appearance feature of the extracted instance image is larger than the dimension of the text feature, the appearance feature needs to be reduced in dimension, such as through principal component analysis (PCA). Of course, dimensionality reduction can also be performed through other methods, such as linear discriminant analysis (LDA), which is not limited here; if the dimension of the appearance feature of the extracted instance image is smaller than the dimension of the text feature, the appearance feature needs to be increased in dimension (that is, the appearance feature is mapped from the original feature space to a higher-dimensional feature space through some mathematical transformation), such as through feature crossing.
[0071] Once appearance features with the same dimensions as text features are obtained, they can be concatenated with the text features, and the image generation model to be trained can then be trained based on the concatenated features. Since the concatenated appearance and text features are fed as a whole into the image generation model to be trained, this helps the model to better learn the interactive relationship between the appearance and text features, thereby improving the accuracy of the images generated by the trained image generation model.
[0072] In some embodiments, when extracting the appearance features of the example images, a single encoder may be used for extraction, or more than one encoder may be used for extraction and then concatenated. When extracting using at least two different types of encoders, the appearance features of the example images in the sample pairs are extracted, and the extracted appearance features are dimensionally processed according to the dimensions of the text features to obtain the appearance features having the same dimensions as the text features, including:
[0073] Extracting appearance features from the instance images in the sample pairs using at least two different types of encoders;
[0074] The extracted appearance features are concatenated according to the dimensions of the text features to obtain the appearance features with the same dimensions as the text features.
[0075] Specifically, different types of encoders refer to encoders with different functions. When two different types of encoders are used to extract appearance features from the same instance image, the obtained appearance features usually have some differences.
[0076] Optionally, considering that the extracted appearance features will subsequently be used together with the text features to train the image generation model to be trained, the core idea of the Contrastive Language-Image Pre-training (CLIP) encoder is to use a large amount of paired data of images and texts for pre-training to learn the alignment relationship between images and texts, that is, the appearance features extracted by the CLIP encoder have a certain relationship with the text. Therefore, in an embodiment of the present application, the encoder used to extract appearance features includes a CLIP encoder.
[0077] Optionally, to improve the accuracy of the extracted appearance features, an encoder (assuming that the encoder is called an appearance encoder) can be pre-trained, and the trained appearance encoder is used to extract the appearance features of the instance image. The training process can be as follows:
[0078] Collect multiple images of a large number of different instances, such as those taken from different angles, with different lighting conditions, and / or with different degrees of occlusion. Multiple images of the same instance constitute positive samples, while images of different instances constitute negative samples. Use these positive and negative samples to train the appearance encoder to obtain a trained appearance encoder.
[0079] Optionally, the encoder of an embodiment of the present application includes a CLIP encoder and a trained appearance encoder, which are used to extract appearance features from an example image respectively. The two extracted appearance features are then concatenated using a multilayer perceptron (MLP) to obtain a concatenated appearance feature having the same dimension as the text feature corresponding to the example image. The main function of the MLP is to achieve approximation of complex functions and classification or regression tasks by learning the mapping relationship between the two input appearance features and the target output (the output dimension is the same as the text feature dimension).
[0080] In some embodiments, the above S13, training the image generation model to be trained based on the above text features and the above appearance features corresponding to the above text features, includes:
[0081] Extracting image features from the text features according to the image generation model to be trained to obtain a first image feature;
[0082] Extracting image features from the text features and the appearance features corresponding to the text features according to the image generation model to be trained to obtain second image features;
[0083] Calculating a semantic alignment loss based on the first image feature and the second image feature to obtain a semantic alignment loss calculation result;
[0084] The above-mentioned image generation model to be trained is trained according to the above-mentioned semantic alignment loss calculation results.
[0085] Specifically, in the process of the image generation model to be trained generating an image based on the text features, it extracts the corresponding image features from the text features, and adds noise (such as Gaussian noise) to the extracted image features so that the image features after adding noise are close to pure noise, and then restores the image features from the pure noise to obtain the image corresponding to the image features. Among them, the operation of adding noise can be 1 or more times. When noise is added multiple times, the above-mentioned first image feature (or second image feature) can be the image feature obtained after any one of the noise additions, but the first image feature and the second image feature are different image features obtained after the same noise addition. For example, the first image feature is an image feature extracted from the text feature and obtained after the first noise addition, and the second image feature is an image feature extracted from the text feature and the appearance feature and obtained after the first noise addition.
[0086] Since the semantic alignment loss result is obtained after calculating the semantic alignment loss between the first image feature obtained by the image generation model to be trained by extracting image features from text features and the second image feature obtained by the image generation model to be trained by extracting image features from the above-mentioned text features and corresponding appearance features, that is, the semantic alignment loss result reflects the difference in image features extracted by the image generation model to be trained when the input has only text features and when the input has both text features and appearance features. Therefore, training the image generation model to be trained according to the calculation result of the semantic alignment loss is conducive to improving the ability of the trained image generation model to generate image features with more appearance characteristics.
[0087] In some embodiments, considering that the dimension of the image feature may be greater than 1, and different dimensions may include different image features, in order to improve the accuracy of the semantic alignment loss calculation result, the dimension of the image feature may be combined for calculation. In this case, the semantic alignment loss calculation result obtained by calculating the semantic alignment loss based on the first image feature and the second image feature includes:
[0088] A semantic alignment loss is calculated based on the text feature, the first image feature, the second image feature, and the dimension of the first image feature to obtain a semantic alignment loss calculation result.
[0089] Among them, the dimension of the first image feature is the same as the dimension of the second image feature. Calculating the semantic alignment loss in combination with the dimension of the first image feature is equivalent to calculating the semantic alignment loss in combination with the dimension of the second image feature.
[0090] Specifically, the semantic alignment loss calculation result can be calculated using the following formula:
[0091]
[0092] Among them, L align Represents the calculation result of semantic alignment loss, represents the square of the 2-norm. The Softmax function is a function that converts an input vector into a probability distribution. K is a text feature (or an image feature extracted from the text feature without adding noise). d is the dimension of the first image feature or the dimension of the second image feature. Q t represents the first image feature, Q ta represents the second image feature, Indicates Q t The transpose of Indicates Q ta The transpose of .
[0093] In order to more clearly describe the image generation model training method provided in the embodiment of the present application, Figure 2 Provide a description.
[0094] exist Figure 2 In , the top branch (assuming it is the first branch) is equivalent to a diffusion model. The extended model includes a text encoder and an image generation model to be trained. The image generation model to be trained is located at Figure 2 The book-shaped location in the middle.
[0095] exist Figure 2 In the first branch, the image generation model to be trained extracts image features based on the text features, adds noise to the image features, and obtains the first image features (i.e., Q t ), the image generated according to the second image feature is G t .
[0096] In the second branch (i.e., the next branch of the first branch), the text features and the appearance features are spliced to obtain the spliced features. The image generation model to be trained extracts the image features based on the spliced features, and performs noise processing on the image features to obtain the second image features (i.e., Q ta ), the image generated according to the first image feature is G ta .
[0097] In the generated G t and G ta Before, Q t and Q ta Calculate the semantic alignment loss to obtain the corresponding semantic alignment loss calculation result, and adjust the parameters of the image generation model to be trained according to the semantic alignment loss calculation result.
[0098] In addition to training the image generation model to be trained using the semantic alignment loss calculation results, the image generation model to be trained can also be trained using the appearance consistency loss.
[0099] Specifically, a large number of images of different instances are collected, and an encoder is trained based on these images to obtain an appearance encoder.
[0100] Obtain an instance image corresponding to a text prompt word, extract the appearance features of the instance in the instance image through the appearance encoder, and extract the appearance features of the instance in the instance image through the CLIP encoder. The two appearance features extracted by the appearance encoder and the CLIP encoder are spliced through MLP to obtain spliced appearance features. The spliced appearance features will be further spliced with the text features extracted by the text encoder. The post-splicing processing is detailed in the description of the second branch and will not be repeated here.
[0101] In order to make the appearance of the image generated by the image generation model to be trained as consistent as possible with the appearance of the instance image, the appearance features extracted by the appearance encoder for the instance image (assuming they are called the first appearance features to be evaluated) are also compared with the appearance features extracted by the appearance encoder for G ta The extracted appearance features (assuming they are called second appearance features to be evaluated) are subjected to appearance consistency loss calculation, and the image generation model to be trained is trained according to the appearance consistency loss result obtained by the calculation.
[0102] It should be pointed out that in Figure 2 In the embodiment, there are only two encoders for extracting appearance features, namely the CLIP encoder and the appearance encoder. In actual situations, the encoders for extracting appearance features can also be other encoders or other numbers of encoders, which are not limited here.
[0103] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0104] After the image generation model is trained, images corresponding to the text prompt words can be generated according to the trained image generation model.
[0105] Figure 3 A flow chart of an image generation method provided in an embodiment of the present application is shown, and is described in detail as follows:
[0106] S31, receiving text information.
[0107] The text information here is also called text prompts. These text prompts include descriptive information about the image, which guides the image generation model to generate an image that matches the description. These text prompts can include specific information about the scene, style, details, and emotion.
[0108] S32: extracting features of the text information.
[0109] In the screenshot, the features of the text information can be extracted through a text encoder to obtain text features.
[0110] S33, using the trained image generation model to generate an image corresponding to the extracted features.
[0111] Among them, the trained image generation model here is a model trained according to the image generation model training method provided in the embodiment of the present application.
[0112] In the embodiment of the present application, since the trained image generation model is a model trained according to the image generation model training method provided in the embodiment of the present application, and the trained image generation model combines appearance features during the training process, when an image corresponding to the extracted features is generated according to the trained image generation model, it is beneficial to improve the matching degree between the appearance of the generated image and the appearance of the text information, that is, it is beneficial to improve the accuracy of the generated image.
[0113] Corresponding to the image generation model training method described in the above embodiment, Figure 4 A structural block diagram of the image generation model training device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0114] Reference Figure 4 The image generation model training device 4 includes: a training set acquisition module 41, a feature extraction module 42, and a model training module 43.
[0115] A training set acquisition module 41 is configured to acquire a training set, wherein the training set includes a plurality of sample pairs, each of which includes a text prompt word and an example image corresponding to the text prompt word;
[0116] A feature extraction module 42 is configured to extract features of the text prompt words in the sample pairs in the training set to obtain text features, and to extract appearance features of the instance images in the sample pairs to obtain appearance features;
[0117] The model training module 43 is used to train the image generation model to be trained based on the above text features and the above appearance features corresponding to the above text features to obtain a trained image generation model.
[0118] In an embodiment of the present application, after obtaining a training set, for each sample pair in the training set, features of the text prompt words in the sample pair are extracted to obtain text features, and appearance features of the instance images in the sample pair are extracted to obtain appearance features. The image generation model to be trained is then trained based on the text features and the appearance features corresponding to the text features, thereby obtaining a trained image generation model. Because the image generation model to be trained is trained based on the text features and the appearance features corresponding to the text features, the trained image generation model not only considers the text features of the text prompt words but also the appearance features of the text prompt words when generating images, thereby facilitating improved appearance matching between the generated image and the text prompt words, i.e., improving the accuracy of the generated image.
[0119] In some embodiments, when extracting the appearance features of the instance images in the sample pairs to obtain the appearance features, the feature extraction module 42 is specifically configured to:
[0120] The appearance features of the instance images in the sample pairs are extracted, and the extracted appearance features are dimensionally processed according to the dimensions of the text features to obtain the appearance features having the same dimensions as the text features.
[0121] In some embodiments, the extracting of appearance features of the example images in the sample pairs and performing dimension processing on the extracted appearance features according to the dimensions of the text features to obtain the appearance features having the same dimensions as the text features includes:
[0122] Extracting appearance features from the instance images in the sample pairs using at least two different types of encoders;
[0123] The extracted appearance features are concatenated according to the dimensions of the text features to obtain the appearance features with the same dimensions as the text features.
[0124] In some embodiments, when the model training module 43 trains the image generation model to be trained based on the text features and the appearance features corresponding to the text features, it is specifically configured to:
[0125] Extracting the appearance features of the instance images in the sample pair to obtain a first appearance feature to be evaluated;
[0126] Generate an image to be evaluated based on the image generation model to be trained, the text features, and the appearance features corresponding to the text features;
[0127] Extracting the appearance features of the image to be evaluated to obtain a second appearance feature to be evaluated;
[0128] An appearance consistency loss is calculated based on the first appearance feature to be evaluated and the second appearance feature to be evaluated, and the image generation model to be trained is trained based on the calculated appearance consistency loss result.
[0129] In some embodiments, when the model training module 43 trains the image generation model to be trained based on the text features and the appearance features corresponding to the text features, it is specifically configured to:
[0130] Extracting image features from the text features according to the image generation model to be trained to obtain a first image feature;
[0131] Extracting image features from the text features and the appearance features corresponding to the text features according to the image generation model to be trained to obtain second image features;
[0132] Calculating a semantic alignment loss based on the first image feature and the second image feature to obtain a first loss calculation result;
[0133] The image generation model to be trained is trained according to the first loss calculation result.
[0134] In some embodiments, the calculation of the semantic alignment loss based on the first image feature and the second image feature to obtain a first loss calculation result includes:
[0135] A semantic alignment loss is calculated based on the text feature, the first image feature, the second image feature, and the dimension of the first image feature to obtain a first loss calculation result.
[0136] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0137] Corresponding to the image generation method described in the above embodiment, Figure 5 A structural block diagram of an image generating device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0138] Reference Figure 5 The image generation device 5 includes: a text information receiving module 51, a feature extraction module 52 and an image generation module 53.
[0139] The text information receiving module 51 is used to receive text information.
[0140] The feature extraction module 52 is used to extract features of the text information.
[0141] The image generation module 53 is configured to generate an image corresponding to the extracted features using the trained image generation model.
[0142] In the embodiment of the present application, since the trained image generation model is a model trained according to the image generation model training method provided in the embodiment of the present application, and the trained image generation model combines appearance features during the training process, when an image corresponding to the extracted features is generated according to the trained image generation model, it is beneficial to improve the matching degree between the appearance of the generated image and the appearance of the text information, that is, it is beneficial to improve the accuracy of the generated image.
[0143] Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 6 As shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 6 Only one processor is shown in the figure), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60, wherein the processor 60 implements the steps of any of the above-mentioned method embodiments when executing the computer program 62.
[0144] The electronic device 6 can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device can include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that Figure 6 This is merely an example of the electronic device 6 and does not constitute a limitation on the electronic device 6 . The electronic device 6 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device 6 may also include input and output devices, network access devices, etc.
[0145] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0146] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. In other embodiments, the memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 6. Furthermore, the memory 61 may also include both an internal storage unit of the electronic device 6 and an external storage device. The memory 61 is used to store an operating system, an application program, a boot loader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or is about to be output.
[0147] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0148] An embodiment of the present application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps of any of the above-mentioned method embodiments when executing the computer program.
[0149] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0150] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps of the above-mentioned method embodiments when executing the computer program product.
[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0152] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0154] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0155] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0156] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for training an image generation model, characterized in that: include: Acquire a training set, wherein the training set includes a plurality of sample pairs, wherein the sample pairs include text prompt words and instance images corresponding to the text prompt words; For the sample pairs in the training set, extracting features of text prompt words in the sample pairs to obtain text features, and extracting appearance features of instance images in the sample pairs to obtain appearance features; The image generation model to be trained is trained according to the text features and the appearance features corresponding to the text features to obtain a trained image generation model.
2. The image generation model training method according to claim 1, wherein: The extracting the appearance features of the instance images in the sample pair to obtain the appearance features includes: The appearance features of the instance images in the sample pair are extracted, and the extracted appearance features are dimensionally processed according to the dimensions of the text features to obtain the appearance features having the same dimensions as the text features.
3. The image generation model training method according to claim 2, wherein: The step of extracting the appearance features of the instance images in the sample pair and performing dimension processing on the extracted appearance features according to the dimensions of the text features to obtain the appearance features having the same dimensions as the text features includes: Extracting appearance features from instance images in the sample pairs using at least two encoders of different types; The extracted appearance features are spliced according to the dimension of the text feature to obtain the appearance feature with the same dimension as the text feature.
4. The image generation model training method according to claim 1, wherein: The step of training the image generation model to be trained based on the text features and the appearance features corresponding to the text features includes: Extracting appearance features of the instance images in the sample pair to obtain first appearance features to be evaluated; Generate an image to be evaluated based on the image generation model to be trained, the text features, and the appearance features corresponding to the text features; Extracting the appearance feature of the image to be evaluated to obtain a second appearance feature to be evaluated; An appearance consistency loss is calculated according to the first appearance feature to be evaluated and the second appearance feature to be evaluated, and the image generation model to be trained is trained according to the calculated appearance consistency loss result.
5. The image generation model training method according to any one of claims 1 to 4, wherein: The step of training the image generation model to be trained based on the text features and the appearance features corresponding to the text features includes: Extracting image features from the text features according to the image generation model to be trained to obtain first image features; Extracting image features from the text features and the appearance features corresponding to the text features according to the image generation model to be trained to obtain second image features; Calculating a semantic alignment loss based on the first image feature and the second image feature to obtain a semantic alignment loss calculation result; The image generation model to be trained is trained according to the semantic alignment loss calculation result.
6. The image generation model training method according to claim 5, wherein: The calculating of the semantic alignment loss according to the first image feature and the second image feature to obtain a semantic alignment loss calculation result includes: A semantic alignment loss is calculated based on the text feature, the first image feature, the second image feature, and the dimension of the first image feature to obtain a semantic alignment loss calculation result.
7. An image generation method, characterized in that: include: Receive text messages; extracting features of the text information; An image corresponding to the extracted features is generated using the trained image generation model as described in any one of claims 1 to 6.
8. An image generation model training device, characterized in that: include: A training set acquisition module, configured to acquire a training set, wherein the training set includes a plurality of sample pairs, each of which includes a text prompt word and an example image corresponding to the text prompt word; a feature extraction module configured to extract, for each sample pair in the training set, features of text prompt words in the sample pair to obtain text features, and to extract appearance features of instance images in the sample pair to obtain appearance features; The model training module is used to train the image generation model to be trained based on the text features and the appearance features corresponding to the text features to obtain a trained image generation model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, characterized in that The invention comprises a computer program, which, when being executed, enables the method according to any one of claims 1 to 7 to be performed.
Citation Information
Patent Citations
Virtual object face reconstruction method, face reconstruction network training method and device
CN116228943A
Text-to-image generation method based on fine-grained semantic reward
CN116883530A
Zero sample image segmentation model training method and device based on multiple modes
CN117788981A
Image generation model training method and device and style image generation method and device
CN118037872A
Training method for generating picture description model, picture description generation method, device, equipment, medium and program product
CN119580034A