Image generation method and device, electronic device and storage medium
By clustering the training data set and building a plug-in model matrix, dynamically matching and fusion of plug-in models, the problem of users in the existing technology spending a lot of effort to select models is solved, and high-quality text-generated images are achieved.
Patent Information
- Application Number
- CN202410430963.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-04-10
AI Technical Summary
Existing text-generating image technology is difficult to effectively handle users' diverse creative needs, which leads to users spending a lot of effort to select appropriate models from multiple plug-in models, affecting the quality and user experience of generated images.
An image generation method based on dynamic plug-in model matrix is proposed. By clustering, the training data set is constructed, and multiple plug-in models are adaptively matched and fused according to the input text to generate a description image.
It realizes adaptively dynamic selection and integration of plug-in models based on user text description, which significantly improves the quality and user experience of generated images and reduces the energy consumption of users when selecting models.
Smart Images

Figure CN118245627B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an image generating method and device, an electronic device and a storage medium. Background Art
[0002] Text to Image Generation refers to the process of converting text into images using artificial intelligence technology. This technology can output a real image that conforms to the semantic description based on a given input text description by a computer model. Moreover, this technology has a wide range of applications, such as virtual reality, game development, medical image analysis, advertising creativity, etc. With the rapid development of artificial intelligence technology, the text-to-image generation platform still has huge room for improvement in lowering the creation threshold, improving the convenience of services, and enhancing the quality of generated images. Summary of the invention
[0003] The present disclosure proposes a technical solution for image generation.
[0004] According to one aspect of the present disclosure, there is provided an image generation method, comprising: determining a fused plug-in model for an input text according to a plug-in model matrix, the plug-in model matrix comprising N trained plug-in models, N being an integer greater than 1, the fused plug-in model being a fusion result of M plug-in models matched from the plug-in model matrix based on the input text, M being an integer greater than 1 and less than N; inputting the input text into an image generation model including the fused plug-in model to generate a description image of the input text.
[0005] In a possible implementation, the method further includes: obtaining a plug-in model matrix, wherein obtaining the plug-in model matrix includes: obtaining a training data set, the training data set including multiple sample data, each sample data including image-text pair data consisting of a sample image and a text description; clustering the training data set to obtain N training data subsets; and training the plug-in model to be trained based on the N training data subsets to obtain the plug-in model matrix.
[0006] In one possible implementation, clustering is performed on the training data set to obtain N training data subsets, including: performing feature extraction on the training data set to determine a feature vector of each sample data in the training data set; clustering is performed on the feature vector of each sample data in the training data set to obtain N feature vector clusters and a center vector of each feature vector cluster, wherein the center vector is used to indicate the center of the feature vector cluster; and determining N training data subsets and a center vector corresponding to each training data subset based on the N feature vector clusters and the center vector of each feature vector cluster.
[0007] In a possible implementation, each plug-in model corresponds to a center vector, and the center vector is the center vector corresponding to a training data subset obtained by training the plug-in model. The method of determining a fused plug-in model for an input text according to a plug-in model matrix includes: performing feature extraction on the input text to determine a target vector; matching M center vectors from N center vectors according to a similarity between the target vector and each center vector; determining M plug-in models corresponding to the M center vectors from the plug-in model matrix; and performing weighted summation processing on the M plug-in models to obtain a fused plug-in model.
[0008] In a possible implementation, the extracting features from the training data set to determine the feature vector of each sample data in the training data set includes: obtaining a pre-trained feature extraction model, where the feature extraction model is used to map images and texts to a shared feature space; extracting features from the training data set according to the feature extraction model to determine the feature vector of a sample image in each sample data in the training data set;
[0009] In a possible implementation, the extracting features of the input text and determining the target vector includes: obtaining a pre-trained feature extraction model, where the feature extraction model is used to map images and text to a shared feature space; and extracting features of the input text according to the feature extraction model to determine the target vector.
[0010] In a possible implementation, each trained plug-in model is trained by a different subset of training data, and the training process of the plug-in model includes: obtaining a first model based on the plug-in model to be trained and a pre-trained base model; inputting the text description of each sample data in the training data subset into the first model to obtain a predicted image; while keeping the parameters of the base model in the first model unchanged, training the plug-in model in the first model according to the loss of the predicted image and the sample image in the sample data to obtain a trained plug-in model.
[0011] In a possible implementation, obtaining a training data set includes: obtaining an original training data set; performing deduplication processing on the original training data set to obtain a deduplication training data set; and screening the deduplication training data set according to a pre-trained screening model to obtain a training data set.
[0012] According to one aspect of the present disclosure, an image generating device is provided, comprising: a determining module, for determining a fused plug-in model for an input text according to a plug-in model matrix, the plug-in model matrix comprising N trained plug-in models, N being an integer greater than 1, the fused plug-in model being a fusion result of M plug-in models matched from the plug-in model matrix based on the input text, M being an integer greater than 1 and less than N; and a generating module, for inputting the input text into an image generating model including the fused plug-in model, to generate a description image of the input text.
[0013] In a possible implementation, the device also includes an acquisition module: used to obtain a plug-in model matrix, wherein the acquisition module is used to: obtain a training data set, the training data set includes multiple sample data, each sample data includes image-text pair data consisting of a sample image and a text description; cluster the training data set to obtain N training data subsets; and train the plug-in model to be trained based on the N training data subsets to obtain a plug-in model matrix.
[0014] In one possible implementation, clustering is performed on the training data set to obtain N training data subsets, including: performing feature extraction on the training data set to determine a feature vector of each sample data in the training data set; clustering is performed on the feature vector of each sample data in the training data set to obtain N feature vector clusters and a center vector of each feature vector cluster, wherein the center vector is used to indicate the center of the feature vector cluster; and determining N training data subsets and a center vector corresponding to each training data subset based on the N feature vector clusters and the center vector of each feature vector cluster.
[0015] In a possible implementation, the determination module is used to: when each plug-in model corresponds to a central vector, and the central vector is the central vector corresponding to a training data subset obtained by training the plug-in model, perform feature extraction on the input text to determine a target vector; match M central vectors from N central vectors according to the similarity between the target vector and each central vector; determine M plug-in models corresponding to the M central vectors from the plug-in model matrix; and perform weighted sum processing on the M plug-in models to obtain a fused plug-in model.
[0016] In a possible implementation, the extracting features from the training data set to determine the feature vector of each sample data in the training data set includes: obtaining a pre-trained feature extraction model, where the feature extraction model is used to map images and texts to a shared feature space; extracting features from the training data set according to the feature extraction model to determine the feature vector of a sample image in each sample data in the training data set;
[0017] In a possible implementation, the extracting features of the input text and determining the target vector includes: obtaining a pre-trained feature extraction model, where the feature extraction model is used to map images and text to a shared feature space; and extracting features of the input text according to the feature extraction model to determine the target vector.
[0018] In a possible implementation, each trained plug-in model is trained by a different subset of training data, and the training process of the plug-in model includes: obtaining a first model based on the plug-in model to be trained and a pre-trained base model; inputting the text description of each sample data in the training data subset into the first model to obtain a predicted image; while keeping the parameters of the base model in the first model unchanged, training the plug-in model in the first model according to the loss of the predicted image and the sample image in the sample data to obtain a trained plug-in model.
[0019] In a possible implementation, obtaining a training data set includes: obtaining an original training data set; performing deduplication processing on the original training data set to obtain a deduplication training data set; and screening the deduplication training data set according to a pre-trained screening model to obtain a training data set.
[0020] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the above method.
[0021] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor.
[0022] In the disclosed embodiment, a fused plug-in model can be determined for the input text according to the plug-in model matrix, the fused plug-in model being the fusion result of at least one plug-in model matched from the plug-in model matrix based on the input text; the input text is input into an image generation model including the fused plug-in model to generate a description image of the input text. In this way, according to the text description input by the user, multiple plug-in models that meet the text description requirements can be adaptively and dynamically selected from the acquired plug-in model matrix, and the selected multiple plug-in models can be fused to obtain a fused plug-in model. The image generation model including the fused plug-in model can complete a variety of text generation image tasks, support users to freely use any text description for creation, save the energy spent by users in selecting from a large number of plug-in models, and greatly improve the quality of the generated image.
[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used to illustrate the technical solutions of the present disclosure together with the specification.
[0025] Figure 1 A flowchart of an image generating method according to an embodiment of the present disclosure is shown.
[0026] Figure 2 A schematic diagram showing an image generating method according to an embodiment of the present disclosure.
[0027] Figure 3 A schematic diagram showing the effect of the image generating method according to an embodiment of the present disclosure.
[0028] Figure 4 A block diagram of an image generating device according to an embodiment of the present disclosure is shown.
[0029] Figure 5 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0030] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0031] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0032] The term "and / or" herein is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.
[0033] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0034] In related technologies, large models are usually used to convert text descriptions into images. Large models are machine learning models with large-scale parameters and complex computing structures, such as neural network models (NN), diffusion models, stable diffusion models, etc., which contain a large number of parameters. Large models are trained using a large number of text-image pairs. Although they can cover a variety of generated content, there is no guarantee that large models can handle all situations and are still not comprehensive.
[0035] Model plug-ins are a way to extend the functionality of a large model. By adding plug-in models, new features can be added to the large model or existing features can be improved without modifying the code of the large model itself. In addition, the plug-in model only needs a small number of training samples (such as dozens) to train a small number of additional model weights based on the large model, which can quickly achieve a certain generation effect, such as style, characters, objects, postures, etc. The same large model can be used to train plug-in models with different functions. In this way, the same large model can be used with different single or multiple plug-in models to achieve a variety of text generation image tasks, such as generating ink paintings, product photography, Hanfu, blind boxes, portraits, etc.
[0036] Most of the text-generated image platforms in the related art provide multiple trained plug-in models for users to choose from. Users need to refer to the example images of each plug-in model and try multiple times before they can pick out satisfactory results. For example, if you need to achieve a high-quality ink painting style text-generated image task, the related technology adopts a solution from problem to method, manually collecting a small number of high-quality ink painting pictures and training the plug-in model. In this way, when users are faced with a large number of well-provided plug-in models, they are often confused and need to spend a lot of time to find a plug-in model that meets their creative needs from multiple models. Due to the limited number of plug-in models provided, if there is no plug-in model corresponding to the user's needs, the user's drawing task cannot be met. In this case, the plug-in model provided by the platform may determine the user's creative content, rather than the user freely creating any content and automatically using the corresponding plug-in model to enhance the effect. This model violates the most fundamental cause and effect relationship of creation.
[0037] In view of this, the present disclosure proposes an image generation method based on a dynamic plug-in model matrix, which can adaptively and dynamically select multiple plug-in models that meet the text description requirements from the acquired plug-in model matrix according to the user's text description, and fuse the selected multiple plug-in models to obtain a fused plug-in model. Such a fused plug-in model can assist the large model as the base model to generate a new image generation model containing the fused plug-in model, and the image generation model can complete a variety of text generation image tasks and improve the quality of the generated image.
[0038] Figure 1 A flowchart of an image generation method according to an embodiment of the present disclosure is shown. Figure 1 As shown, the image generation method includes: in step S11, according to the plug-in model matrix, a fusion plug-in model is determined for the input text, the plug-in model matrix includes N trained plug-in models, N is an integer greater than 1, and the fusion plug-in model is a fusion result of M plug-in models matched from the plug-in model matrix based on the input text, and M is an integer greater than 1 and less than N.
[0039] In step S12, the input text is input into an image generation model including the fusion plug-in model to generate a description image of the input text.
[0040] For example, the fusion plug-in model can be added to a pre-trained base model to obtain an image generation model that includes the fusion plug-in model. Although the base model can also generate images based on the input text, the images generated by the base model are obviously inferior to the description images generated by the image generation model that includes the fusion plug-in model in terms of content details, light and shadow effects, artistic aesthetics, picture texture and other aspects.
[0041] In a possible implementation, the image generation method may be executed by an electronic device such as a terminal device or a server, the terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, etc., and other processing devices may be a server or a cloud server, etc. In some possible implementations, the image generation method may be implemented by a processor calling a computer-readable instruction stored in a memory. Alternatively, the method may be executed by a server.
[0042] In a possible implementation, the image generation method can be implemented by a processor calling computer-readable instructions stored in a memory. In one example, the processor can be a general-purpose processor such as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), etc., or an artificial intelligence processor such as an artificial intelligence (AI) chip.
[0043] In a possible implementation, the plug-in model matrix may be obtained in step S11. Assuming that the execution subject is an electronic device, the electronic device may obtain the plug-in model matrix in a variety of ways.
[0044] Optionally, the electronic device may obtain the plug-in model matrix through supervised learning of a large amount of training data. For example, the electronic device may first obtain a training data set with a large amount of sample data, divide the training data set into several training data subsets through an unsupervised clustering method, and construct a plug-in model matrix based on the plug-in model trained from each training data subset.
[0045] Optionally, the electronic device may call a program code having a function of generating a plug-in model matrix through an application programming interface (Application Programming Interface, API) to obtain the plug-in model matrix.
[0046] Optionally, the electronic device may directly read a pre-stored plug-in model matrix from a database, or the electronic device may also receive a plug-in model matrix sent by other devices.
[0047] It should be understood that the above method is an illustrative example, and the embodiment of the present disclosure does not limit the method of obtaining the plug-in model matrix.
[0048] In a possible implementation, the plug-in model matrix may include N (N≥1) trained plug-in models, each plug-in model is used to assist the base model to complete a certain drawing task; wherein, the number N of plug-in models included in the plug-in model matrix may be determined by the number of clusters of the training data set, and the present disclosure does not impose any specific restrictions on the number N of plug-in models included in the plug-in model matrix. wherein, the plug-in model may be composed of one or more weight matrices, and the embodiments of the present disclosure do not impose any restrictions on the structure of the plug-in model.
[0049] In a possible implementation, the M (N>M≥1) plug-in models with the most relevant semantics can be adaptively matched for any input text from the N plug-in models included in the plug-in model matrix, and the M plug-in models are fused to obtain a fused plug-in model suitable for the input text. In step S12, the fused plug-in model is added to a pre-trained base model to obtain an image generation model suitable for the input text. The input text is input into the image generation model containing the fused plug-in model to generate a description image of the input text.
[0050] Optionally, the base model may include a large model for generating an image based on the input text, which may include but is not limited to: Neural Network Model (NN), Convolutional Neural Network Model (CNN), Deep Neural Network Model (DNN), Autoregressive Network Model (Autoregressive Model), Diffusion Model, Stable Diffusion Model, etc. The present disclosure does not impose specific restrictions on the model structure of the base model.
[0051] Optionally, the fused plug-in model can be added to the base model in series to obtain an image generation model containing the fused plug-in model. For example, the fused plug-in model can be connected between any two adjacent network layers of the base model; for another example, the fused plug-in model can be connected after the output layer of the base model. Alternatively, the fused plug-in model can be added to the base model in parallel to obtain an image generation model containing the fused plug-in model. For example, the fused plug-in model can be connected in parallel to one or several network layers of the base model. The embodiments of the present disclosure do not limit the manner of adding the fused plug-in model, and can be set according to the actual application scenario.
[0052] In this way, compared with the generated image of the input text directly generated by the base model, the description image generated by the image generation model that includes the fusion plug-in model is significantly superior to the generated image without using the fusion plug-in model in terms of content details, light and shadow effects, artistic aesthetics, and picture texture.
[0053] Moreover, compared with the related art that provides multiple plug-in models for users to choose from, users need to spend time selecting plug-in models, and the user's creative content will also be restricted by the selected plug-in model. It may happen that the plug-in model provided by the platform determines the user's creative content, causing the user to violate the causal order of creation. In contrast, the embodiments of the present disclosure support users to freely use any text description for creation, and can match multiple plug-in models from the plug-in model matrix based on the input text to form a fused plug-in model, saving the user the energy spent on selecting from a large number of plug-in models.
[0054] Figure 2 A schematic diagram of an image generation method according to an embodiment of the present disclosure is shown below. Figure 2 The image generation method of the embodiment of the present disclosure is described in detail by taking FIG.
[0055] In a possible implementation, a fusion plug-in model can be determined for the input text according to the plug-in model matrix in step S11, wherein obtaining the plug-in model matrix may include: obtaining a training data set, wherein the training data set includes a plurality of sample data, each sample data includes image-text pair data consisting of a sample image and a text description; clustering the training data set to obtain N training data subsets and a center vector corresponding to each training data subset; and training the plug-in model to be trained according to the N training data subsets to obtain the plug-in model matrix.
[0056] In a possible implementation, obtaining a training data set includes: obtaining an original training data set; performing deduplication processing on the original training data set to obtain a deduplication training data set; and screening the deduplication training data set according to a pre-trained screening model to obtain a training data set.
[0057] In order to cover a variety of generated content as comprehensively as possible, a massive amount of original training data set is obtained. For example, the original training data set may include millions of sample data, each of which includes a sample image and a corresponding text description, i.e., image-text pair data.
[0058] The training data set is a data set used to train and fit the plug-in model. The purpose is to allow the plug-in model to learn the inherent rules and patterns of the training data set. Deduplication in the training set is to prevent the plug-in model from overfitting due to repeated sample data, thereby improving the generalization ability of the plug-in model. Therefore, after obtaining the original training data set, the original training data set can be deduplicated, and the same sample data can be deleted to obtain a deduplicated training data set.
[0059] After deduplication processing of the original training data set, in order to screen out better quality sample data, a pre-trained screening model can be used to score each sample image in the deduplication training data set, and retain the image-text pair data with a score greater than a preset threshold, so as to screen out a training data set with high artistic aesthetics, rich details, and high-quality image quality from the deduplication training data set.
[0060] The trained screening model can be based on computer vision and deep learning technology, and can predict the aesthetic score corresponding to the image by analyzing the color, composition, details, etc. of the image. The screening model can be trained on a large number of generated images and corresponding manually annotated scoring data sets to obtain a trained screening model. The screening model can be a deep neural network including multiple convolutional layers. The embodiments of the present disclosure do not limit the structure and training method of the screening model.
[0061] In this way, a massive training data set can be obtained, and the training data set includes high-quality data with rich details, which is conducive to training a better plug-in model.
[0062] In one possible implementation, clustering is performed on the training data set to obtain N training data subsets, including: performing feature extraction on the training data set to determine a feature vector of each sample data in the training data set; clustering is performed on the feature vector of each sample data in the training data set to obtain N feature vector clusters and a center vector of each feature vector cluster, wherein the center vector is used to indicate the center of the feature vector cluster; and determining N training data subsets and a center vector corresponding to each training data subset based on the N feature vector clusters and the center vector of each feature vector cluster.
[0063] like Figure 2 As shown, it is known that the training data set includes a large amount of sample data, each sample data includes a sample image and a corresponding text description, and feature extraction can be performed on each sample image in the training data set to extract the visual features of each sample image, such as the color features, shape features, texture features, local features, global features, etc. of the sample image, to obtain a feature vector corresponding to each sample image. The feature vector can be a multi-dimensional numerical vector, for example, the feature vector can be a 768-dimensional numerical vector, and the embodiment of the present disclosure does not specifically limit the dimension of the feature vector.
[0064] In one possible implementation, in order to better represent and analyze each sample data in the training data set, a pre-trained feature extraction model can be obtained, and the feature extraction model is used to map images and texts to a shared feature space. For example, the feature extraction model can use massive images and texts to train data, and can map data in both image and text modalities to a shared feature space, so that images and texts with similar visual content and semantic meanings have closer vector representations in the feature space. In addition, the feature extraction model maps images and texts to a shared feature space, so that the feature extraction model can understand the semantic relationship between images and texts, and realize unsupervised joint learning between images and texts, which can be used for various visual and language tasks.
[0065] The feature extraction model may include an image encoder and a text encoder, the image encoder is used to extract a feature vector of an image, and the text encoder is used to extract a feature vector of a text, wherein the dimension of the feature vector extracted by the image encoder is the same as the dimension of the feature vector extracted by the text encoder. It should be understood that when the feature extraction model satisfies the requirement that images and texts can be mapped to a shared feature space, the embodiments of the present disclosure do not limit the structure of the feature extraction model.
[0066] In a possible implementation, a pre-trained feature extraction model is obtained, and feature extraction may be performed on the training data set according to the feature extraction model to determine a feature vector of a sample image in each sample data in the training data set.
[0067] Since images have more information than text, the same image can correspond to multiple texts. For example, if the image content is a flower, possible text descriptions may include: the color of the flower is red, the flower is in full bloom, the flower is beautiful, and the fragrance is overflowing. It can be seen that for each sample data in the training data set, the sample image will have more information than the text description. In order to make the feature vector more comprehensive and better represent the sample data, the feature extraction model can be used to extract features from the sample image in each sample data, and the feature vector of the sample image can be used as the feature vector of the sample data. Clustering the training data set based on such feature vectors can improve the accuracy of clustering the training data set.
[0068] Optionally, k-means clustering can be performed on the feature vector of each sample data in the training data set to obtain N feature vector clusters (for example, feature vector cluster 1 to feature vector cluster N), and the center vector of each feature vector cluster (for example, center vector 1 to center vector N), that is, center vector 1 can be used to indicate the center of feature vector cluster 1, center vector 2 is used to indicate the center of feature vector cluster 2, and so on. Center vector N can be used to indicate the center of feature vector cluster N.
[0069] It should be understood that in addition to the k-means clustering method, the embodiments of the present disclosure may also adopt other unsupervised clustering methods, such as clustering methods based on graph theory, agglomerative hierarchical clustering, density-based non-parametric clustering (Mean Shift Clustering), density-based clustering methods (Density Based Spatial Clustering of Applications with Noise, DBSCAN), etc. The embodiments of the present disclosure do not limit the specific clustering methods.
[0070] Since each sample image can correspond to a feature vector, the sample data including the sample image and the corresponding text description also correspond to a feature vector for each sample data. According to the one-to-one correspondence between the sample data and the feature vector, the sample data corresponding to each feature vector in the feature vector cluster 1 can be clustered to obtain the training data subset 1, and the center vector 1 of the feature vector cluster 1 is used as the center vector of the training data subset 1; the sample data corresponding to each feature vector in the feature vector cluster 2 is clustered to obtain the training data subset 2, and the center vector 2 of the feature vector cluster 2 is used as the center vector of the training data subset 2; and so on, the sample data corresponding to each feature vector in the feature vector cluster N is clustered to obtain the training data subset N, and the center vector N of the feature vector cluster N is used as the center vector of the training data subset N.
[0071] In this way, based on the feature vectors of each sample data in the training data set, an unsupervised clustering method can be used to cluster millions of sample data into multiple training data subsets. The feature vectors in each training data subset are closer in value, so the corresponding sample images are more similar in visual information such as content and style. For example, most of the images in a certain training data subset are cat images, while most of the images in another training data subset are dog images. Since the unsupervised clustering method can not only output the training data subset to which each feature vector belongs, but also derive the center vector of each category, the dimension of the center vector is the same as the dimension of the feature vector, and the center vector can be used as the summary and representative of the content of the training data subset, that is, the content concentrated in the sample data in the training data subset can be analyzed through the center vector. In this way, it is beneficial to divide the massive high-quality image-text data in the training data set into multiple training data subsets in the content concentration, and it is beneficial to disassemble the large model fine-tuning task based on big data into the plug-in model training task based on small data, thereby improving the efficiency of the image generation method of the embodiment of the present disclosure, so that the generated description image is more in line with human preferences in terms of aesthetics and authenticity.
[0072] like Figure 2 As shown, N training data subsets (for example, training data subset 1 to training data subset N) are obtained, and the plug-in model to be trained can be trained according to these N training data subsets to obtain N trained plug-in models (for example, plug-in model 1 to plug-in model N), and a plug-in model matrix can be constructed according to these N trained plug-in models.
[0073] Since the sample data in each training data subset are more similar in visual information such as content and style, the embodiments of the present disclosure use the sample data (such as image-text pair data) of each training data subset to separately train a corresponding plug-in model. For example, most of the sample images in a certain training data subset are in the style of ink painting. The plug-in model trained using this training data subset can be understood as containing drawing experience and reference works in the style of ink painting. When used, this plug-in model can assist the base model in generating higher quality ink painting images. Compared with the related art that adopts a solution path from problem to method, it is to manually collect training images with the goal of ink painting style. Therefore, the resulting plug-in model has a clear artificial definition, and the expected effect of the plug-in model can be described very clearly in advance. Unlike related technologies, the embodiments of the present disclosure do not establish the goals of each plug-in model in advance. The expected effect of each plug-in model is automatically determined by the image content of the training data subset in the corresponding unsupervised clustering subclass, and the content can be described and analyzed by the center vector of the unsupervised clustering subclass. Therefore, the plug-in model matrix composed of N plug-in models trained by this method in the embodiments of the present disclosure can be regarded as an N×K two-dimensional array. The first dimension N of the plug-in model matrix is the number of categories of unsupervised clustering, and the second dimension K is the dimension of the feature vector, for example, including K=768. The embodiments of the present disclosure do not limit the specific values of N and K, and can be set according to the actual application scenario.
[0074] In a possible implementation, each trained plug-in model is trained by a different subset of training data, and the training process of the plug-in model includes: obtaining a first model based on the plug-in model to be trained and a pre-trained base model; inputting the text description of each sample data in the training data subset into the first model to obtain a predicted image; while keeping the parameters of the base model in the first model unchanged, training the plug-in model in the first model according to the loss of the predicted image and the sample image in the sample data to obtain a trained plug-in model.
[0075] Exemplarily, assuming that the training data subset 1 is used to train the plug-in model, the plug-in model to be trained can be added to the pre-trained base model by a parallel method, a serial method, a learnable prompt method, a hybrid method, etc., to obtain a first model. The size of the plug-in model is much smaller than the base model. For example, the model size of the plug-in model can be 4 to 150 megabytes, and the model size of the base model can be 2 gigabytes. The embodiments of the present disclosure do not limit the specific sizes of the plug-in model and the base model.
[0076] The text description of each sample data in the training data subset 1 is input into the first model to obtain a predicted image; wherein each text description can be input into the first model individually to obtain a predicted image of the current text description; or multiple text descriptions can be input into the first model in batch data to obtain multiple predicted images corresponding to multiple text descriptions, and the embodiments of the present disclosure are not limited to this.
[0077] After obtaining the predicted image of the text description in the sample data, the loss between the predicted image and the sample image can be calculated according to a preset loss function, and the loss is used to indicate the deviation information between the predicted image and the sample image. Among them, the loss function includes, for example, the mean square error function (MSE), the structural similarity metric function (SSIM), etc., and the embodiments of the present disclosure do not limit the type of loss function.
[0078] While keeping the parameters of the base model in the first model unchanged, the model parameters of the plug-in model in the first model are updated using the loss of the predicted image and the sample image until the model training convergence condition is reached. The plug-in model obtained when the model training convergence condition is reached is used as the trained plug-in model 1.
[0079] The above-mentioned condition for reaching the convergence of model training may be that the number of training iteration operations reaches a preset number of training times. Alternatively, the condition for reaching the convergence of model training may also be that the current loss is less than a preset threshold. It should be understood that the preset number of training times and the preset threshold may be pre-set in combination with the training speed and accuracy of the network in actual applications, and the present disclosure does not impose any specific restrictions on this.
[0080] It should be understood that the training methods of other plug-in models 2 to N can refer to the training method of plug-in model 1, and the embodiments of the present disclosure will not be repeated. In addition, in practical applications, N plug-in models can be trained in parallel based on N training data subsets.
[0081] In this way, the local plug-in model parameters are adjusted while keeping most of the parameters of the first model unchanged. Since the gradient of the weight parameters of the first model does not need to be recalculated, the amount of calculation required for training is greatly reduced, and the consumption of hardware resources is reduced.
[0082] like Figure 2 As shown, the plug-in model matrix composed of plug-in model 1 to plug-in model N can correspond to a central vector group, and the plug-in models in the same row have a corresponding relationship with the central vectors. For example, plug-in model 1 corresponds to central vector 1, plug-in model 2 corresponds to central vector 2, and so on. Plug-in model N corresponds to central vector N.
[0083] In a possible implementation, a fused plug-in model can be determined for any input text according to a plug-in model matrix in step S11, wherein the plug-in model matrix includes a plurality of trained plug-in models, each plug-in model corresponds to a center vector, and the center vector is a center vector corresponding to a subset of training data for training the plug-in model. Step S11 may include: performing feature extraction on the input text to determine a target vector; matching M center vectors from N center vectors according to a similarity between the target vector and each center vector; determining M plug-in models corresponding to the M center vectors from the plug-in model matrix; and performing weighted summation processing on the M plug-in models to obtain a fused plug-in model.
[0084] For any input text, in order to match the fusion plug-in model related to its semantic description from the plug-in model matrix, a pre-trained feature extraction model can be obtained, and the feature extraction model is used to map the image and text to a shared feature space; according to the feature extraction model, the input text is feature extracted to determine the target vector. The feature extraction model is the same as the feature extraction model used to extract the feature vector of the sample image in the training data set, which is conducive to improving the accuracy of matching the target vector with the center vector.
[0085] The feature extraction model is used to determine the semantic feature vector of the input text, and the semantic feature vector is used as the target vector, which has the same dimension as the center vector corresponding to each plug-in model. Considering that the feature extraction model can map images and texts to a shared feature space, the cosine similarity between the center vector of the cluster subclass corresponding to each plug-in model and the target vector is calculated, and the M center vectors with the highest similarity can be taken. These M center vectors are closer to the target vector, which can be understood as the content of the corresponding training data subset is closer to the target vector. These M plug-in models can be used for weighted summation processing to obtain a fusion plug-in model.
[0086] For example, assuming the input text is "cute cat, pink, ink painting", the top three plug-in models in similarity after matching correspond to these three contents. Although the base model cannot perform all text-to-image tasks, it can dynamically perform targeted assistance and enhancement based on the plug-in model matrix, and dynamically match different plug-in models for different input texts. Moreover, when using multiple matched plug-in models at the same time, the cosine similarity between the center vector of each plug-in model and the target vector can be summed and normalized to obtain the corresponding weight of each plug-in model, so that the weighted sum of multiple plug-in models can be performed to obtain the fused plug-in model after fusion.
[0087] In this way, through the plug-in matrix, for any input text, by calculating the cosine similarity between the target vector of the input text and the center vectors of each training data subset obtained by unsupervised clustering of the training data set, the most relevant multiple plug-in models are adaptively matched to the input text to form a fusion plug-in model to assist and enhance text generation.
[0088] In step S11, a fusion plug-in model is determined for a certain input text, and in step S12, an image generation model is determined based on the fusion plug-in model and the base model; the input text is input into the image generation model to generate a description image of the input text. In this way, the fusion plug-in model can be used to assist the input text in generating a description image.
[0089] Figure 3 A schematic diagram showing the effect of the image generation method according to an embodiment of the present disclosure is shown as follows: Figure 3 As shown, the first row is an image generated using the relevant technology, and the second row is a description image generated using the image generation method of the embodiment of the present disclosure. By comparing these two rows of images, it can be seen that the image generation method of the embodiment of the present disclosure can adaptively match multiple plug-in models with the most semantic relevance from the plug-in model matrix according to the text description, and obtain a fused plug-in model through weighted fusion. The description image generated based on the fused plug-in model is significantly better than the result of not using the fused plug-in model in terms of content details, light and shadow effects, artistic aesthetics, and picture texture. Furthermore, due to the image generation method of the embodiment of the present disclosure, a plug-in model matrix composed of tens of thousands of plug-in models can be trained on massive data. The plug-in model matrix can provide a higher matching accuracy for any text description, thereby retrieving multiple corresponding plug-in models, and assisting in improving the generation effect of text raw images based on the fusion results of these plug-in models.
[0090] In summary, the embodiments of the present disclosure can perform unsupervised clustering processing on the training data set based on the visual feature vectors of each sample data in the training data set extracted by the feature extraction model, and obtain N training data subsets and the central vector of each training data subset. And according to the N training data subsets, N plug-in models are trained to form a plug-in model matrix. For any input text, according to the target vector representing the semantic feature and the central vector representing the visual feature of each training data subset, the most relevant multiple plug-in models are dynamically matched from the plug-in model matrix, and the multiple plug-in models are fused and used to generate text images.
[0091] In this way, compared with the related art that provides a limited number of plug-in models for users to choose from, users can only passively determine the creation content based on these plug-in models, which has great limitations. The user's creation content is limited by which plug-in models are used, which violates the causal order of creation; in contrast, the embodiments of the present disclosure support users to freely use any text description for creation, and save the energy spent on selecting from a large number of plug-in models. Moreover, compared with the related art that provides a limited number of discrete plug-in models for users to choose from, it is still impossible to achieve comprehensive text generation; in contrast, the embodiments of the present disclosure can construct a plug-in model matrix composed of thousands of plug-in models, making full use of the prior knowledge contained in a large number of high-quality image-text pairs, so that users can freely create any content, and automatically match the corresponding prior knowledge for corresponding effect enhancement. The matching process can be based on numerical vector similarity calculation, which is more accurate and automated.
[0092] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not repeat them. It can be understood by those skilled in the art that in the above-mentioned method of the specific implementation method, the specific execution order of each step should be determined according to its function and possible internal logic.
[0093] In addition, the present disclosure also provides an image generating device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any image generating method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be repeated here.
[0094] Figure 4 A block diagram of an image generating device according to an embodiment of the present disclosure is shown as follows: Figure 4 As shown, the device includes: a determination module 41, used to determine a fused plug-in model for an input text according to a plug-in model matrix, the plug-in model matrix includes N trained plug-in models, N is an integer greater than 1, and the fused plug-in model is a fusion result of M plug-in models matched from the plug-in model matrix based on the input text, M is an integer greater than 1 and less than N; a generation module 42, used to input the input text into an image generation model containing the fused plug-in model to generate a description image of the input text.
[0095] In a possible implementation, the device also includes an acquisition module: used to obtain a plug-in model matrix, wherein the acquisition module is used to: obtain a training data set, the training data set includes multiple sample data, each sample data includes image-text pair data consisting of a sample image and a text description; cluster the training data set to obtain N training data subsets; and train the plug-in model to be trained based on the N training data subsets to obtain a plug-in model matrix.
[0096] In one possible implementation, clustering is performed on the training data set to obtain N training data subsets, including: performing feature extraction on the training data set to determine a feature vector of each sample data in the training data set; clustering is performed on the feature vector of each sample data in the training data set to obtain N feature vector clusters and a center vector of each feature vector cluster, wherein the center vector is used to indicate the center of the feature vector cluster; and determining N training data subsets and a center vector corresponding to each training data subset based on the N feature vector clusters and the center vector of each feature vector cluster.
[0097] In a possible implementation, the determination module 41 is used to: when each plug-in model corresponds to a central vector, and the central vector is the central vector corresponding to a training data subset obtained by training the plug-in model, perform feature extraction on the input text to determine a target vector; match M central vectors from N central vectors according to the similarity between the target vector and each central vector; determine M plug-in models corresponding to the M central vectors from the plug-in model matrix; and perform weighted sum processing on the M plug-in models to obtain a fused plug-in model.
[0098] In a possible implementation, the extracting features from the training data set to determine the feature vector of each sample data in the training data set includes: obtaining a pre-trained feature extraction model, where the feature extraction model is used to map images and texts to a shared feature space; extracting features from the training data set according to the feature extraction model to determine the feature vector of a sample image in each sample data in the training data set;
[0099] In a possible implementation, the extracting features of the input text and determining the target vector includes: obtaining a pre-trained feature extraction model, where the feature extraction model is used to map images and text to a shared feature space; and extracting features of the input text according to the feature extraction model to determine the target vector.
[0100] In a possible implementation, each trained plug-in model is trained by a different subset of training data, and the training process of the plug-in model includes: obtaining a first model based on the plug-in model to be trained and a pre-trained base model; inputting the text description of each sample data in the training data subset into the first model to obtain a predicted image; while keeping the parameters of the base model in the first model unchanged, training the plug-in model in the first model according to the loss of the predicted image and the sample image in the sample data to obtain a trained plug-in model.
[0101] In a possible implementation, obtaining a training data set includes: obtaining an original training data set; performing deduplication processing on the original training data set to obtain a deduplication training data set; and screening the deduplication training data set according to a pre-trained screening model to obtain a training data set.
[0102] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0103] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0104] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.
[0105] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0106] The electronic device may be provided as a terminal, a server, or a device in other forms.
[0107] Figure 5 1 is a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Figure 5, the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0108] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (Mac OS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open source Unix-like operating system (FreeBSD TM ) or similar.
[0109] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0110] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0111] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, (but not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.
[0112] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0113] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0114] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0115] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0116] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0117] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0118] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK) and the like.
[0119] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0120] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0121] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0122] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. An image generation method, characterized in that: The method comprises: Determine a fusion plug-in model for the input text according to a plug-in model matrix, wherein the plug-in model matrix includes N trained plug-in models, where N is an integer greater than 1, and the fusion plug-in model is a fusion result of M plug-in models matched from the plug-in model matrix based on the input text according to a target vector representing a semantic feature of the input text and a center vector representing a visual feature of each training data subset, wherein each trained plug-in model is obtained by training with a different training data subset, and M is an integer greater than 1 and less than N; The input text is input into the image generation model including the fusion plug-in model to generate a description image of the input text.
2. The method according to claim 1, characterized in that The method further includes: obtaining a plug-in model matrix, wherein obtaining the plug-in model matrix includes: Acquire a training data set, wherein the training data set includes a plurality of sample data, each sample data includes image-text pair data consisting of a sample image and a text description; Performing clustering processing on the training data set to obtain N training data subsets; The plug-in model to be trained is trained according to N training data subsets to obtain a plug-in model matrix.
3. The method according to claim 2, characterized in that The training data set is clustered to obtain N training data subsets, including: Performing feature extraction on the training data set to determine a feature vector for each sample data in the training data set; Performing clustering processing on the feature vector of each sample data in the training data set to obtain N feature vector clusters and a central vector of each feature vector cluster, wherein the central vector is used to indicate the center of the feature vector cluster; According to the N feature vector clusters and the central vector of each feature vector cluster, N training data subsets and the central vector corresponding to each training data subset are determined.
4. The method according to claim 1, characterized in that Each plug-in model corresponds to a central vector, and the central vector is the central vector corresponding to the training data subset obtained by training the plug-in model. The method of determining the fusion plug-in model for the input text according to the plug-in model matrix includes: Extracting features from the input text to determine a target vector; According to the similarity between the target vector and each center vector, M center vectors are matched from the N center vectors; Determining M plug-in models corresponding to the M center vectors from the plug-in model matrix; A weighted sum process is performed on the M plug-in models to obtain a fused plug-in model.
5. The method according to claim 3, characterized in that: The step of extracting features from the training data set to determine a feature vector for each sample data in the training data set includes: Obtain a pre-trained feature extraction model, wherein the feature extraction model is used to map images and texts to a shared feature space; According to the feature extraction model, feature extraction is performed on the training data set to determine a feature vector of a sample image in each sample data in the training data set.
6. The method according to claim 4, characterized in that The step of extracting features from the input text and determining a target vector includes: Obtain a pre-trained feature extraction model, wherein the feature extraction model is used to map images and texts to a shared feature space; According to the feature extraction model, feature extraction is performed on the input text to determine a target vector.
7. The method according to any one of claims 2 to 6, characterized in that: The training process of the plug-in model includes: Obtain a first model according to the plug-in model to be trained and the pre-trained base model; Inputting the text description of each sample data in the training data subset into the first model to obtain a predicted image; When the parameters of the base model in the first model remain unchanged, the plug-in model in the first model is trained according to the loss of the predicted image and the sample image in the sample data to obtain a trained plug-in model.
8. An image generating device, characterized in that: The device comprises: A determination module, used to determine a fused plug-in model for an input text according to a plug-in model matrix, wherein the plug-in model matrix includes N trained plug-in models, where N is an integer greater than 1, and the fused plug-in model is a fusion result of M plug-in models matched from the plug-in model matrix based on the input text according to a target vector representing a semantic feature of the input text and a center vector representing a visual feature of each training data subset, wherein each trained plug-in model is obtained by training with a different training data subset, and M is an integer greater than 1 and less than N; The generation module is used to input the input text into the image generation model including the fusion plug-in model to generate a description image of the input text.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image generation method and device based on interaction, electronic equipment and storage medium
CN116306588A
Content quality control method and device, storage medium and electronic equipment
CN117075993A
Image generation method and device, equipment and medium
CN117456033A
Systems and methods for hierarchical text-conditional image generation
US11922550B1