Image generation method, image generation model training method, computing device, computer readable storage medium, and computer program product
By combining the visual language processing unit and the diffusion model of the image generation model, the problem of insufficient consistency between the image generation model and the input text is solved, and high-quality, flexible image generation of multiple control signals is achieved.
Patent Information
- Application Number
- PCT/CN2024/142787
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-31
AI Technical Summary
The existing image generation model has major problems in the consistency between the generated image and the input text, and the introduction of more control signals requires additional models and training, which is not flexible and general enough.
The visual language processing unit of the image generation model performs text description of the target image sample, obtains the target image description text, and uses the target image generation sample data to train the model to ensure the consistency of the generated image and the input text, and combines the visual language model and the diffusion model to generate images.
The image generation model improves the consistency between the generated image and the input text, and realizes flexible input and control of a variety of control signals, and the generated image quality is high and meets user needs.
Smart Images

Figure CN2024142787_31072025_PF_FP_ABST
Abstract
Description
Image generation method, image generation model training method, computing device, computer-readable storage medium, and computer program product
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on January 25, 2024, with application number 202410112463.0 and application name “Image generation method, image generation model training method, computing device, computer-readable storage medium, computer program product”, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The embodiments of the present disclosure relate to the field of computer technology, and in particular to an image generation method, an image generation model training method, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0003] Current image generation models, such as those that generate images from text and those that control image generation using various conditions, often use representation models obtained through contrastive learning, such as clip (Contrastive Language-Image Pre-Training, a pre-training model based on contrasting text-image pairs) to extract text representations. However, the representations obtained through contrastive learning cannot accurately represent longer texts and complex descriptive relationships. At the same time, additional models and training are still required to handle multiple control signals, which is not friendly to new tasks and new scenarios.
[0004] That is, the current image generation model is limited by the expressive power of its own text representation model, and there are still major problems in the consistency between the generated image and the input text. Therefore, there is an urgent need for an image generation method that can improve the consistency between image and text by understanding the image and text. Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide an image generation method. One or more embodiments of the present disclosure also relate to an image generation model training method, an image generation apparatus, an image generation model training apparatus, a computing device, a computer-readable storage medium, a computer program, and a computer program product to address technical deficiencies in the prior art.
[0006] According to a first aspect of an embodiment of the present disclosure, there is provided an image generation method, comprising:
[0007] Determine image generation data; use an image generation model to process the image generation data to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model.
[0008] According to a second aspect of an embodiment of the present disclosure, a method for training an image generation model is provided, comprising:
[0009] Determine a target image sample, and use an image generation model to perform a text description on the target image sample to obtain multiple initial image description texts of the target image sample; determine the target image description text based on the matching relationship between the target image sample and each initial image description text; determine target image generation sample data based on the target image sample and the target image description text; and train the image generation model based on the target image generation sample data.
[0010] According to a third aspect of an embodiment of the present disclosure, there is provided an image generation method, which is applied in a cloud, comprising:
[0011] Receive image generation data sent by the client;
[0012] The image generation data is processed using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model; the target image is displayed to the user through a user interaction interface of the client.
[0013] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image generating apparatus, including:
[0014] The data determination module is configured to determine the image generation data; the image acquisition module is configured to process the image generation data using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on the target image sample and the target image description text, the target image description text is determined based on the matching relationship between the target image sample and the initial image description text, and the initial image description text is obtained by using the image generation model to perform a text description of the target image sample.
[0015] According to a fifth aspect of an embodiment of the present disclosure, there is provided an image generation model training device, comprising:
[0016] The determination module is configured to determine a target image sample and, using an image generation model, perform a text description on the target image sample to obtain a plurality of initial image description texts of the target image sample; the text acquisition module is configured to determine the target image description text based on the matching relationship between the target image sample and each initial image description text; the data determination module is configured to determine the target image generation sample data based on the target image sample and the target image description text; the training module is configured to train and obtain the image generation model based on the target image generation sample data.
[0017] According to a sixth aspect of an embodiment of the present disclosure, there is provided an image generation device, which is applied in a cloud, comprising:
[0018] A receiving module is configured to receive image generation data sent by a client; a determining module is configured to process the image generation data using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model; a display module is configured to display the target image to the user through a user interaction interface of the client.
[0019] According to a seventh aspect of an embodiment of the present disclosure, there is provided a computing device, including:
[0020] A memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image generation method or image generation model training method.
[0021] According to an eighth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image generation method or image generation model training method.
[0022] According to a ninth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned image generation method or image generation model training method.
[0023] An embodiment of the present disclosure provides an image generation method, comprising: determining image generation data; processing the image generation data using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model.
[0024] Based on this, the image generation method uses the text and image understanding capabilities of the image generation model to perform a text description on the target image sample, obtain the initial image description text of the target image sample, and determine the target image description text based on the matching relationship between the target image sample and the initial image description text. The target image description text is consistent with the content contained in the target image sample. Therefore, when the image generation model is obtained by training the target image sample and the target image description text, the image generation data is processed by the image generation model to obtain the target image, which can ensure that the description of the target image is consistent with that of the image generation data, thereby achieving a technical effect of consistency between image and text. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] FIG1 is a schematic diagram of a scenario of an image generation method provided by an embodiment of the present disclosure;
[0026] FIG2 is a flow chart of an image generation method provided by one embodiment of the present disclosure;
[0027] FIG3 is an architecture diagram of an image generation model provided by one embodiment of the present disclosure;
[0028] FIG4 is a flowchart of an image generation method applied to the cloud provided by one embodiment of the present disclosure;
[0029] FIG5 is a flowchart of an image generation model training method provided by one embodiment of the present disclosure;
[0030] FIG6 is a flowchart of a processing process of an image generation model training method provided by one embodiment of the present disclosure;
[0031] FIG7 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure;
[0032] FIG8 is a schematic structural diagram of an image generation model training device provided by one embodiment of the present disclosure;
[0033] FIG9 is a schematic structural diagram of an image generation device applied to a cloud according to an embodiment of the present disclosure;
[0034] FIG10 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0036] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0037] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0038] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0039] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model (Foundation Model), which is pre-trained by using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization capabilities, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.
[0040] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0041] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0042] Diffusion models: A type of generative model used to gradually generate images or videos from noisy samples.
[0043] Latent diffusion models: A diffusion model that uses latent space to achieve generation, which can reduce computational complexity.
[0044] LVLM: Large Visual-Language model, large-scale visual language model.
[0045] MOE:Mixture of Experts, mixed expert model.
[0046] FFN: Feed-Forward Network, fully connected linear layer, refers to the fully connected feedforward neural network layer, which is an important component of the Transformer model.
[0047] Token: The basic unit of model processing. The text and image input to the model will be divided into a sequence of tokens.
[0048] With large-scale data collection and large-scale training of deep learning models, the diffusion model has made significant improvements in image generation tasks, and the generated images are also very beautiful. However, due to the limitations of the expressive power of the text representation model itself, there are still major problems in the consistency between the generated images and the input text. In addition, the introduction of more control signals often requires additional models and additional training, and the method is not universal and flexible enough.
[0049] Multimodal large models have made great progress in their understanding capabilities. They can not only understand language and output knowledge more accurately, but also understand and recognize images. Therefore, multimodal large models can be combined with image generation tasks to generate images that are more in line with requirements and more beautiful. At the same time, multiple control signals can be input and controlled through only one image generation model.
[0050] In the present disclosure, an image generation method is provided. The present disclosure also relates to an image generation model training method, an image generation device, an image generation model training device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0051] Referring to FIG1 , FIG1 shows a schematic diagram of a scenario of an image generation method provided by an embodiment of the present disclosure.
[0052] Specifically, the image generation method is implemented by applying the client 102 and the cloud 104. The client 102 is used to send image generation data to the cloud 104. The image generation data can be only image description text, or an initial image and prompt text; for example, the image description text can be "Generate an image of a puppy running in the sun", the initial image can be the image uploaded by the user through the user interaction interface as shown in Figure 1, and the prompt text is the text entered by the user "Change the style of the sky in the above picture"; in actual application, the user can enter the image description text and prompt text in the client 102 by text or voice. If voice is used, the client 102 will also include corresponding voice processing parts, such as voice analysis, voice-to-text, voice synthesis and other modules, which are used to convert the user's voice into text. The present disclosure does not impose any restrictions on this.
[0053] An image generation model is obtained by training in the cloud 104. Specifically, the image generation model is obtained by training with target image generation sample data. The target image generation sample data is determined by a target image sample and an image description text. That is, the target image sample is determined, and a visual language processing unit of the image generation model is used to perform a text description on the target image sample to obtain an image description text of the target image sample. The image description text is consistent with the content contained in the target image sample. Therefore, when the image generation model is obtained by training with the target image generation sample data, the image generation model can generate a target image consistent with the description of the image generation data after processing the image generation data.
[0054] Specifically, the image generation data is input into the image generation model, and the visual language processing unit of the image generation model is used to obtain the initial feature vector corresponding to the image generation data. The initial feature vector is vector-mapped using the vector mapping layer of the image generation model to obtain the target feature vector corresponding to the image processing unit of the image generation model; the target feature vector is processed using the image processing unit of the image generation model to obtain the target image.
[0055] The target image is displayed to the user through the user interaction interface of the client 102. For example, the target image is the image shown in FIG1 with the sky style in the picture changed, and the text "The sky style in the picture has been changed" can also be displayed to the user to improve the user's interactive experience.
[0056] The client 102 may include a browser, an APP (Application), or a web application such as an H5 (Hyper Text Markup Language 5, version 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client may be based on a software development kit (SDK) of a corresponding service provided by the cloud, such as one developed based on a real-time communication (RTC) SDK. The client may be deployed in an electronic device and may rely on the device to run or on certain APPs in the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, or a personal computer. Various other types of applications may also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0057] Cloud 104 can be understood as a distributed system consisting of multiple servers, storage devices, network equipment, virtualization technology, data centers, and related software and services. Servers include physical servers and cloud servers, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that Cloud 104 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. Cloud 104 can also be a server in a distributed system or a server integrated with blockchain. Cloud 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms, or intelligent cloud computing servers or intelligent cloud hosts with artificial intelligence technology.
[0058] It is worth noting that the image generation method provided in the embodiment of the present disclosure can be executed by the cloud 104. In other embodiments of the present disclosure, the image generation model can be deployed in the client 102, so that the client 102 can also have similar functions as the cloud 104, thereby executing the image generation method provided in the embodiment of the present disclosure; in other embodiments, the image generation method provided in the embodiment of the present disclosure can also be jointly executed by the client 102 and the cloud 104.
[0059] The image generation method provided by the embodiment of the present disclosure utilizes the text and image understanding capabilities of the visual language processing unit in the image generation model to perform a text description on the target image sample and obtain the target image description text of the target image sample. The target image description text is consistent with the content contained in the target image sample. Therefore, when the image generation model is obtained by training the target image sample and the image description text, the image generation data is processed by the image generation model to obtain the target image, which can ensure that the description of the target image is consistent with that of the image generation data, thereby achieving a technical effect of consistency between image and text.
[0060] Referring to FIG. 2 , FIG. 2 shows a flow chart of an image generation method provided by an embodiment of the present disclosure, which specifically includes the following steps.
[0061] Step 202: Determine image generation data.
[0062] The image generation data may be understood as data input by a user through a user interaction interface of a client, including but not limited to text, an initial image, and prompt text related to the initial image.
[0063] Specifically, when the image generation data is text, the text can be used to generate an image whose display content is consistent with the text description, such as the text can be "Generate an image of a lake at the foot of a snow-capped mountain"; or when the image generation data is an initial image and prompt text related to the initial image, the prompt text can be used to edit, modify, and perform other operations on the initial image; for example, if the initial image is "a man in a black suit" and the prompt text is "change the clothes to blue", the "black suit" in the initial image can be changed to "blue suit" based on the prompt text, thereby obtaining an image of "a man in a blue suit".
[0064] In actual applications, the image generation data sent by the user through the user interaction interface of the client can be received, and when the image generation data is determined, the image generation data is input into the image generation model.
[0065] Step 204: Use an image generation model to process the image generation data to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model.
[0066] The image generation model can be understood as a model for generating images; the target image can be understood as an image generated by the image generation model and consistent with the description of the image generation data.
[0067] Specifically, the image generation model is obtained by training target image generation sample data, and the target image generation sample data includes target image samples and image description text of the target image samples; the target image samples can be understood as high-resolution and beautiful images, and the visual language processing unit of the image generation model is used to perform text description on the target image samples to obtain the initial image description text, wherein the visual language processing unit of the image generation model can be understood as the visual language model LVLM, which is a large multimodal model with powerful text and image understanding capabilities, and can understand and process the relationship between text and images. Therefore, the image understanding capability of the visual language processing unit can be used to obtain the initial image description text of the target image sample.
[0068] In practical applications, the visual language processing unit can be used to perform text descriptions on the target image sample multiple times to obtain multiple initial image description texts. The matching model can then be used to obtain the matching relationship between the target image sample and each initial image description text, thereby determining the target image description text from the multiple initial image description texts. The target image description text is consistent with the description of the content displayed in the target image sample. For example, if the target image sample is an image of "giraffe walking on the grassland", the target image description text can be "This image shows a moment of an elegant giraffe on the African grassland. The giraffe is eye-catching with its iconic long neck and spotted fur. It is leisurely strolling on the grassland, showing the harmony and tranquility in nature. In the background, the vast blue sky is boundless, dotted with a few sparse trees, adding a sense of three-dimensionality and layering to the picture. The sun shines through the clouds onto the giraffe, creating a warm and charming light and shadow effect."
[0069] Based on this, the image generation model trained according to the target image samples and the target image description text can obtain the target image whose displayed content is consistent with the description of the image generation data when processing the image generation data.
[0070] In one or more embodiments of the present disclosure, when the target data identifier set corresponding to the image generation data contains the target data identifier, the target feature extraction layer is used for processing without changing the original feature extraction layer in the image generation model. This means that the original ability to understand text and images and provide text responses can be maintained, and the image generation capability can be improved while effectively retaining its original capabilities. The specific implementation method is as follows:
[0071] The method of processing the image generation data using the image generation model to obtain a target image includes:
[0072] Determining a target data identifier set corresponding to the image generation data using the image generation model;
[0073] In a case where the target data identification set includes the target data identification, using a target feature extraction layer to obtain an initial feature vector corresponding to the target data identification set;
[0074] Vector mapping is performed on the initial feature vector to obtain a target feature vector, and a target image is obtained based on the target feature vector.
[0075] Among them, the target data identification set can be understood as the token sequence used to input the feature extraction layer; the target data identification can be understood as the learning image generation tokens (Learn image generation tokens); the target feature extraction layer can be understood as the image generation fully connected linear layer (Image Generation FFN), which is used to process the image generation data containing the image generation intention; the initial feature vector can be understood as the vector containing the features of the image generation data; the target feature vector can be understood as the feature vector used to generate the target image.
[0076] Specifically, the input image generation data is preprocessed by word segmentation, encoding, prediction and other preprocessing operations to obtain a token sequence corresponding to the image generation data for input into the feature extraction layer. Then, the image generation model is used to determine whether the image generation data contains image generation intention. In the case that the image generation data contains image generation intention, the obtained target data identification set contains Learn image generation tokens. If so, the image generation data containing image generation intention is processed using an image generation fully connected linear layer to obtain a feature vector corresponding to the target data identification set containing the image generation data. The feature vector is then vector mapped, such as by linear transformation, dimension change, etc., to a feature vector for generating an image, thereby obtaining a target image corresponding to the image generation data.
[0077] The image generation method provided by the embodiment of the present disclosure can use the target feature extraction layer for processing when the target data identification set corresponding to the image generation data contains the target data identification, thereby improving the image generation capability without changing the original feature extraction layer, that is, maintaining its original capability.
[0078] In one or more embodiments of the present disclosure, the image generation model includes a visual language processing unit. In order to obtain a more specific and detailed data identification set corresponding to the image generation data, thereby obtaining a more accurate initial feature vector corresponding to the image generation data, after determining the initial data identification set corresponding to the image generation data, the initial data identification set is predicted to obtain a target data identification set corresponding to the image generation data. The specific implementation method is as follows:
[0079] The visual language processing unit using the image generation model determines a target data identifier set corresponding to the image generation data, including:
[0080] The determining, using the image generation model, a target data identifier set corresponding to the image generation data includes:
[0081] Determining an initial data identification set corresponding to the image generation data, and inputting the initial data identification set into the visual language processing unit;
[0082] Using the visual language processing unit, predicting the initial data identifier set to obtain a predicted data identifier set;
[0083] A target data identification set corresponding to the image generation data is obtained according to the initial data identification set and the predicted data identification set.
[0084] The initial data identifier set can be understood as the original token sequence corresponding to the image generation data; the predicted data identifier set can be understood as the predicted tokens obtained by the visual language processing unit using autoregressive methods to predict the original tokens, continuously updating its understanding of the original tokens, and adjusting subsequent predictions based on the generated tokens. The target data identifier set can be understood as the target token sequence containing the predicted tokens.
[0085] Specifically, when the image generation data is text, the text tokens corresponding to the image generation data can be words, numbers, punctuation marks, and special characters obtained after word segmentation of the image generation data, and the visual language model can predict the next possible token based on these text tokens using an autoregressive method, thereby obtaining a target token sequence corresponding to the image generation data that contains the predicted tokens, and each token is usually converted into a vector representation; this vector can be generated through word embedding or an embedding layer (Embedding layer), which can obtain the semantic and grammatical information of the token.
[0086] In the case where the image generation data is an initial image and prompt text related to the initial image, for such multimodal input, the specific implementation of the prompt text is similar to the above. Image segmentation and other processing can be performed on the initial image to obtain image tokens corresponding to the initial image; and the prompt text tokens corresponding to the prompt text and the image tokens corresponding to the initial image can be further fused or jointly encoded to obtain a token vector sequence after the prompt text and the initial image are fused to capture the interaction and association between the two modalities.
[0087] For example, the image generation data is an image description text: "A brown Labrador running on the grass". This text can be decomposed into a series of tokens, such as "dog", "brown", "Labrador", "grass" and "running", which constitute the original token sequence.
[0088] The visual language processing unit is used to analyze and predict the original token sequence. For example, it can predict more specific color information (such as "dark brown"), environmental details (such as "sunny grass") or action features (such as "running with open mouth"); thereby obtaining a target data identification set, such as "a dark brown Labrador dog running with its mouth open on sunny grass."
[0089] The image generation method provided by the embodiment of the present disclosure can predict the initial data identification set through the visual language processing unit, thereby obtaining a more specific and detailed target data identification set, so as to subsequently obtain a more accurate initial feature vector based on the target data identification set.
[0090] In one or more embodiments of the present disclosure, the image generation data includes a target text or a prompt text and an initial image, the target text is a text description of the initial image, and the prompt text is used to edit the initial image;
[0091] The determining of a target data identifier set corresponding to the image generation data by using an image generation model includes:
[0092] Determining a first initial data identification set corresponding to the target text, and inputting the first initial data identification set into the visual language processing unit;
[0093] Using the visual language processing unit, predicting the first initial data identifier set to obtain a first predicted data identifier set;
[0094] Obtaining a first target data identification set corresponding to the image generation data according to the first initial data identification set and the first predicted data identification set; or
[0095] Determining a second initial data identification set corresponding to the prompt text and the initial image, and inputting the second initial data identification set into the visual language processing unit;
[0096] Using the visual language processing unit, predicting the second initial data identifier set to obtain a second predicted data identifier set;
[0097] A second target data identification set corresponding to the image generation data is obtained according to the second initial data identification set and the second predicted data identification set.
[0098] The target text can be understood as the text input by the user through the user interaction interface of the client; the prompt text and the initial image can be understood as the image and text input by the user through the user interaction interface of the client, and the prompt text is related to the content of the initial image.
[0099] The first initial data identification set can be understood as the original token sequence corresponding to the target text; the first predicted data identification set can be understood as the predicted token sequence corresponding to the target text; and the first target data identification set can be understood as the target token sequence corresponding to the target text and containing the predicted token.
[0100] The second initial data identification set can be understood as the original token sequence corresponding to the prompt text and the initial image; the first predicted data identification set can be understood as the predicted token sequence corresponding to the prompt text and the initial image; the second target data identification set can be understood as the target token sequence corresponding to the prompt text and the initial image, which contains the predicted tokens.
[0101] Specifically, when the visual language processing unit of the image generation model is a large multimodal model, it can not only understand text and images, but also perform cross-modal understanding and reasoning, that is, understand the relationship between text and image. Therefore, the image generation data can be not only the target text, but also the initial image and prompt text. The target text is pre-processed by word segmentation, prediction, etc. to obtain a first target data identification set corresponding to the target text, or the prompt text is pre-processed by word segmentation, prediction, etc., and the initial image is subjected to image segmentation and other operations to obtain a second target data identification set corresponding to the prompt text and the initial image. The specific implementation can be referred to the above embodiment, which will not be repeated here.
[0102] In the image generation method provided by the embodiment of the present disclosure, the image generation data may include target text, or prompt text and an initial image; thereby achieving the goal of obtaining a target image not only by using the target text, but also by using the initial image and the prompt text.
[0103] In one or more embodiments of the present disclosure, when the image generation data includes target text, or prompt text, and an initial image, it is determined whether the first target data identifier set contains the target data identifier, and whether the second target data identifier set contains the target data identifier, respectively, to obtain a first initial feature vector corresponding to the first target data identifier set, or a second initial feature vector corresponding to the second target data identifier set. The specific implementation method is as follows:
[0104] When the target data identifier set includes the target data identifier, obtaining the initial feature vector corresponding to the target data identifier set by using the target feature extraction layer includes:
[0105] In a case where the first target data identification set includes a target data identification, using a target feature extraction layer in the visual language processing unit to obtain a first initial feature vector corresponding to the first target data identification set; or
[0106] In a case where the second target data identification set includes a target data identification, a target feature extraction layer in the visual language processing unit is used to obtain a second initial feature vector corresponding to the second target data identification set.
[0107] The first initial feature vector can be understood as including the target text of the image generation intention and the corresponding initial feature vector; the second initial feature vector can be understood as including the prompt text of the image generation intention and the initial image and the corresponding initial feature vector.
[0108] For example, in the case where the image generation data is prompt text and an initial image, such as the prompt text is: "Change the picture to a picture of an orange cat playing with a ball of yarn on green grass" and an initial image, the initial image shows a scene of a cat running on green grass; in this case, it can be determined that the image generation data contains the image generation intention, and thus it can be determined that the second target data identifier set contains the target data identifier, so that through the target feature extraction layer, the extracted features of the initial image and the prompt text can be fused together to form a second initial feature vector; the second initial feature vector contains key information about the content of the initial image and the prompt text description.
[0109] The image generation method provided by the embodiment of the present disclosure can utilize the target feature extraction layer to extract text descriptions that can help the image generation model understand the target text, or understand and fuse the prompt text and the information of the initial image, so as to obtain a more accurate initial feature vector when the image generation data contains the image generation intention.
[0110] In one or more embodiments of the present disclosure, when it is determined that the image generation data does not contain the image generation intent, the text feature extraction layer and decoding layer in the visual language processing unit are used to obtain the target text corresponding to the image generation data. The specific implementation method is as follows:
[0111] After determining the target data identification set corresponding to the image generation data, the method further includes:
[0112] When it is determined that the target data identifier set does not include the target data identifier, using the text feature extraction layer in the visual language processing unit to obtain a text feature vector corresponding to the target data identifier set;
[0113] The text feature vector is decoded using a decoding layer in the visual language processing unit to obtain a target text.
[0114] Among them, the text feature extraction layer can be understood as the original feature extraction layer of the visual language model, which is used to obtain the target text corresponding to the image generation data that does not contain the image generation intention; the target text can be understood as the text response corresponding to the user's input.
[0115] Specifically, when the image generation data input by the user does not contain the image generation intention, for example, the text input by the user is "What does a big model usually refer to?", the text feature extraction layer in the visual language processing unit is used to obtain a text feature vector for replying to the user input, and then the decoding layer in the visual language processing unit is used to decode the text feature vector to obtain a text reply corresponding to the user input, such as "Big models usually refer to models in the field of machine learning and artificial intelligence with the following characteristics: 1. A huge number of parameters. These models contain millions, billions or even more parameters, far exceeding traditional machine learning models. 2. Complex structure..."
[0116] Alternatively, the user inputs an initial image and a prompt text, and the prompt text is "Describe this image". In this case, the obtained target text may be a text description of the initial image.
[0117] The image generation method provided by the embodiment of the present disclosure does not change the original feature extraction layer in the visual language processing unit, and can maintain its original ability to understand text and images and provide text replies, and can improve the image generation capability while effectively retaining its original capability.
[0118] In one or more embodiments of the present disclosure, the image generation model includes a vector mapping layer, and the image processing unit is used to generate an image. Therefore, when an initial feature vector is obtained, the vector mapping layer of the image generation model is used to perform vector mapping on the initial feature vector to obtain a target feature vector that matches the image processing unit, so that the target feature vector can be processed by the image processing unit. The specific implementation method is as follows:
[0119] The processing of the initial feature vector to obtain a target feature vector includes:
[0120] The initial feature vector is vector mapped by utilizing the vector mapping layer of the image generation model to obtain a target feature vector corresponding to the image processing unit of the image generation model.
[0121] The vector mapping layer can be understood as a mapping layer that processes the initial feature vector, such as performing linear transformation on the initial feature vector; the image processing unit can be understood as a unit for generating an image, such as a diffusion model.
[0122] Specifically, after obtaining the initial feature vector, the vector mapping layer of the image generation model is used to process the initial feature vector so that it matches the representation of the image processing unit to obtain the target feature vector.
[0123] In practical applications, the initial feature vector output by the visual language processing unit needs to match the representation of the diffusion model to ensure seamless connection between the two processing units. Therefore, through a vector mapping layer, the initial feature vector output by the visual language processing unit is processed by linear transformation, changing the vector dimension, etc. to obtain the target feature vector that can be processed by the image processing unit.
[0124] The image generation method provided by the embodiment of the present disclosure aligns the visual language processing unit and the image processing unit through a vector mapping layer, providing a basis for subsequent image generation and effectively improving the consistency and accuracy of image generation results.
[0125] In one or more embodiments of the present disclosure, vector mapping is performed on the first initial feature vector corresponding to the target text, the initial image, and the second initial feature vector corresponding to the prompt text, thereby obtaining the corresponding first target feature vector and second target feature vector. The specific implementation method is as follows:
[0126] The processing of the initial feature vector to obtain a target feature vector includes:
[0127] Performing vector mapping on the first initial feature vector using a vector mapping layer of the image generation model to obtain a first target feature vector corresponding to the image processing unit; or
[0128] The second initial feature vector is vector mapped by utilizing the vector mapping layer of the image generation model to obtain a second target feature vector corresponding to the image processing unit.
[0129] Specifically, a vector mapping layer is used to perform vector mapping on the first initial feature vector corresponding to the target text, such as performing linear transformation, changing the vector dimension, etc., so as to obtain the first target feature vector corresponding to the image processing unit, so that the image processing unit can obtain the target image corresponding to the target text based on the first target feature vector corresponding to the target text.
[0130] Alternatively, a vector mapping layer is used to perform vector mapping on the initial image and the second initial feature vector corresponding to the prompt text, such as performing linear transformation, changing the vector dimension, etc., so as to obtain a second target feature vector corresponding to the image processing unit, so that the image processing unit can obtain a target image that is edited and modified according to the prompt text based on the second target feature vector.
[0131] The image generation method provided by the embodiment of the present disclosure performs vector mapping on the first initial eigenvector or the second initial eigenvector through a vector mapping layer to obtain a first target eigenvector or a second target eigenvector corresponding to the image processing unit, thereby facilitating the subsequent effective improvement of the consistency and accuracy of image generation.
[0132] In one or more embodiments of the present disclosure, the image processing unit of the image generation model is used to obtain a target image with high quality and high definition. The specific implementation method is as follows:
[0133] The step of obtaining a target image according to the target feature vector includes:
[0134] The target feature vector is processed using the image processing unit of the image generation model to obtain a target image.
[0135] Specifically, the image processing unit of the image generation model is used, that is, the diffusion model is used. For example, a denoiser is used to denoise the target feature vector and a generated random noise image, and the target feature vector is used as a control variable to control the generation result of the diffusion model in a cross-attention manner; after multiple iterations, each iteration will make the image clearer and gradually approach the content described by the target feature vector, thereby obtaining a clear target image; ControlNet (a deep learning architecture) can also be used to train an additional encoder to achieve conditional control of the diffusion model. This additional encoder allows the diffusion model to guide the generation process according to specific input conditions or control signals, thereby achieving more accurate and flexible generation of target images in different modalities (such as images, texts), and the generated target image reflects the characteristics of the target feature vector and is consistent with the description of the image generation data.
[0136] The image generation method provided by the embodiment of the present disclosure can obtain a clear, high-quality target image through the image processing unit of the image generation model.
[0137] In one or more embodiments of the present disclosure, the first target feature vector corresponding to the target text, the initial image, and the second target feature vector corresponding to the prompt text are processed to obtain the corresponding first target image and second target image. The specific implementation method is as follows:
[0138] The step of obtaining a target image according to the target feature vector includes:
[0139] Processing the first target feature vector using an image processing unit of the image generation model to obtain a first target image; or
[0140] The second target feature vector is processed using an image processing unit of the image generation model to obtain a second target image.
[0141] Specifically, after using the image processing unit to process the first target feature vector corresponding to the target text, an image consistent with the target text description is obtained; and after using the image processing unit to process the initial image and the second target feature vector corresponding to the prompt text, an image obtained is obtained by editing and modifying the initial image according to the prompt text.
[0142] For example, when the target text is "a cat sitting on a red sofa", the diffusion model starts with a pure noise image. With each iteration, the noise gradually decreases, and the image content is generated according to the guidance of the first target feature vector corresponding to the target text, until a clear image of a cat sitting on a red sofa is generated.
[0143] When the initial image is a cat sitting on a blue sofa and the prompt text is "Change the blue sofa to a red sofa", the diffusion model restores the initial image and then adjusts the color of the sofa based on the second target feature vector corresponding to the initial image and the prompt text. The blue sofa in the final output image is changed to a red sofa, while the position and shape of the cat remain unchanged.
[0144] The image generation method provided by the embodiment of the present disclosure can generate different images when the image generation data is different, that is, it can generate an image consistent with the target text description, or it can edit and modify the initial image only according to the prompt text. It can meet the different needs of users and has more diverse applicable scenarios.
[0145] In one or more embodiments of the present disclosure, after generating a target image, if it is determined that the image generation data contains instructions for generating a text description for the target image, the image understanding capability of the visual language processing unit can be used to generate a target image description text corresponding to the target image. The specific implementation method is as follows:
[0146] After obtaining the target image, the method further includes:
[0147] When it is determined that the image generation data includes a text description generation instruction for the target image, the target image description text corresponding to the target image is generated by using the visual language processing unit.
[0148] Among them, the text description generation instruction can be understood as an instruction for textually describing the generated target image; the target image description text can be understood as the description text corresponding to the target image, which may include information such as "location, main objects, and what is being done".
[0149] Specifically, when the user inputs image generation data, he or she can also inject text description generation instructions for the target image into the image generation data. The text description generation instructions can be sent in the form of input text. For example, if the image generation data input by the user includes a text description generation instruction such as "After generating the picture, add a paragraph of text to the picture", the text description generation instruction can also be sent by clicking the prompt button on the user interaction interface. There is no limitation here.
[0150] When it is determined that the image generation data includes a text description generation instruction for the target image, the visual language processing unit is used to extract features of the target image and generate a target image description text corresponding to the target image.
[0151] For example, when the target image is a picture of a beautiful beach sunset, the generated target image description text can be "On the beach of time, sunset is a gentle farewell to every day; golden light weaves romance, and waves gently tell stories."
[0152] The image generation method provided by the embodiment of the present disclosure can use the image understanding ability of the visual language processing unit to describe the generated target image, thereby generating a reply with both pictures and text for the user, meeting the diverse needs of users and supporting diverse application scenarios.
[0153] In one or more embodiments of the present disclosure, to improve efficiency and save resources, after generating a target image, the target image may be displayed to the user. If the user is satisfied with the target image, an instruction to generate a text description of the target image may be sent. The specific implementation method is as follows:
[0154] After obtaining the target image, the method further includes:
[0155] Displaying the target image to the user through the user interaction interface of the client;
[0156] When a text description generation instruction for the target image is received from the user, the visual language processing unit is used to generate a target image description text corresponding to the target image.
[0157] Specifically, when a target image is obtained, the target image can be first displayed to the user through the user interaction interface of the client, so that when the user is satisfied with the target image, the user can send a text description generation instruction for the target image to the server through the client. Based on this, it is avoided that when the user is not satisfied with the target image, the target image description text is generated, which wastes computer resources and reduces processing efficiency.
[0158] The image generation method provided by the embodiment of the present disclosure can first display the target image to the user through the user interaction interface of the client, and then receive the text description generation instruction for the target image sent by the user, which can reduce the interaction process with the user and improve efficiency.
[0159] In one or more embodiments of the present disclosure, after generating a target image description text corresponding to a target image, the target image and the target image description text can be displayed to the user together, so that the user can obtain a response that meets the user's needs and matches the image and text. The specific implementation method is as follows:
[0160] After generating the target image description text corresponding to the target image, the method further includes:
[0161] The target image and the target image description text corresponding to the target image are displayed to the user through the user interaction interface of the client.
[0162] Specifically, after obtaining the target image description text corresponding to the target image, the target image and the target image description text may be displayed to the user through a user interaction interface of the client.
[0163] The image generation method provided by the embodiment of the present disclosure displays the target image and the target image description text to the user together, which enables the user to clearly obtain a response to the image and text matching, facilitates the user to operate the target image and the target image description text, and improves visual appeal.
[0164] In one or more embodiments of the present disclosure, the image generation model can also be updated and trained based on the target image and the target image description text, so that the trained image generation model can understand different dialogue scenarios and control instructions and generate corresponding images and text accordingly. The specific implementation method is as follows:
[0165] The image generation model is updated and trained according to the target image and the target image description text.
[0166] Specifically, the image generation model is updated and trained based on the real dialogue scenario with the user, the obtained target image, and the target image description text; in actual applications, the image generation model can be updated and trained based on the image generation data input by the user, a series of control instructions input by the user through dialogue, and the target image and target image description text, so that the image generation model can understand different dialogue scenarios and control instructions.
[0167] The image generation method provided by the embodiment of the present disclosure can adjust the image generation model according to the actual data of interaction with the user, covering the diverse control of image text generation tasks, thereby supporting more diverse user needs and application scenarios and improving user experience.
[0168] When the visual language processing unit is a visual language model LVLM, the dialogue, text, and image understanding capabilities of the visual language model can be utilized to understand multiple control signals input by the user. Multiple control signals can be understood as influencing and controlling specific attributes or features of the generated image through inputs of multiple different types or sources in the image generation task. For example, control signals include text prompts, composition guidance, scene types, etc.
[0169] In practical applications, the image generation data input by the user can include control signals such as layout, proportion, scene type, emotional color, etc., so as to accurately respond to the user's intentions and needs based on various control signals and the text and image understanding capabilities of the visual language model; of course, the user's interactive feedback is also a control signal. For example, when the user receives the target image, the user can gradually adjust certain attributes of the target image or provide more specific guidance through dialogue.
[0170] By combining multiple control signals with visual language models, the image generation model can respond to user intentions and needs more flexibly and accurately, generating images that are both compliant and beautiful. It can achieve the input and control of multiple control signals through a single image generation model, completing diverse and general image generation capabilities.
[0171] The image generation method provided by the embodiment of the present disclosure utilizes the text and image understanding capabilities of the visual language processing unit in the image generation model, and can obtain a more accurate initial feature vector of the image generation data through the target feature extraction layer. The initial feature vector of the image generation data is processed through the vector mapping layer of the image generation model to obtain the image processing unit of the image generation model and the corresponding target feature vector, thereby realizing the use of the image processing unit to process the target feature vector and obtain a target image that is consistent with the description of the image generation data and has high quality.
[0172] Referring to FIG3 , FIG3 shows an architecture diagram of an image generation model provided by an embodiment of the present disclosure.
[0173] Specifically, the architecture of the image generation model is described in detail by taking the image generation data as the initial image and the prompt text as an example.
[0174] The architecture of the image generation model provided by the embodiments of the present disclosure samples the structure of MOE, and includes a visual language model and a diffusion model; wherein the visual language model can be understood as the visual language processing unit in the above embodiments, and the diffusion model can be understood as the image processing unit in the above embodiments; MOE is a deep learning architecture in which multiple expert networks exist in parallel, and a gating network (also called a routing network or controller) determines which expert network should be used in a specific situation; this architecture allows the model to process complex and diverse input data; based on this, when the input image generation data contains image generation intent, the diffusion model can be called for image generation based on the visual language model.
[0175] A fully connected linear layer for image generation is introduced into the visual language model to learn image generation capabilities, while other parameters of the visual language model remain fixed. This allows the image generation capabilities of the image generation model to be improved while effectively retaining the original conversation and comprehension capabilities of the visual language model. Specifically, a fully connected linear layer for image generation is added after each self-attention layer of the visual language model.
[0176] In practical applications, the initial image and prompt text are input into the visual language model to obtain tokens of the initial image and prompt text. When it is determined that the tokens of the initial image and prompt text are generated by learning the image, the tokens of the initial image and prompt text are passed through the self-attention layer and the image generation fully connected linear layer to obtain the initial feature vectors corresponding to the initial image and prompt text, and the initial feature vectors are vector-mapped. The mapped target feature vectors are input into the diffusion model. The target feature vectors can be used as control vectors to control the generation results of the diffusion model in a cross-attention manner, thereby obtaining the target image.
[0177] The image generation method provided by the embodiments of the present disclosure, with the help of the text and image understanding capabilities of the visual language model, can understand user needs through dialogue, can generate images that are more consistent with text descriptions, and can edit and generate images according to the method of multiple rounds of natural language dialogue, thereby improving the consistency between images and texts; when the recognition and understanding scope of the visual language model for images is not limited to depth maps, skeleton maps, contour maps, etc., a wider range of control signal understanding can be achieved. In addition, combined with the text itself, multiple cross-modal signal controls can also be achieved. Therefore, in terms of generation, not only image generation can be achieved, but also the composite generation of images and texts can be achieved.
[0178] Referring to FIG4 , FIG4 shows a flowchart of an image generation method applied in the cloud provided by an embodiment of the present disclosure, which specifically includes the following steps.
[0179] Step 402: Receive image generation data sent by the client.
[0180] Step 404: Process the image generation data using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model.
[0181] Step 406: Display the target image to the user through the user interaction interface of the client.
[0182] The specific implementation method can be found in the above embodiments and will not be described again here.
[0183] The image generation method applied to the cloud provided by the embodiment of the present disclosure receives image generation data sent by the client, utilizes the text and image understanding capabilities of the visual language processing unit in the image generation model, and obtains a more accurate initial feature vector of the image generation data through the target feature extraction layer, and obtains the image processing unit and the corresponding target feature vector of the image generation model through the vector mapping layer of the image generation model, thereby realizing the use of the image processing unit to obtain a target image that is consistent with the description of the image generation data and has high quality, and presents the target image to the user through the user interaction interface, thereby improving the user's interactive experience.
[0184] The above is a schematic diagram of an image generation method applied to the cloud in accordance with this embodiment. It should be noted that the technical solution of this image generation method applied to the cloud shares the same concept as the technical solution of the aforementioned image generation method. For details not described in detail in the technical solution of the image generation method applied to the cloud, please refer to the description of the technical solution of the aforementioned image generation method.
[0185] Referring to FIG5 , FIG5 shows a flowchart of an image generation model training provided by an embodiment of the present disclosure, which specifically includes the following steps.
[0186] Step 502: Determine a target image sample, and use an image generation model to perform a text description on the target image sample to obtain a plurality of initial image description texts of the target image sample.
[0187] The target image sample can be understood as a high-resolution, beautiful image; the initial description text can be understood as a description text that uses an image generation model to provide a textual description of the target image sample.
[0188] On the basis that the target image sample has high resolution and beautiful appearance, when the image generation model is trained using the target image sample, it can generate a target image with high resolution and beautiful appearance.
[0189] Specifically, the text and image understanding capabilities of the visual language processing unit in the image generation model are used to perform text descriptions on clear and beautiful target image samples to obtain multiple initial image description texts of the target image samples.
[0190] Step 504: Determine the target image description text according to the matching relationship between the target image sample and each initial image description text.
[0191] The matching relationship can be understood as whether the target image sample and the initial image description text are consistent.
[0192] Specifically, the matching degree between the target image sample and each initial image description text can be determined. The matching degree is used to indicate the degree of description consistency between the target image sample and the initial image description text. The initial image description texts are sorted from high to low according to the matching degree, and the first initial image description text in the sequence is determined as the target image description text.
[0193] Step 506: Determine target image generation sample data based on the target image sample and the target image description text.
[0194] The target image generation sample data can be understood as training data for the image generation model.
[0195] Specifically, after obtaining the image description text of the target image sample, target image generation sample data can be determined based on the target image sample and the image description text, which are training data with consistent images and texts.
[0196] In one or more embodiments, an image generation model may be trained using text and images consistent with the text description, so that the image generation model can generate images using the text; an image generation model may also be trained using text and associated images, or images consistent with the text description, so that the image generation model can generate another image using the text and the image. Specific implementation methods are as follows:
[0197] The step of determining target image generation sample data based on the target image sample and the target image description text includes:
[0198] Using the target image description text as an image generation training sample, using the target image sample as an image generation training label, and determining target image generation sample data based on the image generation training sample and the image generation training label; or
[0199] Determining an associated image sample corresponding to the target image sample;
[0200] The target image description text and the associated image sample are used as image generation training samples, the target image sample is used as an image generation training label, and target image generation sample data is determined based on the image generation training sample and the image generation training label.
[0201] Among them, the associated image sample can be understood as an image that is similar to the target image sample, and can be an image with a certain part of the target image sample changed. For example, when the target image sample is "a man in a black suit", the associated image sample can be "a man in a blue suit", and the color of the clothes in the target image sample is changed.
[0202] Specifically, the image description text and the target image samples corresponding to the image description text are used as training data to train the image generation model, wherein the image description text is used as the image generation training sample and the target image sample is used as the image generation training label, so that the image generation model can generate an image consistent with the text description based on the text in subsequent applications.
[0203] Alternatively, an associated image sample corresponding to the target image sample is determined, and the image generation model is trained using the image description text, the associated image sample, and the target image sample, so that the image generation model can generate images based on the text and the image, and edit and modify the image based on the text in subsequent applications.
[0204] The image generation model training method provided by the embodiment of the present disclosure can use a variety of training data to train the image generation model so that the image generation model can be applicable to more scenarios in subsequent applications, that is, it can not only use text to generate images, but also use text and images to generate images that have been edited and modified.
[0205] Step 508: Generate sample data according to the target image, and train to obtain the image generation model.
[0206] In one or more embodiments of the present disclosure, when it is necessary to ensure alignment between the visual language processing unit and the image processing unit, a large amount of training data is required to train the initial image generation model. The training data can be obtained from network resources or public databases to ensure the training efficiency of the initial image generation model. The specific implementation method is as follows:
[0207] Before determining the target image sample, the method further includes:
[0208] Determine the initial image generation sample data and the initial image sample label;
[0209] Determine a target sample data identifier set corresponding to the initial image generation sample data using the initial image generation model;
[0210] In a case where the target sample data identification set includes the target sample data identification, using a target feature extraction layer to obtain an initial feature vector corresponding to the target sample data identification set;
[0211] Performing vector mapping on the initial feature vector to obtain a target feature vector, and obtaining a target sample image based on the target feature vector;
[0212] The initial image generation model is trained according to the target sample image and the initial image sample label to obtain the image generation model.
[0213] The initial image generation sample data can be understood as sample data including text, or text and images; the initial image sample label can be understood as the image corresponding to the first image generation sample data. Specifically, the initial image generation sample data and initial image sample label can be obtained from online resources or public databases, and are not limited here. The initial image generation model can be understood as an untrained model architecture that includes a visual language processing unit and an image processing unit.
[0214] Specifically, the visual language processing unit of the initial image generation model is used to determine the target sample data identification set corresponding to the initial image generation sample data through the data identification prediction layer in the visual language processing unit; when it is determined that the target sample data identification set contains the target sample data identification, the target feature extraction layer in the visual language processing unit is used to obtain the initial feature vector corresponding to the target sample data identification set; the vector mapping layer of the image generation model is used to perform vector mapping on the initial feature vector to obtain the target feature vector, and based on the target feature vector, the target sample image is obtained, and based on the target sample image and the initial image sample label, the initial image generation model is trained to obtain the image generation model.
[0215] In one or more embodiments of the present disclosure, to ensure efficient training of the image generation model, when alignment between the visual language processing unit and the image processing unit is required, only the data identification prediction layer and the vector mapping layer are trained, and other model parameters are frozen. The specific implementation is as follows:
[0216] The step of training the image generation model according to the target sample image and the initial image sample label includes:
[0217] Freezing model parameters in the image generation model except for the data identification prediction layer and the vector mapping layer;
[0218] The data identification prediction layer and the vector mapping layer in the image generation model are trained according to the target sample image and the initial image sample label.
[0219] Specifically, when training the initial image generation model, a large amount of training data is required. When the training data is obtained through different channels, the quality of the training data is also uneven. Therefore, at this time, it is sufficient to ensure that the visual language processing unit and the image processing unit are aligned, and the generation quality of the target image is not required; that is, at this time, it is possible to ensure that when the initial image generation sample data contains the image generation intention, the target sample data identification set determined by the data identification prediction layer must include the target sample data identification, that is, it can be identified that the first image generation sample data contains the image generation intention; and through the vector mapping layer, the target feature vector obtained can be processed by the image processing unit.
[0220] In practical applications, when training the initial image generation model, all model parameters except the data identification prediction layer and the vector mapping layer in the image generation model are frozen. The data identification prediction layer and the vector mapping layer in the initial image generation model are trained according to the target sample image and the initial image sample label. That is, the only parameters trained are the image generation tokens and the mapped target feature vector.
[0221] In the image generation model training method provided by the embodiment of the present disclosure, the alignment between the visual language processing unit and the image processing unit is the basis for subsequent image and text generation. Therefore, the data identification prediction layer and the vector mapping layer in the initial image generation model are first trained. By freezing some model parameters, the training efficiency of the initial image generation model can be improved when there is a lot of training data.
[0222] In one or more embodiments of the present disclosure, to improve the training effect of the image generation model, it is necessary to clean and filter the training data to select high-quality training data, thereby obtaining an image generation model with better training effect based on the high-quality training data. The specific implementation method is as follows:
[0223] The step of determining the target image description text according to the matching relationship between the target image sample and each initial image description text includes:
[0224] Using a matching model, determining a matching relationship between the target image sample and each initial image description text;
[0225] According to the matching relationship, the target image description text of the target image sample is determined from the multiple initial image description texts.
[0226] Among them, the target image sample can be understood as a high-resolution, beautiful image selected from the above-mentioned initial image sample labels; the matching model can be understood as a clip model, which is used to match text and images; the image description text of the target image sample can be understood as the description text in the initial description text that is consistent with the description of the target image sample.
[0227] Specifically, in order to improve the generation effect of the target image of the image generation model, it is necessary to improve the quality of the initial image sample labels. In practical applications, high-resolution and beautiful images in the initial image sample labels can be first determined as target image samples.
[0228] In practical applications, Dalle3 (a deep learning model) can achieve fine-grained control over generated images by deeply understanding and learning the complex relationship between natural language and vision, ensuring that the output target image accurately reflects the detailed requirements provided by the user; the visual language processing unit of the image generation model can also be used to perform fine-grained descriptions of target image samples, such as using instructions such as "describe all the contents in the picture in detailed language" and "describe the picture in detail according to where it is, the main objects, what they are doing, and what is in the background" to obtain the initial description text corresponding to the second image sample label, and multiple initial description texts can be generated for the same image sample label; and using the matching model, the initial description text with a higher degree of match with the target image sample is selected from multiple initial description texts as the image description text of the target image sample.
[0229] The image generation model training method provided by the embodiments of the present disclosure can obtain high-quality target image samples by cleaning and screening the labels of the initial image samples, and can use the matching model to obtain the target image description text corresponding to the target image samples, so that the initial image generation model can be subsequently trained with the target image description text and the target image samples, so that the trained image generation model can generate higher-quality target images.
[0230] In one or more embodiments of the present disclosure, the target image generation sample data includes an image generation training sample and an image generation training label corresponding to the image generation training sample;
[0231] Generating sample data according to the target image and training to obtain the image generation model includes:
[0232] Determine a target sample data identifier set corresponding to the image generation training sample using an initial image generation model;
[0233] In a case where the target sample data identification set includes the target sample data identification, using a target feature extraction layer to obtain an initial feature vector corresponding to the target sample data identification set;
[0234] Performing vector mapping on the initial feature vector to obtain a target feature vector, and obtaining a target sample image based on the target feature vector;
[0235] The initial image generation model is trained according to the target sample image and the image generation training label to obtain the image generation model.
[0236] The specific implementation can be found in the above embodiments, which will not be described again here.
[0237] The image generation model training method provided by the embodiment of the present disclosure can use the generated data to train the image generation model, that is, the image can be described in fine-grained form by the visual language processing unit, and the description can be cleaned with a matching model, thereby ensuring that the image description contains as much information as possible while being more consistent with the image content, thereby improving the fine-grained control of the image processing unit by the visual language processing unit, and the control of fine-grained details helps to improve the quality and authenticity of the generated image.
[0238] 6 , which shows a flowchart of a processing process of an image generation model training method provided by one embodiment of the present disclosure, specifically comprising the following steps:
[0239] According to the architecture diagram of the image generation model in Figure 3, it can be seen that the image generation model consists of a visual language model and a diffusion model.
[0240] Step 602: Model alignment stage.
[0241] In this stage, the output of the visual language model is adjusted so that its output feature vector matches the representation of the diffusion model. This process involves mapping the initial feature vector so that the mapped feature vector can be processed by the diffusion model to ensure seamless connection between the visual language model and the diffusion model.
[0242] Model alignment is the basis for subsequent image and text generation, and can effectively improve the consistency and accuracy of the generated results. This process is the initial training stage of the image generation model, and there is a lot of training data. Therefore, only the data identification prediction layer and the vector mapping layer in the image generation model can be trained. The training parameters are small, which can ensure training efficiency. With sufficient training data, the effective alignment of the visual language model and the diffusion model can be completed.
[0243] Step 604: Generate a quality improvement phase.
[0244] Based on the model alignment stage, training of the fully connected linear layer and diffusion model for image generation is added to optimize and improve the model performance.
[0245] During this stage, the amount of training data is reduced, but the data quality is improved, using data with higher aesthetics and resolution for training. The fully connected linear layer for image generation is responsible for learning complex image-text relationships on large-scale datasets to improve the accuracy of text-to-image mapping. The diffusion model focuses on generating image details, ensuring that the generated images are rich in details and high in quality. By combining the two for training, more realistic and accurate image and text content can be generated.
[0246] Step 606: Adjustment and expansion phase.
[0247] The image generation model is carefully adjusted through real text-image dialogue data from user interactions.
[0248] This stage focuses on training the image generation model to understand different dialogue scenarios and control instructions, and to generate corresponding images and text based on different dialogue scenarios and control instructions.
[0249] For example, users can input the control command of "sunset on the beach", and the image generation model can generate images and text that match this scene; for example, the initial image can be edited and controlled by the control command of changing the color of the clothes in the image; in addition, this stage covers the diversified scene control of image and text generation tasks, that is, supporting more diverse user needs and application scenarios, such as providing users with not only generated target images, but also providing users with replies together with images and texts.
[0250] The image generation model training method provided by the embodiment of the present disclosure fully utilizes the understanding ability of the visual language model and the image generation ability of the diffusion model through three-stage training. First, the alignment of the visual language model and the diffusion model is completed; then the generation quality of the image generation model is improved; it can also support diverse scene control and dialogue capabilities, and can support various tasks such as dialogue control image generation and graphic text generation.
[0251] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of an image generation device. FIG7 shows a schematic structural diagram of an image generation device provided by an embodiment of the present disclosure. As shown in FIG7 , the device includes:
[0252] The data determination module 702 is configured to determine the image generation data; the image acquisition module 704 is configured to process the image generation data using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on the target image sample and the target image description text, the target image description text is determined based on the matching relationship between the target image sample and the initial image description text, and the initial image description text is obtained by using the image generation model to perform a text description of the target image sample.
[0253] Optionally, the image acquisition module 704 is further configured to:
[0254] Utilize the image generation model to determine the target data identification set corresponding to the image generation data; when the target data identification set includes the target data identification, utilize the target feature extraction layer of the image generation model to obtain the initial feature vector corresponding to the target data identification set, wherein the initial feature vector is a vector containing the features of the image generation data; perform vector mapping on the initial feature vector to obtain the target feature vector, and obtain the target image based on the target feature vector.
[0255] Optionally, the image acquisition module 704 is further configured to:
[0256] Determine an initial data identification set corresponding to the image generation data, and input the initial data identification set into the visual language processing unit; use the visual language processing unit to predict the initial data identification set to obtain a predicted data identification set; and obtain a target data identification set corresponding to the image generation data based on the initial data identification set and the predicted data identification set.
[0257] Optionally, the image acquisition module 704 is further configured to:
[0258] The initial feature vector is vector mapped by utilizing the vector mapping layer of the image generation model to obtain the target feature vector corresponding to the image processing unit of the image generation model.
[0259] Optionally, the image acquisition module 704 is further configured to:
[0260] The target feature vector is processed using an image processing unit of the image generation model to obtain the target image.
[0261] Optionally, the image acquisition module 704 is further configured to:
[0262] Determine a first initial data identification set corresponding to the image generation data, and input the first initial data identification set into the visual language processing unit; use the visual language processing unit to predict the first initial data identification set to obtain a first predicted data identification set; obtain a first target data identification set corresponding to the image generation data based on the first initial data identification set and the first predicted data identification set; or, determine a second initial data identification set corresponding to the image generation data, and input the second initial data identification set into the visual language processing unit; use the visual language processing unit to predict the second initial data identification set to obtain a second predicted data identification set; obtain a second target data identification set corresponding to the image generation data based on the second initial data identification set and the second predicted data identification set.
[0263] Optionally, the image acquisition module 704 is further configured to:
[0264] When the first target data identification set includes target data identification, the target feature extraction layer in the visual language processing unit is used to obtain the first initial feature vector corresponding to the first target data identification set; or, when the second target data identification set includes target data identification, the target feature extraction layer in the visual language processing unit is used to obtain the second initial feature vector corresponding to the second target data identification set.
[0265] Optionally, the image acquisition module 704 is further configured to:
[0266] Using the vector mapping layer of the image generation model, the first initial feature vector is vector mapped to obtain a first target feature vector corresponding to the image processing unit; or, using the vector mapping layer of the image generation model, the second initial feature vector is vector mapped to obtain a second target feature vector corresponding to the image processing unit.
[0267] Optionally, the image acquisition module 704 is further configured to:
[0268] Processing the first target feature vector using an image processing unit of the image generation model to obtain a first target image; or
[0269] The second target feature vector is processed using an image processing unit of the image generation model to obtain a second target image.
[0270] The device further comprises:
[0271] The first text generation module is configured to generate a target image description text corresponding to the target image by using the visual language processing unit when it is determined that the image generation data contains a text description generation instruction for the target image.
[0272] The device further comprises:
[0273] A second text generation module displays the target image to the user through a user interaction interface of the client;
[0274] When a text description generation instruction for the target image is received from the client, the visual language processing unit is used to generate a target image description text corresponding to the target image.
[0275] The device further comprises:
[0276] The display module is configured to display the target image and the target image description text corresponding to the target image to the user through the user interaction interface of the client.
[0277] The device further comprises:
[0278] a target text obtaining module configured to, when determining that the target data identifier set does not include the target data identifier, obtain a text feature vector corresponding to the target data identifier set using a text feature extraction layer in the visual language processing unit;
[0279] The text feature vector is decoded using a decoding layer in the visual language processing unit to obtain a target text.
[0280] The device further comprises:
[0281] The training module is configured to update and train the image generation model according to the target image and the target image description text.
[0282] The image generation device provided by the embodiment of the present disclosure utilizes the text and image understanding capabilities of the visual language processing unit in the image generation model, and can obtain a more accurate initial feature vector of the image generation data through the target feature extraction layer. The initial feature vector of the image generation data is processed through the vector mapping layer of the image generation model, and the image processing unit and the corresponding target feature vector of the image generation model can be obtained, thereby realizing the use of the image processing unit to process the target feature vector and obtain a target image that is consistent with the description of the image generation data and has high quality.
[0283] The above is a schematic diagram of an image generation device according to this embodiment. It should be noted that the technical solution of this image generation device and the technical solution of the aforementioned image generation method are based on the same concept. For details not described in detail in the technical solution of the image generation device, please refer to the description of the technical solution of the aforementioned image generation method.
[0284] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of an image generation model training device. FIG8 shows a schematic structural diagram of an image generation model training device provided by an embodiment of the present disclosure. As shown in FIG8 , the device includes:
[0285] The determination module 802 is configured to determine a target image sample and use an image generation model to perform a text description on the target image sample to obtain a plurality of initial image description texts of the target image sample;
[0286] The text obtaining module 804 is configured to determine the target image description text according to the matching relationship between the target image sample and each initial image description text;
[0287] The data determination module 806 is configured to determine target image generation sample data according to the target image sample and the target image description text;
[0288] The training module 808 is configured to generate sample data according to the target image and obtain the image generation model through training.
[0289] Optionally, the data determination module 806 is further configured to:
[0290] Using the target image description text as an image generation training sample, using the target image sample as an image generation training label, and determining target image generation sample data based on the image generation training sample and the image generation training label; or
[0291] Determining an associated image sample corresponding to the target image sample;
[0292] The image description text and the associated image sample are used as image generation training samples, the target image sample is used as an image generation training label, and target image generation sample data is determined based on the image generation training sample and the image generation training label.
[0293] Optionally, the text obtaining module 804 is further configured to:
[0294] Using a matching model, determining a matching relationship between the target image sample and each initial image description text;
[0295] According to the matching relationship, the image description text of the target image sample is determined from the multiple initial image description texts.
[0296] Optionally, the training module 808 is further configured to:
[0297] Determine a target sample data identifier set corresponding to the image generation training sample using an initial image generation model;
[0298] In a case where the target sample data identification set includes the target sample data identification, using a target feature extraction layer to obtain an initial feature vector corresponding to the target sample data identification set;
[0299] Performing vector mapping on the initial feature vector to obtain a target feature vector, and obtaining a target sample image based on the target feature vector;
[0300] The initial image generation model is trained according to the target sample image and the image generation training label to obtain the image generation model.
[0301] The device further comprises:
[0302] The initial training module is configured to determine initial image generation sample data and initial image sample labels; use the initial image generation model to determine the target sample data identification set corresponding to the initial image generation sample data; when it is determined that the target sample data identification set contains the target sample data identification, use the target feature extraction layer to obtain the initial feature vector corresponding to the target sample data identification set; perform vector mapping on the initial feature vector to obtain the target feature vector, and obtain the target sample image based on the target feature vector; train the initial image generation model based on the target sample image and the initial image sample label to obtain the image generation model.
[0303] The above is a schematic diagram of an image generation model training device according to this embodiment. It should be noted that the technical solution of this image generation model training device and the technical solution of the aforementioned image generation model training method are based on the same concept. For details not described in detail in the technical solution of the image generation model training device, please refer to the description of the technical solution of the aforementioned image generation model training method.
[0304] The image generation model training device provided by the embodiment of the present disclosure can use the generated data to train the image generation model, that is, the image can be described in fine-grained form by the visual language processing unit, and the description can be cleaned with a matching model, thereby ensuring that the image description contains as much information as possible while being more consistent with the image content, thereby improving the fine-grained control of the image processing unit by the visual language processing unit, and the control of fine-grained details helps to improve the quality and authenticity of the generated image.
[0305] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of an image generation device applied to the cloud. FIG9 shows a schematic structural diagram of an image generation device applied to the cloud provided by an embodiment of the present disclosure. As shown in FIG9 , the device includes:
[0306] Receiving module 902, configured to receive image generation data sent by the client;
[0307] An image acquisition module 904 is configured to process the image generation data using an image generation model to generate a target image, wherein the image generation model is obtained by training target image generation sample data, the target image generation sample data is determined based on a target image sample and a target image description text, the target image description text is determined based on a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by performing a text description of the target image sample using the image generation model;
[0308] The display module 906 is configured to display the target image to the user through the user interaction interface of the client.
[0309] The image generation device applied to the cloud provided by the embodiment of the present disclosure receives image generation data sent by the user through the user interaction interface of the client, utilizes the text and image understanding capabilities of the visual language processing unit in the image generation model, and obtains a more accurate initial feature vector of the image generation data through the target feature extraction layer, and obtains the image processing unit and the corresponding target feature vector of the image generation model through the vector mapping layer of the image generation model, thereby realizing the use of the image processing unit to obtain a target image that is consistent with the description of the image generation data and has high quality, and displays the target image to the user through the user interaction interface, thereby improving the user's interactive experience.
[0310] The above is a schematic diagram of an image generation device for cloud-based applications according to this embodiment. It should be noted that the technical solution of this image generation device for cloud-based applications shares the same concept as the technical solution of the aforementioned image generation method for cloud-based applications. For details not described in detail in the technical solution of the image generation device for cloud-based applications, please refer to the description of the technical solution of the aforementioned image generation method for cloud-based applications.
[0311] Figure 10 shows a block diagram of a computing device 1000 according to one embodiment of the present disclosure. Components of the computing device 1000 include, but are not limited to, a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 via a bus 1030, and a database 1050 is used to store data.
[0312] The computing device 1000 also includes an access device 1040 that enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0313] In one embodiment of the present disclosure, the aforementioned components of the computing device 1000 and other components not shown in FIG10 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG10 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0314] Computing device 1000 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1000 may also be a mobile or stationary server.
[0315] Among them, the processor 1020 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image generation method or image generation model training method.
[0316] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solution of the aforementioned image generation method or image generation model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned image generation method or image generation model training method.
[0317] An embodiment of the present disclosure also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image generation method or image generation model training method.
[0318] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned image generation method or image generation model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned image generation method or image generation model training method.
[0319] An embodiment of the present disclosure further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned image generation method or image generation model training method.
[0320] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product is based on the same concept as the technical solution of the aforementioned image generation method or image generation model training method. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned image generation method or image generation model training method.
[0321] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0322] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0323] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.
[0324] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0325] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. An image generation method, comprising: Determining image generation data; Processing the image generation data by using an image generation model to generate a target image, wherein the image generation model is obtained by training with target image generation sample data, the target image generation sample data is determined according to a target image sample and a target image description text, the target image description text is determined according to a matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by using the image generation model to perform text description on the target image sample.
2. The image generation method according to claim 1, wherein the processing the image generation data by using the image generation model to obtain a target image comprises: Using the image generation model to determine a target data identifier set corresponding to the image generation data; When the target data identifier set contains a target data identifier, using a target feature extraction layer of the image generation model to obtain an initial feature vector corresponding to the target data identifier set, wherein the initial feature vector is a vector containing features of the image generation data; Performing vector mapping on the initial feature vector to obtain a target feature vector, and obtaining the target image according to the target feature vector.
3. The image generation method according to claim 2, wherein the image generation model comprises a vision-language processing unit; The using the image generation model to determine a target data identifier set corresponding to the image generation data comprises: Determining an initial data identifier set corresponding to the image generation data, and inputting the initial data identifier set into the vision-language processing unit; Using the vision-language processing unit to perform prediction on the initial data identifier set to obtain a predicted data identifier set; Obtaining the target data identifier set corresponding to the image generation data according to the initial data identifier set and the predicted data identifier set.
4. The image generation method according to claim 2, wherein the image generation model comprises a vector mapping layer and an image processing unit; The performing vector mapping on the initial feature vector to obtain a target feature vector comprises: Using the vector mapping layer of the image generation model to perform vector mapping on the initial feature vector to obtain the target feature vector corresponding to the image processing unit of the image generation model.
5. The image generation method according to claim 2, wherein the image generation model comprises an image processing unit; The obtaining the target image according to the target feature vector comprises: Using the image processing unit of the image generation model to process the target feature vector to obtain the target image.
6. The image generation method according to claim 2, wherein the image generation model comprises a vision-language processing unit, and the image generation data comprises a target text, or a prompt text and an initial image; The using the image generation model to determine a target data identifier set corresponding to the image generation data comprises: Determine the first initial data identifier set corresponding to the target text, and input the first initial data identifier set into the vision-language processing unit; Use the vision-language processing unit to predict the first initial data identifier set to obtain a first predicted data identifier set; Obtain a first target data identifier set corresponding to the image generation data according to the first initial data identifier set and the first predicted data identifier set; or Determine a second initial data identifier set corresponding to the prompt text and the initial image, and input the second initial data identifier set into the vision-language processing unit; Use the vision-language processing unit to predict the second initial data identifier set to obtain a second predicted data identifier set; Obtain a second target data identifier set corresponding to the image generation data according to the second initial data identifier set and the second predicted data identifier set.
7. The image generation method according to claim 6, wherein when the target data identifier set contains a target data identifier, using the target feature extraction layer of the image generation model to obtain an initial feature vector corresponding to the target data identifier set, includes: When the first target data identifier set contains the target data identifier, use the target feature extraction layer in the vision-language processing unit to obtain a first initial feature vector corresponding to the first target data identifier set; Or When the second target data identifier set contains the target data identifier, use the target feature extraction layer in the vision-language processing unit to obtain a second initial feature vector corresponding to the second target data identifier set.
8. The image generation method according to claim 7, wherein the image generation model includes a vector mapping layer; The vector mapping of the initial feature vector to obtain a target feature vector includes: Use the vector mapping layer of the image generation model to perform vector mapping on the first initial feature vector to obtain a first target feature vector corresponding to the image processing unit; Or Use the vector mapping layer of the image generation model to perform vector mapping on the second initial feature vector to obtain a second target feature vector corresponding to the image processing unit.
9. The image generation method according to claim 8, wherein the image generation model includes an image processing unit; The obtaining of the target image according to the target feature vector includes: Use the image processing unit of the image generation model to process the first target feature vector to obtain a first target image; Or Use the image processing unit of the image generation model to process the second target feature vector to obtain a second target image.
10. The image generation method according to claim 3, after obtaining the target image, further includes: When it is determined that the image generation data contains a text description generation instruction for the target image, use the vision-language processing unit to generate a target image description text corresponding to the target image.
11. The image generation method according to claim 3, after obtaining the target image, further comprising: Displaying the target image to the user through the user interaction interface of the client; When receiving a text description generation instruction for the target image sent by the client, using the vision-language processing unit to generate a target image description text corresponding to the target image.
12. The image generation method according to claim 10 or 11, after generating the target image description text corresponding to the target image, further comprising: Displaying the target image and the target image description text corresponding to the target image to the user through the user interaction interface of the client.
13. The image generation method according to claim 3, after using the image generation model to determine the target data identifier set corresponding to the image generation data, further comprising: When determining that the target data identifier set does not include a target data identifier, using the text feature extraction layer in the vision-language processing unit to obtain a text feature vector corresponding to the target data identifier set; Using the decoding layer in the vision-language processing unit to decode the text feature vector to obtain a target text.
14. The image generation method according to claim 10 or 11, further comprising: Updating and training the image generation model according to the target image and the target image description text.
15. An image generation model training method, comprising: Determining a target image sample, and using an image generation model to perform text description on the target image sample to obtain a plurality of initial image description texts of the target image sample; Determining a target image description text according to the matching relationship between the target image sample and each initial image description text; Determining target image generation sample data according to the target image sample and the target image description text; Training to obtain the image generation model according to the target image generation sample data.
16. The image generation model training method according to claim 15, the determining target image generation sample data according to the target image sample and the target image description text includes: Using the target image description text as an image generation training sample, using the target image sample as an image generation training label, and determining the target image generation sample data according to the image generation training sample and the image generation training label; or Determining an associated image sample corresponding to the target image sample; Using the target image description text and the associated image sample as an image generation training sample, using the target image sample as an image generation training label, and determining the target image generation sample data according to the image generation training sample and the image generation training label.
17. The image generation model training method according to claim 15 or 16, the determining a target image description text according to the matching relationship between the target image sample and each initial image description text includes: Using a matching model, determine the matching relationships between the target image sample and each initial image description text; According to the matching relationships, determine the target image description text of the target image sample from the multiple initial image description texts.
18. The method for training an image generation model according to any one of claims 15 to 17, wherein the target image generation sample data includes an image generation training sample and an image generation training label corresponding to the image generation training sample; The training to obtain the image generation model according to the target image generation sample data includes: Using an initial image generation model, determine a set of target sample data identifiers corresponding to the image generation training sample; When the set of target sample data identifiers contains a target sample data identifier, use a target feature extraction layer to obtain an initial feature vector corresponding to the set of target sample data identifiers; Perform vector mapping on the initial feature vector to obtain a target feature vector, and obtain a target sample image according to the target feature vector; According to the target sample image and the image generation training label, train the initial image generation model to obtain the image generation model.
19. The method for training an image generation model according to any one of claims 15 to 18, before determining the target image sample, further includes: Determine initial image generation sample data and an initial image sample label; Using the initial image generation model, determine a set of target sample data identifiers corresponding to the initial image generation sample data; When the set of target sample data identifiers contains a target sample data identifier, use a target feature extraction layer to obtain an initial feature vector corresponding to the set of target sample data identifiers; Perform vector mapping on the initial feature vector to obtain a target feature vector, and obtain a target sample image according to the target feature vector; According to the target sample image and the initial image sample label, train the initial image generation model to obtain the image generation model.
20. An image generation method, applied to the cloud, includes: Receive image generation data sent by a client; Use an image generation model to process the image generation data to generate a target image, wherein the image generation model is trained by target image generation sample data, the target image generation sample data is determined according to a target image sample and a target image description text, the target image description text is determined according to the matching relationship between the target image sample and an initial image description text, and the initial image description text is obtained by using the image generation model to perform text description on the target image sample; Display the target image to a user through a user interface of the client.
21. A computing device, includes: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the image generation method according to any one of claims 1 to 14 or the image generation model training method according to any one of claims 15 to 19 are implemented.
22. A computer-readable storage medium stores a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the image generation method according to any one of claims 1 to 14 or the image generation model training method according to any one of claims 15 to 19 are implemented.
23. A computer program product includes a computer program. When the computer program is executed by a processor, the steps of the image generation method according to any one of claims 1 to 14 or the image generation model training method according to any one of claims 15 to 19 are implemented.
Citation Information
Patent Citations
Image generation method
CN116778011A
Image generation method and device, electronic equipment, storage medium and program product
CN117315070A
Image generation method and data processing method for image generation
CN117409109A
Image generation method and device, electronic equipment, storage medium and program product
CN117437317A
Image generation method and apparatus, and device and medium
WO2023221363A1