Image generation method, electronic device, and computer readable storage medium

By acquiring text information and enhancing mark information, and using the image generation model to generate multi-modal image, the problems of low accuracy of image generation and limited range in the prior art are solved, and high-quality, multi-object image generation is achieved.

WO2025112588A1PCT designated stage expired Publication Date: 2025-06-05ALIBABA (CHINA) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/107883
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-07-26
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the prior art, images are generated only through text prompts, resulting in low accuracy of generated images and limited image generation range.

Method used

By obtaining multimodal prompt information, including text information and enhanced marking information, the image generation model is used to generate multimodal image to generate the target image. Enhanced marking information is used to determine the positional features and image features of the target object.

Benefits of technology

It is realized that high-quality target images including multiple target objects are accurately generated through multi-modal prompt information, and the accuracy and generation range of image generation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024107883_05062025_PF_FP_ABST
    Figure CN2024107883_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical fields of large model technology and image processing. Disclosed are an image generation method, an electronic device, and a computer readable storage medium. The method comprises: acquiring multi-modal prompt information, wherein the multi-modal prompt information comprises: text information and augmented token information, the text information is used for describing image content to be generated, said image content comprises: at least one target object, and the augmented token information is used for determining position features and image features of the at least one target object; and using an image generation model to perform multi-modal image generation on the multi-modal prompt information to obtain a target image, wherein the image generation model is used for using a multi-modal image generation mode to generate the target image. The present disclosure solves the technical problems in the related art of low accuracy of generated images and limited image generation range caused by using only text prompt to generate the corresponding images.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, electronic device, and computer-readable storage medium Technical Field

[0001] The present disclosure relates to the fields of large model technology and image processing technology, and in particular to an image generation method, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of artificial intelligence, significant progress has been made in the field of text-to-image generation, which can generate corresponding high-quality images based on text prompts, thereby realizing the function of automatic image generation.

[0003] Currently, only textual cues are used to describe objects to generate corresponding images. Since it is difficult to describe physics with textual cues, the accuracy of the generated images is low. In addition, this method usually requires fine-tuning or only supports the use of a single object as a constraint, so the scope of image generation is relatively limited.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far.

[0005] Summary of the Invention

[0006] The embodiments of the present disclosure provide an image generation method, an electronic device, and a computer-readable storage medium to at least solve the technical problem in the related art that corresponding images are generated only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0007] According to one aspect of an embodiment of the present disclosure, an image generation method is provided, including: obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of the at least one target object; using an image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.

[0008] According to another aspect of an embodiment of the present disclosure, another image generation method is provided, which provides a graphical user interface through a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene, including: in response to a first control operation performed on the graphical user interface, inputting text information, wherein the text information is used to describe the image content to be generated, and the image content includes: at least one target object; in response to a second control operation performed on the graphical user interface, generating position features and image features of at least one target object based on the text information to obtain enhanced marking information, and using an image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the enhanced marking information is used to determine the position features and image features of at least one target object, and the image generation model is used to generate the target image using a multimodal image generation method, and the multimodal prompt information includes: text information and enhanced marking information; and displaying the target image in the graphical user interface.

[0009] According to another aspect of an embodiment of the present disclosure, another image generation method is provided, including: receiving a currently input multimodal conversation request, wherein the information carried in the multimodal conversation request includes: multimodal conversation text information and multimodal conversation enhanced tag information, the multimodal conversation text information is used to describe the content of the multimodal conversation image to be generated, the multimodal conversation image content includes: at least one object, and the multimodal conversation enhanced tag information is used to determine the position features and image features of the at least one object; using an image generation model to perform multimodal image generation on the multimodal conversation request to obtain a multimodal conversation image, wherein the image generation model is used to generate the multimodal conversation image using a multimodal image generation method; and feeding back a multimodal conversation response, wherein the information carried in the multimodal conversation response includes: the multimodal conversation image.

[0010] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes any one of the above-mentioned image generation methods when running.

[0011] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is further provided, which includes a stored executable program, wherein when the executable program runs, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned image generation methods.

[0012] In the embodiment of the present disclosure, by obtaining multimodal prompt information including text information and enhanced mark information, it is possible to determine the image content to be generated based on the text information, determine the position features and image features of at least one target object in the image content based on the enhanced mark information, and then use the image generation model to perform multimodal image generation on the multimodal prompt information to obtain the target image, thereby achieving the purpose of accurately generating high-quality target images including multiple target objects through multimodal prompt information. In addition, it is possible to achieve more precise control over the generated image, generate images with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby achieving more accurate, more diverse, and high-quality image generation that supports multiple target objects, so as to meet the technical effect of image generation needs in different fields, thereby solving the technical problem in the related art of generating corresponding images only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0013] It is easy to note that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present disclosure, and do not constitute a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0015] FIG1 is a schematic diagram of an application scenario of an image generation method according to Embodiment 1 of the present disclosure;

[0016] FIG2 is a flow chart of an image generating method according to Embodiment 1 of the present disclosure;

[0017] FIG3 is a flow chart of another image generating method according to Embodiment 1 of the present disclosure;

[0018] FIG4 is a flow chart of an image generating method according to Embodiment 2 of the present disclosure;

[0019] FIG5 is a flow chart of an image generating method according to Embodiment 3 of the present disclosure;

[0020] FIG6 is a schematic diagram of a human-computer dialogue scenario according to Example 3 of the present disclosure;

[0021] FIG7 is a schematic structural diagram of an image generating device according to Embodiment 4 of the present disclosure;

[0022] FIG8 is a schematic structural diagram of another image generating device according to Embodiment 4 of the present disclosure;

[0023] FIG9 is a schematic structural diagram of another image generating device according to Embodiment 4 of the present disclosure;

[0024] FIG10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present disclosure. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] The technical solution provided by the present disclosure is mainly implemented using large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. The large model can also be called a cornerstone model / foundation model. It is pre-trained by using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as large-scale language models (LLMs) and multi-modal pre-training models.

[0028] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned with a small number of samples, so that the large model can be applied to different tasks. For example, the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0029] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are subject to the following explanations:

[0030] Diffusion Model: A mathematical model used to describe the diffusion process, and is a commonly used generative model. In the disclosed embodiments, the diffusion model can denoise a noisy image based on the diffusion process, thereby generating a generated image.

[0031] Constituency tree: A tree-like structure used to represent the structure of natural language sentences, consisting of nodes and edges, where nodes represent words or phrases in a sentence and edges represent the grammatical relationships between these words or phrases. In the constituency tree, sentences are broken down into different phrase structures, such as noun phrases, verb phrases, adjective phrases, etc. These phrase structures can be further decomposed into smaller phrase structures, down to the smallest word level. The constituency tree can help understand the structure and grammatical relationships of sentences, facilitate syntactic and semantic analysis, and can also be used for natural language processing tasks such as syntactic analysis, machine translation, and information extraction. By analyzing the constituency tree, we can better understand the grammatical structure and semantic meaning of a sentence.

[0032] Natural Language Processing (NLP) parser: A tool used for natural language processing that converts natural language text into structured data and performs tasks such as grammatical analysis, part-of-speech tagging, and named entity recognition. Typically based on machine learning and deep learning technologies, NLP parsers can automatically analyze text and extract information from it. They can help understand and process large amounts of natural language data, such as speech recognition, sentiment analysis, and text classification.

[0033] Text Encoder: A model or tool used to convert text data into numerical representations. In the field of natural language processing (NLP), text encoders are often used to convert text data into vectors or matrices for machine learning and deep learning tasks such as text classification, semantic similarity calculation, and sentiment analysis.

[0034] Fine-tuning refers to further training and adjusting a pre-trained model on a specific task to improve its performance and adaptability. Pre-trained models are typically trained on large datasets, while fine-tuning involves fine-tuning the model on a specific, smaller dataset to better adapt it to the specific task. Fine-tuning is commonly used in deep learning models in fields such as natural language processing and computer vision.

[0035] In related text-to-image generation methods, text-to-image models typically use word embeddings extracted from textual cues as conditions. However, the high level of abstraction and limited information density of text make it difficult to accurately describe objects, resulting in low precision in generated images. Furthermore, these methods often require fine-tuning or only support the use of a single object as a constraint, limiting the scope of image generation.

[0036] Currently, the BLIP-Diffusion model has been proposed, which can generate corresponding images from mixed images and text, but it only supports the generation of single objects and the accuracy of the generated images is low. In addition, the KOSMOS-G model has been proposed, which can also generate corresponding images from mixed images and text, but the accuracy of the generated images is still low.

[0037] The related art method of generating corresponding images based on text prompts has the following defects.

[0038] Defect 1: The corresponding image is generated by describing the object only through text prompts. Since it is difficult to describe physics through text prompts, the accuracy of the generated image is low.

[0039] Disadvantage 2: Usually requires fine-tuning or only supports using a single object as a constraint, resulting in high model training costs and a limited range of image generation;

[0040] Defect 3: The BLIP-Diffusion model and the KOSMOS-G model support mixed generation of images and text, but the accuracy of the generated images is low, and the BLIP-Diffusion model only supports single-object generation.

[0041] With respect to the above-mentioned defects, no effective solution has been proposed before the present disclosure.

[0042] Example 1

[0043] According to an embodiment of the present disclosure, an image generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0044] Considering the huge number of model parameters of the large model and the limited computing resources of the mobile terminal, the above-mentioned image generation method provided by the embodiment of the present disclosure can be applied to the application scenario shown in Figure 1, but is not limited thereto. In the application scenario shown in Figure 1, the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client devices 20 here can include but are not limited to: smart phones, tablet computers, laptops, PDAs, personal computers, smart home devices, car-mounted devices, etc. The client device 20 can interact with the user through a graphical user interface to call the large model, thereby implementing the method provided by the embodiment of the present disclosure.

[0045] In an embodiment of the present disclosure, a system composed of a client device and a server can perform the following steps: the client device executes steps such as obtaining multimodal prompt information input by a user in a graphical user interface and sending the multimodal prompt information to the server; the server executes steps such as using an image generation model to generate a multimodal image for the obtained multimodal prompt information, thereby obtaining a target image, and returning the target image to the client device. It should be noted that when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiment of the present disclosure can be performed in the client device.

[0046] In the above operating environment, the present disclosure provides an image generation method as shown in Figure 2. Figure 2 is a flow chart of an image generation method according to Example 1 of the present disclosure. As shown in Figure 2, the method may include the following steps:

[0047] Step S21: acquiring multimodal prompt information, wherein the multimodal prompt information includes text information and enhanced markup information, wherein the text information is used to describe the image content to be generated, the image content includes at least one target object, and the enhanced markup information is used to determine the position characteristics and image characteristics of the at least one target object;

[0048] Step S22: Using an image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.

[0049] The text information can be understood as a text prompt, that is, the image content described in natural language, which is used to describe the image content to be generated. Exemplarily, the natural language can be Chinese, English, Japanese, etc., which are not limited here. In the embodiment of the present disclosure, the image content described by the text information may include at least one target object, and the target object can be understood as a person, animal or object in the image content to be generated, that is, the present disclosure can support multi-object image generation. Exemplarily, the target object can include real people such as boys, girls, doctors, and teachers, and can include virtual characters such as game players and non-player characters (NPCs), and can include animals such as cats, dogs, peacocks, and elephants, and can also include objects such as tables, cars, grass, trees, and train stations, which are not limited here.

[0050] For example, the text information can generate text information corresponding to the user's image generation request. If the user wishes to obtain an image of a cat and a dog on grass, the corresponding text information can be "a cat and a dog on the grass." This text information includes three target objects: a cat, a dog, and grass. It is understood that this text information can also be replaced with "a dog and a cat on the grass" or "there is a cat and a dog on the grass." This disclosure does not limit the description method of the text information.

[0051] Considering that describing an object only through textual hints will lead to inaccurate image generation, the present disclosure uses enhanced tag information as additional information on the basis of providing textual information, and generates a target image through the textual information and the enhanced tag information, thereby improving the accuracy of image generation.

[0052] The enhanced tag information may include coordinate information and image information, wherein the coordinate information may be the coordinates of at least one target object included in the text information, and is used to determine the positional features of the at least one target object. For example, if a user wishes to obtain an image of a cat and a dog on grass, the coordinate information may include the coordinates of the cat (e.g., [coordinates: 12, 15, 100, 200]), the coordinates of the dog (e.g., [coordinates: 100, 150, 220, 240]), and the coordinates of the grass (e.g., [coordinates: 500, 500, 500, 500]).

[0053] Exemplarily, the coordinates of at least one target object may be coordinates in a two-dimensional rectangular coordinate system, which can represent the coordinate range of the target object, or coordinates in a two-dimensional coordinate system or a three-dimensional coordinate system. The present disclosure does not limit the coordinate system used.

[0054] The image information may be an image of at least one target object included in the text information, and is used to determine the image features of the at least one target object in the text information. For example, if a user wishes to obtain an image of a cat and a dog on grass, the image information may include an image of a cat (e.g., [Image: an image of a cat uploaded by the user]), an image of a dog (e.g., [Image: an image of a dog uploaded by the user]), and an image of grass (e.g., [Image: an image of grass uploaded by the user]).

[0055] Multimodal prompt information can be understood as prompt information in multiple modes, including the above-mentioned text information and enhanced mark information. In the embodiment of the present disclosure, multimodal prompt information includes prompt information in text mode, coordinate mode, and image mode.

[0056] Exemplarily, multimodal prompt information of the same target object can be bundled together to avoid affecting other objects or global conditions. That is, the present disclosure introduces an augmented token, which contains object-level text, coordinates and image information at the same time, and is used to describe an object in the generated image.

[0057] In the disclosed embodiments, the image generation model is a model suitable for generating images. For example, the image generation model can be a pre-trained diffusion-based image generation model, i.e., a diffusion model. The image generation model can also be an autoregressive model, etc., without limitation herein. The image generation model can accurately generate high-quality target images based on multimodal cue information of multiple target objects, i.e., it can generate high-quality images based on multimodal cues at the multi-object level, thereby achieving precise control over the generated images.

[0058] Exemplarily, if the multimodal prompt information of the above example is input into the image generation model, that is, the input text information: a cat and a dog on the grass, coordinate information: the coordinates of the cat, the coordinates of the dog and the coordinates of the grass, image information: the image of a cat, the image of a dog and the image of the grass, then the image generation model of the present invention can accurately output the corresponding target image according to the multimodal prompt information, that is, the image of a cat and a dog on the grass.

[0059] It can be understood that the image generation model in the embodiment of the present disclosure can maintain the original structure of the pre-trained text-to-image model as much as possible so as to integrate extensions to the existing model, and the image generation model in the embodiment of the present disclosure only changes the input of the model and does not need to change the architecture of the model, so it can maintain the usability of the technology built on the basis of the model.

[0060] In the embodiment of the present disclosure, by obtaining multimodal prompt information including text information and enhanced mark information, it is possible to determine the image content to be generated based on the text information, determine the position features and image features of at least one target object in the image content based on the enhanced mark information, and then use the image generation model to perform multimodal image generation on the multimodal prompt information to obtain the target image, thereby achieving the purpose of accurately generating high-quality target images including multiple target objects through multimodal prompt information, and being able to achieve more precise control over the generated image, generate images with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider.

[0061] The above-mentioned image generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving image generation in the fields of script services, design services, game services, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services. For example: in script services, the required image materials are generated according to the script or storyline; in design services, design drawings are generated according to the user's text, image, and coordinate descriptions; in games, corresponding character images are generated according to the text description of the character image; in e-commerce services, corresponding product images are generated according to the description of the product, etc., which are not limited here.

[0062] By adopting the embodiment of the present disclosure, by obtaining multimodal prompt information including text information and enhanced mark information, it is possible to determine the image content to be generated based on the text information, determine the position features and image features of at least one target object in the image content based on the enhanced mark information, and then use the image generation model to perform multimodal image generation on the multimodal prompt information to obtain the target image, thereby achieving the purpose of accurately generating high-quality target images including multiple target objects through multimodal prompt information. In addition, it is possible to achieve more precise control over the generated image, generate images with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby achieving more accurate, more diverse, and high-quality image generation that supports multiple target objects, so as to meet the technical effect of image generation needs in different fields, thereby solving the technical problem in the related art of generating corresponding images only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0063] In an optional embodiment, in step S21, obtaining multimodal prompt information includes the following method steps:

[0064] Step S211, obtaining text information;

[0065] Step S212, generating at least one target object's position feature and at least one target object's image feature based on the text information to obtain enhanced marking information;

[0066] Step S213: Combine the text information and the enhanced markup information to obtain multimodal prompt information.

[0067] Considering that it may be difficult to obtain multimodal prompt information for all target objects during the inference process, for example, only the textual information of the target object may be obtained, or only the multimodal prompt information for part of the target object may be obtained (for example, when the user wishes to merge real objects and generated objects in an image). Therefore, in the embodiments of the present disclosure, enhanced tag information can be obtained based on textual information, that is, enhanced tag information for missing target objects can be generated, thereby achieving flexible combination of multiple modalities.

[0068] In an embodiment of the present disclosure, when obtaining multimodal prompt information, text information can be obtained first, and then the position features of at least one target object and the image features of at least one target object can be generated based on the text information, that is, object-level coordinates and images can be generated according to the text prompt, thereby obtaining corresponding enhanced marking information, and then the text information and the enhanced marking information can be combined to obtain multimodal prompt information.

[0069] In an optional embodiment, in step S212, generating a position feature and an image feature of at least one target object based on the text information to obtain enhanced marking information includes the following method steps:

[0070] Step S2121: performing part-of-speech analysis on the text information and selecting at least one target segmentation word, wherein the at least one target segmentation word meets a preset part-of-speech requirement and the at least one target segmentation word is used to determine at least one target object;

[0071] Step S2122: generating a location feature of at least one target object based on the text information and the at least one target word segmentation, and generating an image feature of at least one target object based on the text information and the at least one target object location feature;

[0072] Step S2123 : Determine enhanced tag information using at least one target word segment, at least one location feature of a target object, and at least one image feature of a target object.

[0073] It is understandable that the parts of speech in a language include nouns, verbs, adjectives, adverbs, pronouns, numerals, quantifiers, conjunctions, prepositions, auxiliary words, interjections, etc., and nouns are usually used to represent people, things, places, etc., and need to be generated when generating images, so nouns can be used to represent target objects.

[0074] In the disclosed embodiment, the preset part of speech may be a noun, and the at least one target segmented word is at least one noun in the text information. By performing part-of-speech analysis on the text information and selecting at least one target segmented word that is a noun in the text information, at least one target object contained in the text information can be determined, thereby facilitating the acquisition of multimodal prompt information corresponding to each target object.

[0075] For example, a constituency tree can be used to identify objects in text information, that is, to identify target objects in the text information. For example, if the text information is "a cat and a dog on the grass", the text information is analyzed by the constituency tree to obtain "a (determiner) cat (noun) and (conjunction) a (determiner) dog (noun) on (preposition) the (determiner) grass (noun)", and the nouns in the text information, namely cat, dog, and grass, are identified, and the identified nouns cat, dog, and grass are selected as multiple target segmentations.

[0076] In addition, you can also use Conditional Random Fields (CRF), Maximum Entropy Model, etc. to perform part-of-speech tagging on the text through training models, or use Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM), Attention Mechanism, etc. for part-of-speech tagging and part-of-speech recognition, which are not restricted here.

[0077] In an embodiment of the present disclosure, when generating the location features and image features of at least one target object based on text information to obtain enhanced tag information, part-of-speech analysis can be performed on the text information to select at least one target segmentation word with a noun part of speech from the text information. Then, based on the text information and the at least one target segmentation word, the location features and image features of the at least one target object are generated. Furthermore, based on the at least one target segmentation word, the location features of the at least one target object, and the image features of the at least one target object, the enhanced tag information corresponding to each of the at least one target object is determined.

[0078] In an optional embodiment, in step S2122, generating a location feature of at least one target object based on the text information and at least one target word segmentation includes the following method steps:

[0079] Step S21221: Use a position feature generation model to generate position features for text information and at least one target word segmentation to obtain a position feature of at least one target object.

[0080] In the disclosed embodiments, the position feature generation model is a model suitable for generating position features, such as a diffusion model, an autoregressive model, etc. The position feature generation model may be a coordinate model that generates the position coordinates corresponding to each object based on the text content of the text information and the object corresponding to at least one target word.

[0081] In an embodiment of the present disclosure, when generating the location features of at least one target object based on text information and at least one target segmentation word, the text information and the at least one target segmentation word can be input into a location feature generation model, and the location feature generation model can be used to generate location features for the text information and the at least one target segmentation word, thereby obtaining the location features corresponding to the at least one target object, that is, obtaining the coordinate information corresponding to the at least one target object.

[0082] For example, if the text information is "a cat and a dog", the multiple target segmentations are "cat" and "dog". The present disclosure combines the three text features of "a cat and a dog", "cat" and "dog" together and uses them as the input of the position feature generation model, thereby obtaining the coordinates of the cat and the coordinates of the dog according to the output of the position feature generation model.

[0083] It can be seen that in the present disclosure, when coordinate modal information is missing, a position feature generation model can be used to automatically generate position coordinates corresponding to multiple objects based on text prompts and multiple objects, thereby completing the missing coordinate modal information.

[0084] In an optional embodiment, in step S2122, generating an image feature of at least one target object based on the text information and the position feature of at least one target object includes the following method steps:

[0085] Step S21222: Use an image feature generation model to generate image features for the text information and the position features of at least one target object to obtain image features of the at least one target object.

[0086] In the disclosed embodiments, the image feature generation model is a model suitable for generating image features, such as a diffusion model, an autoregressive model, etc. The image feature generation model may be an image feature model capable of generating image features corresponding to at least one target object based on the text content of the text information and the positional features of at least one target object.

[0087] In an embodiment of the present disclosure, when generating image features of at least one target object based on text information and position features of at least one target object, the text information and the position features of at least one target object can be input into an image feature generation model, and the image feature generation model can be used to generate image features for the text information and the coordinates of at least one target object, thereby obtaining image features corresponding to the at least one target object, that is, obtaining image information corresponding to the at least one target object.

[0088] For example, if the text information is "a cat and a dog" and the coordinates of the cat and the dog have been obtained, the present disclosure obtains the image features of the cat and the dog according to the output of the image feature generation model by inputting "a cat and a dog (text information)", "cat (text) + cat coordinates (position features)", and "dog (text) + dog coordinates (position features)" into the image feature generation model.

[0089] It can be seen that in the present disclosure, when image modality information is missing, an image feature generation model can be used to automatically generate image information corresponding to multiple objects based on text prompts and the position coordinates of multiple objects, thereby completing the missing image modality information.

[0090] In other words, it can be seen that the image generation model in the embodiments of the present disclosure not only supports generating target images based on text prompts, coordinate information, and image information, but also supports generating target images based on text prompts alone. In other words, the present disclosure can generate images based on multi-object multi-modal prompts. At the same time, when a modality is missing, the generation model can also be used to supplement the missing modality, thus having a wider range of applications.

[0091] In an optional embodiment, the image generation method further includes the following method steps:

[0092] Step S2124: Use a text encoder to perform text encoding on the text information to obtain global text features, and use a text encoder to perform text encoding on at least one target word to obtain object text features.

[0093] Considering that a text encoder can convert text data into vector or matrix form for machine learning and deep learning tasks, such as part-of-speech tagging, in the embodiments of the present disclosure, when processing text information, a text encoder can be used to perform text encoding processing on the text information to obtain a global text encoding result, i.e., a global text feature. At the same time, a text encoder can also be used to perform text encoding processing on at least one target word segmentation to obtain a noun part encoding result, i.e., an object text feature.

[0094] It is understandable that when processing the image information in the enhanced mark information, an image encoder may be used to perform image feature encoding processing on the image information, thereby obtaining image features, which will not be elaborated herein.

[0095] In an optional embodiment, in step S2123, the enhanced tag information is determined using at least one target word segmentation, at least one location feature of the target object, and at least one image feature of the target object, including the following method steps:

[0096] Step S21231: combining the object text feature, the position feature of at least one target object, and the image feature of at least one target object to obtain enhanced tag information.

[0097] In the disclosed embodiments, when determining enhanced tag information using at least one target segmented word, at least one target object's positional features, and at least one target object's image features, the object text features obtained by performing text encoding processing on the at least one target segmented word, the at least one target object's positional features, and the at least one target object's image features may be combined to obtain the enhanced tag information. For example, the object text features, the at least one target object's positional features, and the at least one target object's image features may be horizontally spliced ​​to obtain the enhanced tag information, which is not a limitation herein.

[0098] For example, if the text information is "a cat and a dog on the grass", the object text features "cat", "dog" and "grass", the position features of at least one target object "coordinates of the cat", "coordinates of the dog" and "coordinates of the grass", and the image features of at least one target object "image of a cat", "image of a dog" and "image of the grass" can be combined to obtain enhanced labeling information.

[0099] In an optional embodiment, in step S213, the text information and the enhanced markup information are combined to obtain multimodal prompt information, including the following method steps:

[0100] Step S2131 : combining global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.

[0101] In the embodiments of the present disclosure, when combining text information with enhanced markup information to obtain multimodal prompt information, global text features, object text features, positional features of at least one target object, and image features of at least one target object may be combined to obtain the multimodal prompt information. For example, the object text features, positional features of at least one target object, and image features of at least one target object may be horizontally spliced ​​together and then vertically spliced ​​together with the global text features to obtain the multimodal prompt information, although this is not a limitation herein.

[0102] For example, if the text information is "a cat and a dog on the grass", the global text feature "a cat and a dog on the grass", the object text features "cat", "dog" and "grass", the position features of at least one target object "coordinates of the cat", "coordinates of the dog" and "coordinates of the grass", and the image features of at least one target object "image of a cat", "image of a dog" and "image of the grass" can be combined to obtain multimodal prompt information.

[0103] In an optional embodiment, in step S21, obtaining multimodal prompt information includes the following method steps:

[0104] Step S211: acquiring text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object;

[0105] Step S212, determining enhanced marking information based on the text information and the additional information;

[0106] Step S213: Combine the text information and the enhanced markup information to obtain multimodal prompt information.

[0107] The text information can be understood as a text prompt input by the user to describe the content of the image to be generated.

[0108] The additional information may be understood as coordinates and / or images input by the user. It is understandable that the additional information includes at least one of the coordinates of at least one target object input by the user and the image of at least one target object input by the user.

[0109] In the embodiment of the present disclosure, by obtaining text information and additional information input by the user, enhanced marking information for determining the position features and image features of at least one target object can be determined based on the obtained text information and additional information, and then multimodal prompt information can be obtained by combining the text information and the enhanced marking information.

[0110] For example, the text information and the enhanced markup information may be vertically spliced ​​to obtain multimodal prompt information, which is not limited here.

[0111] FIG3 is a flowchart of another image generation method according to Example 1 of the present disclosure. As shown in FIG3 , the image generation model of the present disclosure supports taking text information and enhanced markup information as input to generate a target image, that is, it supports combining text modality with coordinate modality and image modality to generate a target image based on multimodal prompt information. In addition, the image generation model of the present disclosure also supports taking only text information as input when the modality is missing, thereby generating corresponding enhanced markup information based on the text information, that is, generating corresponding coordinate modality and image modality based on the text modality, and then generating a target image based on the text modality and the generated coordinate modality and image modality.

[0112] Furthermore, in the image generation process based on multimodal cues, at least one target word (i.e., noun), the coordinates of at least one target object, and the image of at least one target object in the text information can be determined based on the input text information. The text encoder then determines the object text features corresponding to the at least one target word, determines the positional features of the at least one target object based on the coordinates of the at least one target object, and determines the image features of the at least one target object using the image encoder. The object text features, the positional features of the at least one target object, and the image features of the at least one target object are then combined to obtain enhanced tag information. The text information is then embedded and combined with the enhanced tag information, which is then input into the image generation model to ultimately obtain the target image.

[0113] For example, the text information may be "a cat and a dog on the grass." Therefore, it can be determined that the multiple target segmentations include "dog," "cat," and "grass," the coordinates of at least one target object include [coordinates of the dog], [coordinates of the cat], and [coordinates of the grass], and the image of at least one target object includes an image of a dog, an image of a cat, and an image of grass. Then, the object text features corresponding to the multiple target segmentations are determined using a text editor, the position features of at least one target object are determined based on the coordinates of the at least one target object, and the image features of the at least one target object are determined using an image encoder. The object text features, the position features of the at least one target object, and the image features of the at least one target object are then combined to obtain enhanced tag information. The text information "a cat and a dog on the grass" and the enhanced tag information are then combined and input into the image generation model to ultimately obtain the target image.

[0114] In the image generation process based on pure text prompts, to address modality loss, only text information can be input. A word segmentation model is used to determine at least one target word in the text information, and the text encoder is used to determine the object text features corresponding to the at least one target word. A coordinate model is then used to generate positional features for the text information and the at least one target word, obtaining positional features for the at least one target object. An image feature model is then used to generate image features for the text information and the at least one target object's positional features, obtaining image features for the at least one target object. The object text features, the generated positional features for the at least one target object, and the image features for the at least one target object are then combined to obtain enhanced tag information. The text information is then embedded in the image generation model and combined with the enhanced tag information, which is then input into the image generation model to ultimately obtain the target image.

[0115] For example, the text information may be "a cat and a dog on the grass", so it is possible to determine that at least one target segmentation includes "dog", "cat" and "grass" according to the segmentation model, and determine the object text feature corresponding to the at least one target segmentation according to the text encoder. Then, the coordinate model is used to generate position features for the text information and at least one target segmentation, and the position features corresponding to the coordinates of at least one target object ([coordinates of the dog], [coordinates of the cat] and [coordinates of the grass]) are obtained. The image feature model is used to generate image features for the text information and the position features of at least one target object, and the image features corresponding to the image of at least one target object (an image of the dog, an image of the cat and an image of the grass) are obtained. The object text features, the generated position features of at least one target object and the image features of at least one target object are then combined to obtain enhanced tag information, and then the text information is embedded and combined with the enhanced tag information and input into the image generation model to finally obtain the target image.

[0116] As can be seen, given an image-text pair, the present invention obtains object-level text, coordinates, and images, and integrates this information into an "augmented token" for each object. The augmented token is trained as an additional condition along with the textual cues in the diffusion model, enabling the image generation model of the present invention to handle multi-object and multi-modal textual cues.

[0117] Furthermore, to address the problem of missing modalities during inference, that is, to solve the problem of zero-sample image generation, the present disclosure proposes using a coordinate model and an image feature model to generate object-level coordinates and image features based on textual cues. Therefore, the present disclosure can generate target images solely through textual cues or through a combination of various multimodal cues, and can flexibly combine various modalities to generate target images. Furthermore, through a large number of qualitative and quantitative experiments, it has been demonstrated that the present disclosure not only outperforms image generation methods in related technologies, but can also accomplish a wider range of image generation tasks.

[0118] It is easy to understand that the beneficial effects of the image generation method provided by the present disclosure include the following points.

[0119] Beneficial effects (1): Supporting the generation of high-quality images based on multi-modal prompts at the multi-object level, enabling more precise control of the generated images and a wider range of applications;

[0120] Beneficial effect (2): A coordinate model and an image feature model are designed to support the generation of coordinates and image modalities based on text modalities, thereby overcoming the possible modality missing problem;

[0121] Beneficial effect (3): Most of the image generation models in related technologies are generated based on text modalities, while the image generation model disclosed in the present invention can generate target images based on text, image, and coordinate modalities at the same time, and the generated images have higher accuracy;

[0122] Beneficial effect (4): If the image generation model of the related art wants to generate a given object, such as a dog, since there are many breeds of dogs, if you want to generate a specific dog, you generally need to give 3 to 5 pictures and fine-tune the model before the model can learn. However, the image generation model of the present disclosure does not require the fine-tuning process. It only needs to give a dog picture during inference to generate the corresponding breed of dog, that is, it can achieve zero-sample generation. It does not require the participation of the target object during training, so the model training cost is low.

[0123] Beneficial effect (5): The present disclosure adds enhanced tags to text embedding as sampling conditions for the diffusion model to generate images, so that the present disclosure does not need to change the architecture of the diffusion model, but only needs to change the input of the diffusion model, thereby improving the usability of building technology based on the diffusion model.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0125] In addition, it should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0126] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.

[0127] Example 2

[0128] In the operating environment of Example 1, the present disclosure provides an image generation method as shown in FIG4 . A graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. FIG4 is a flow chart of an image generation method according to Example 2 of the present disclosure. As shown in FIG4 , the method includes:

[0129] Step S41, in response to a first control operation performed on a graphical user interface, inputting text information, wherein the text information is used to describe image content to be generated, and the image content includes: at least one target object;

[0130] Step S42, in response to a second control operation performed on the graphical user interface, generating, based on the text information, a position feature of at least one target object and an image feature of at least one target object to obtain enhanced tag information, and performing multimodal image generation on the multimodal prompt information using an image generation model to obtain a target image, wherein the enhanced tag information is used to determine the position feature and the image feature of the at least one target object, and the image generation model is used to generate the target image using a multimodal image generation method, and the multimodal prompt information includes: text information and enhanced tag information;

[0131] Step S43: display the target image in the graphical user interface.

[0132] The graphical user interface in the disclosed embodiments displays at least an image generation scene, in which a user can perform control operations to input textual information describing the content of the image to be generated, control the generation of at least one target object's location features and at least one target object's image features based on the textual information to obtain enhanced tag information, and control the use of an image generation model to generate a multimodal image from the multimodal prompt information to obtain a target image. It is understood that the above-mentioned image generation scene can include, but is not limited to, application scenarios involving image generation in fields such as scriptwriting, design, gaming, e-commerce, education, healthcare, conferences, social networks, financial products, logistics, and navigation.

[0133] The graphical user interface also includes a first control (or a first touch area). When a first touch operation is detected on the first control (or the first touch area), text information input by the user can be obtained. The text information can be input by the user into a text box in the graphical user interface through the first touch operation. The first touch operation can be a click, box selection, check, conditional filtering, etc., which are not limited here.

[0134] The text information can be understood as a text prompt, that is, the image content described in natural language, which is used to describe the image content to be generated. Exemplarily, the natural language can be Chinese, English, Japanese, etc., which are not limited here. In the embodiment of the present disclosure, the image content described by the text information may include at least one target object, and the target object can be understood as a person, animal or object in the image content to be generated, that is, the present disclosure can support multi-object image generation. Exemplarily, the target object can include real people such as boys, girls, doctors, and teachers, and can include virtual characters such as game players and non-player characters (NPCs), and can include animals such as cats, dogs, peacocks, and elephants, and can also include objects such as tables, cars, grass, trees, and train stations, which are not limited here.

[0135] For example, the text information can generate text information corresponding to the user's image generation request. If the user wishes to obtain an image of a cat and a dog on grass, the corresponding text information can be "a cat and a dog on the grass." This text information includes three target objects: a cat, a dog, and grass. It is understood that this text information can also be replaced with "a dog and a cat on the grass" or "there is a cat and a dog on the grass." This disclosure does not limit the description method of the text information.

[0136] The graphical user interface also includes a second control (or a second touch area). When a second touch operation is detected on the second control (or the second touch area), the graphical user interface can generate at least one target object's positional features and at least one target object's image features based on the text information to obtain enhanced marking information, and use an image generation model to generate a multimodal image of the multimodal prompt information to obtain a target image. The second touch operation can be a click, box, check, conditional filter, or other operation, which is not limited here.

[0137] Considering that describing an object only through textual hints will lead to inaccurate image generation, the present disclosure uses enhanced tag information as additional information on the basis of providing textual information, and generates a target image through the textual information and the enhanced tag information, thereby improving the accuracy of image generation.

[0138] The enhanced tag information may include coordinate information and image information, wherein the coordinate information may be the coordinates of at least one target object included in the text information, and is used to determine the positional features of the at least one target object. For example, if a user wishes to obtain an image of a cat and a dog on grass, the coordinate information may include the coordinates of the cat (e.g., [coordinates: 12, 15, 100, 200]), the coordinates of the dog (e.g., [coordinates: 100, 150, 220, 240]), and the coordinates of the grass (e.g., [coordinates: 500, 500, 500, 500]).

[0139] Exemplarily, the coordinates of at least one target object may be coordinates in a two-dimensional rectangular coordinate system, which can represent the coordinate range of the target object, or coordinates in a two-dimensional coordinate system or a three-dimensional coordinate system. The present disclosure does not limit the coordinate system used.

[0140] The image information may be an image of at least one target object included in the text information, and is used to determine the image features of the at least one target object in the text information. For example, if a user wishes to obtain an image of a cat and a dog on grass, the image information may include an image of a cat (e.g., [Image: an image of a cat uploaded by the user]), an image of a dog (e.g., [Image: an image of a dog uploaded by the user]), and an image of grass (e.g., [Image: an image of grass uploaded by the user]).

[0141] Multimodal prompt information can be understood as prompt information in multiple modes, including the above-mentioned text information and enhanced mark information. In the embodiment of the present disclosure, multimodal prompt information includes prompt information in text mode, coordinate mode, and image mode.

[0142] Exemplarily, multimodal prompt information of the same target object can be bundled together to avoid affecting other objects or global conditions. That is, the present disclosure introduces an augmented token, which contains object-level text, coordinates and image information at the same time, and is used to describe an object in the generated image.

[0143] In the disclosed embodiments, the image generation model is a model suitable for generating images. For example, the image generation model can be a pre-trained diffusion-based image generation model, i.e., a diffusion model. The image generation model can also be an autoregressive model, etc., without limitation herein. The image generation model can accurately generate high-quality target images based on multimodal cue information of multiple target objects, i.e., it can generate high-quality images based on multimodal cues at the multi-object level, thereby achieving precise control over the generated images.

[0144] Exemplarily, if the multimodal prompt information of the above example is input into the image generation model, that is, the input text information: a cat and a dog on the grass, coordinate information: the coordinates of the cat, the coordinates of the dog and the coordinates of the grass, image information: the image of a cat, the image of a dog and the image of the grass, then the image generation model of the present invention can accurately output the corresponding target image according to the multimodal prompt information, that is, the image of a cat and a dog on the grass.

[0145] It can be understood that the image generation model in the embodiment of the present disclosure can maintain the original structure of the pre-trained text-to-image model as much as possible so as to integrate extensions to the existing model, and the image generation model in the embodiment of the present disclosure only changes the input of the model and does not need to change the architecture of the model, so it can maintain the usability of the technology built on the basis of the model.

[0146] In an embodiment of the present disclosure, a graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. If the user performs a first control operation on the graphical user interface, the user enters text information for describing the image content of at least one target object to be generated. If the user performs a second control operation on the graphical user interface, such as a submit operation, the position features of at least one target object and the image features of at least one target object can be generated based on the text information, thereby obtaining enhanced marking information. At the same time, the image generation model can be used to perform multimodal image generation on the multimodal prompt information to obtain a target image, and the generated target image can be displayed in the graphical user interface for feedback to the user. The purpose of accurately generating a high-quality target image including multiple target objects through multimodal prompt information is achieved, and more precise control of the generated image can be achieved, generating an image with higher accuracy, and supporting the generation of images of multiple target objects, so that the range of image generation is wider.

[0147] It should be noted that both the first touch operation and the second touch operation can be operations in which a user touches the display screen of the terminal device with a finger and touches the terminal device. The touch operation can include single-point touch and multi-point touch, wherein the touch operation of each touch point can include clicking, long pressing, pressing hard, swiping, etc. The first touch operation and the second touch operation can also be touch operations implemented through input devices such as a mouse and keyboard, which are not limited here.

[0148] The above-mentioned image generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving image generation in the fields of script services, design services, game services, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services. For example: in script services, the required image materials are generated according to the script or storyline; in design services, design drawings are generated according to the user's text, image, and coordinate descriptions; in games, corresponding character images are generated according to the text description of the character image; in e-commerce services, corresponding product images are generated according to the description of the product, etc., which are not limited here.

[0149] According to an embodiment of the present disclosure, a graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. If a user performs a first control operation on the graphical user interface, the user inputs text information for describing the image content of at least one target object to be generated. If the user performs a second control operation on the graphical user interface, such as a submit operation, the position features of at least one target object and the image features of at least one target object can be generated based on the text information, thereby obtaining enhanced marking information. At the same time, an image generation model can be used to perform multimodal image generation on the multimodal prompt information to obtain a target image, and the generated target image is then displayed in the graphical user interface for feedback to the user, thereby achieving the purpose of accurately generating a high-quality target image including multiple target objects through multimodal prompt information, and being able to achieve more precise control over the generated image, generate an image with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby solving the technical problem in the related art of generating corresponding images only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0150] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.

[0151] Example 3

[0152] In the operating environment as in Example 1, the present disclosure provides an image generation method as shown in Figure 5. Figure 5 is a flow chart of an image generation method according to Example 3 of the present disclosure. As shown in Figure 5, the method includes:

[0153] Step S51: Receive a currently input multimodal conversation request, wherein the information carried in the multimodal conversation request includes: multimodal conversation text information and multimodal conversation enhanced tag information; the multimodal conversation text information is used to describe the content of the multimodal conversation image to be generated; the multimodal conversation image content includes: at least one object; and the multimodal conversation enhanced tag information is used to determine the location characteristics and image characteristics of the at least one object;

[0154] Step S52: Using an image generation model to generate a multimodal image for the multimodal conversation request to obtain a multimodal conversation image, wherein the image generation model is used to generate the multimodal conversation image using a multimodal image generation method;

[0155] Step S53: Feedback a multimodal dialogue response, wherein the information carried in the multimodal dialogue response includes: a multimodal dialogue image.

[0156] A multimodal conversation request can be understood as a conversation request initiated by a user to a computer or robot. This multimodal conversation request carries multimodal conversation text information and multimodal conversation enhancement markup information. The multimodal conversation text information can be understood as a text prompt, i.e., a natural language description of the multimodal conversation image content to be generated. For example, the natural language can be Chinese, English, Japanese, etc., without limitation.

[0157] In embodiments of the present disclosure, the multimodal conversation image content described by the multimodal conversation text information may include at least one object. An object can be understood as a person, animal, or object within the multimodal conversation image content to be generated. This means that the present disclosure supports multi-object image generation. For example, objects can include real people such as boys, girls, doctors, and teachers; virtual characters such as game players and non-player characters (NPCs); animals such as cats, dogs, peacocks, and elephants; and objects such as tables, cars, grass, trees, and train stations, without limitation.

[0158] For example, the multimodal conversation text information may be text information corresponding to a multimodal conversation request input by a user. If the user wishes to obtain an image of a cat and a dog on grass, the corresponding multimodal conversation text information may be "a cat and a dog on the grass." This multimodal conversation text information includes three objects: a cat, a dog, and grass. It is understood that this multimodal conversation text information may also be replaced with "a dog and a cat on the grass" or "there is a cat and a dog on the grass." This disclosure does not limit the description of the multimodal conversation text information.

[0159] Considering that describing objects only through multimodal conversation text prompts will lead to inaccurate generation of multimodal conversation images, the present disclosure, on the basis of providing multimodal conversation text information, uses multimodal conversation enhancement tag information as additional information, and jointly generates multimodal conversation images through the multimodal conversation text information and the multimodal conversation enhancement tag information, thereby improving the accuracy of multimodal conversation image generation.

[0160] Multimodal conversation enhancement tagging information may include coordinate information and image information. The coordinate information may be the coordinates of at least one object included in the multimodal conversation text information, used to determine the location features of the at least one object. For example, if a user wishes to obtain an image of a cat and a dog on grass, the coordinate information may include the coordinates of the cat (e.g., [coordinates: 12, 15, 100, 200]), the coordinates of the dog (e.g., [coordinates: 100, 150, 220, 240]), and the coordinates of the grass (e.g., [coordinates: 500, 500, 500, 500]).

[0161] Exemplarily, the coordinates of at least one object may be coordinates in a two-dimensional rectangular coordinate system, which can represent the coordinate range of the target object, or coordinates in a two-dimensional coordinate system or a three-dimensional coordinate system. The present disclosure does not limit the coordinate system used.

[0162] The image information may be an image of at least one object included in the multimodal conversation text information, and is used to determine the image features of the at least one object in the multimodal conversation text information. For example, if a user wishes to obtain an image of a cat and a dog on grass, the image information may include an image of a cat (e.g., [Image: a user-uploaded image of their own cat]), an image of a dog (e.g., [Image: a user-uploaded image of their own dog]), and an image of grass (e.g., [Image: a user-uploaded image of grass]) that the user needs to generate.

[0163] It can be seen that the multimodal dialogue request in the embodiment of the present disclosure carries prompt information of multiple modalities, including the above-mentioned multimodal dialogue text information and multimodal dialogue enhancement marking information, that is, the multimodal prompt information includes prompt information of text modality, coordinate modality and image modality.

[0164] Exemplarily, multiple modal prompt information of the same object can be bundled together to avoid affecting other objects or global conditions. That is, the present disclosure introduces an augmented token, which contains object-level text, coordinates and image information at the same time, and is used to describe an object in the generated image.

[0165] In the disclosed embodiments, the image generation model is a model suitable for generating images. For example, the image generation model can be a pretrained diffusion-based image generation model, i.e., a diffusion model. The image generation model can also be an autoregressive model, etc., without limitation. The image generation model can accurately generate high-quality multimodal conversation images based on multimodal conversation requests. Specifically, it can generate high-quality images based on multimodal prompts at the multi-object level, thereby achieving precise control over the generation of multimodal conversation images.

[0166] Exemplarily, if an image generation model is used to generate a multimodal image for the multimodal conversation request in the above example, that is, the multimodal conversation text information: a cat and a dog on the grass, coordinate information: the coordinates of the cat, the coordinates of the dog and the coordinates of the grass, and image information: an image of a cat, an image of a dog and an image of the grass are input into the image generation model, then the image generation model of the present disclosure can accurately output the corresponding multimodal conversation image according to the multimodal conversation request, that is, an image of a cat and a dog on the grass.

[0167] It can be understood that the image generation model in the embodiment of the present disclosure can maintain the original structure of the pre-trained text-to-image model as much as possible so as to integrate extensions to the existing model, and the image generation model in the embodiment of the present disclosure only changes the input of the model and does not need to change the architecture of the model, so it can maintain the usability of the technology built on the basis of the model.

[0168] A multimodal conversation reply can be understood as the response provided by a computer or robot to a user's multimodal conversation request. This response corresponds to the user's multimodal conversation request. The multimodal conversation reply includes a multimodal conversation image, which is an image containing the to-be-generated multimodal conversation image content described in the multimodal conversation request.

[0169] Steps S51 to S53 described above can be applied to human-computer dialogue scenarios, i.e., scenarios involving a conversation between a user and a computer or robot. As shown in FIG6 , FIG6 is a schematic diagram of a human-computer dialogue scenario according to Embodiment 3 of the present disclosure, in which a user inputs a multimodal dialogue request to a computer or robot, and the computer or robot can provide the user with a multimodal dialogue response corresponding to the multimodal dialogue request.

[0170] It is understandable that the conversation between the user and the computer or robot can be achieved through voice recognition and natural language processing technology, or through text communication, which is not limited here.

[0171] In the disclosed embodiments, by receiving a multimodal conversation request from a user, including multimodal conversation text information and multimodal conversation enhancement markup information, the system can determine the content of a multimodal conversation image to be generated based on the multimodal conversation text information and the positional features and image features of at least one object in the multimodal conversation image content to be generated based on the multimodal conversation enhancement markup information. An image generation model is then used to generate a multimodal image based on the multimodal conversation request, thereby obtaining a multimodal conversation image. A multimodal conversation reply containing the multimodal conversation image is then fed back to the user. This achieves the goal of accurately generating a high-quality multimodal conversation image containing multiple objects based on the multimodal conversation request, while also enabling more precise control over the generated multimodal conversation image, resulting in a more accurate multimodal conversation image. Furthermore, the system supports the generation of multimodal conversation images for multiple objects, expanding the scope of multimodal conversation image generation.

[0172] The above-mentioned image generation method provided by the embodiment of the present disclosure can also be applied to, but is not limited to, application scenarios involving image generation in the fields of script services, design services, game services, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services. For example: in design services, design drawings are generated based on the user's text, image, and coordinate descriptions; in games, corresponding character drawings are generated based on the text description of the character image; in e-commerce services, corresponding product drawings are generated based on the description of the product, etc., which are not limited here.

[0173] According to the disclosed embodiments, by receiving a multimodal conversation request from a user, including multimodal conversation text information and multimodal conversation enhancement markup information, the system can determine the content of the multimodal conversation image to be generated based on the multimodal conversation text information and the positional features and image features of at least one object in the multimodal conversation image content to be generated based on the multimodal conversation enhancement markup information. An image generation model is then used to generate a multimodal image based on the multimodal conversation request, thereby obtaining a multimodal conversation image. A multimodal conversation reply containing the multimodal conversation image is then fed back to the user. This achieves the goal of accurately generating high-quality multimodal conversation images containing multiple objects based on the multimodal conversation request, and enables more precise control over the generated multimodal conversation image, resulting in a more accurate multimodal conversation image. Furthermore, the system supports the generation of multimodal conversation images for multiple objects, expanding the range of multimodal conversation image generation. This addresses the technical issues in related arts where generating corresponding images solely through text prompts results in low image accuracy and a limited image generation range.

[0174] In an optional embodiment, in step S51, receiving a currently input multimodal dialogue request includes the following method steps:

[0175] Step S511, obtaining multimodal conversation text information;

[0176] Step S512: generating a position feature of at least one object and an image feature of at least one object based on the multimodal conversation text information to obtain multimodal conversation enhanced tag information;

[0177] Step S513: Combine the multimodal dialogue text information and the multimodal dialogue enhancement mark information to obtain a multimodal dialogue request.

[0178] In an embodiment of the present disclosure, when receiving a currently input multimodal dialogue request, the multimodal dialogue text information can be obtained first, and then the position features of at least one object and the image features of at least one object can be generated based on the multimodal dialogue text information, that is, object-level coordinates and images are generated according to the multimodal dialogue text prompt, thereby obtaining corresponding multimodal dialogue enhanced tag information, and then the multimodal dialogue text information and the multimodal dialogue enhanced tag information are combined to obtain the multimodal dialogue request.

[0179] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.

[0180] Example 4

[0181] According to an embodiment of the present disclosure, an embodiment of a device for implementing the above-mentioned image generation is also provided. FIG7 is a structural schematic diagram of an image generation device according to embodiment 4 of the present disclosure. As shown in FIG7 , the device includes:

[0182] An acquisition module 701 is configured to acquire multimodal prompt information, wherein the multimodal prompt information includes text information and enhanced markup information, wherein the text information is used to describe the content of an image to be generated, the image content includes at least one target object, and the enhanced markup information is used to determine positional features and image features of the at least one target object;

[0183] The image generation module 702 is configured to use an image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.

[0184] Optionally, the above-mentioned acquisition module 701 is also configured to: acquire text information; generate position features of at least one target object and image features of at least one target object based on the text information to obtain enhanced marking information; and combine the text information and the enhanced marking information to obtain multimodal prompt information.

[0185] Optionally, the above-mentioned acquisition module 701 is also configured to: perform part-of-speech analysis on the text information, select at least one target participle, wherein the at least one target participle meets the preset part-of-speech requirements, and the at least one target participle is used to determine at least one target object; generate position features of at least one target object based on the text information and the at least one target participle, and generate image features of at least one target object based on the text information and the position features of the at least one target object; determine enhanced tag information using the at least one target participle, the position features of the at least one target object, and the image features of the at least one target object.

[0186] Optionally, the acquisition module 701 is further configured to: generate position features for the text information and at least one target word segment using a position feature generation model to obtain a position feature of at least one target object.

[0187] Optionally, the acquisition module 701 is further configured to: generate image features from the text information and the position features of the at least one target object using an image feature generation model to obtain image features of the at least one target object.

[0188] Optionally, it also includes: an encoding module, which is configured to: use a text encoder to perform text encoding on text information to obtain global text features, and use a text encoder to perform text encoding on at least one target word to obtain object text features.

[0189] Optionally, the acquisition module 701 is further configured to: combine the object text feature, the position feature of at least one target object, and the image feature of at least one target object to obtain enhanced marking information.

[0190] Optionally, the acquisition module 701 is further configured to: combine global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.

[0191] Optionally, the above-mentioned acquisition module 701 is also configured to: acquire text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object; determine enhanced marking information based on the text information and additional information; combine the text information and the enhanced marking information to obtain multimodal prompt information.

[0192] By adopting the embodiment of the present disclosure, by obtaining multimodal prompt information including text information and enhanced mark information, it is possible to determine the image content to be generated based on the text information, determine the position features and image features of at least one target object in the image content based on the enhanced mark information, and then use the image generation model to perform multimodal image generation on the multimodal prompt information to obtain the target image, thereby achieving the purpose of accurately generating high-quality target images including multiple target objects through multimodal prompt information. In addition, it is possible to achieve more precise control over the generated image, generate images with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby achieving more accurate, more diverse, and high-quality image generation that supports multiple target objects, so as to meet the technical effect of image generation needs in different fields, thereby solving the technical problem in the related art of generating corresponding images only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0193] It should be noted that the acquisition module 701 and the image generation module 702 correspond to steps S21 and S22 in Example 1. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the modules can also run in the server 10 provided in Example 1.

[0194] According to an embodiment of the present disclosure, another embodiment of a device for implementing the above-mentioned image generation is also provided. Figure 8 is a schematic structural diagram of another image generation device according to Example 4 of the present disclosure. A graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. As shown in Figure 8, the device includes:

[0195] The first response module 801 is configured to respond to a first control operation performed on the graphical user interface and input text information, wherein the text information is used to describe the image content to be generated, and the image content includes: at least one target object;

[0196] The second response module 802 is configured to, in response to a second control operation performed on the graphical user interface, generate, based on the text information, a position feature of at least one target object and an image feature of at least one target object to obtain enhanced tag information, and perform multimodal image generation on the multimodal prompt information using an image generation model to obtain a target image, wherein the enhanced tag information is used to determine the position feature and the image feature of the at least one target object, and the image generation model is used to generate the target image using a multimodal image generation method, and the multimodal prompt information includes: text information and enhanced tag information;

[0197] The display module 803 is configured to display the target image in the graphical user interface.

[0198] According to an embodiment of the present disclosure, a graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. If a user performs a first control operation on the graphical user interface, the user inputs text information for describing the image content of at least one target object to be generated. If the user performs a second control operation on the graphical user interface, such as a submit operation, the position features of at least one target object and the image features of at least one target object can be generated based on the text information, thereby obtaining enhanced marking information. At the same time, an image generation model can be used to perform multimodal image generation on the multimodal prompt information to obtain a target image, and the generated target image is then displayed in the graphical user interface for feedback to the user, thereby achieving the purpose of accurately generating a high-quality target image including multiple target objects through multimodal prompt information, and being able to achieve more precise control over the generated image, generate an image with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby solving the technical problem in the related art of generating corresponding images only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0199] It should be noted that the first response module 801, the second response module 802, and the display module 803 correspond to steps S41 to S43 in Example 2. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 2. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also run in the server 10 provided in Example 1.

[0200] According to an embodiment of the present disclosure, another embodiment of a device for implementing the above-mentioned image generation is also provided. FIG9 is a structural schematic diagram of another image generation device according to embodiment 4 of the present disclosure. As shown in FIG9 , the device includes:

[0201] Receiving module 901 is configured to receive a currently input multimodal conversation request, wherein the information carried in the multimodal conversation request includes: multimodal conversation text information and multimodal conversation enhancement mark information, wherein the multimodal conversation text information is used to describe the content of the multimodal conversation image to be generated, the multimodal conversation image content includes: at least one object, and the multimodal conversation enhancement mark information is used to determine the location characteristics and image characteristics of the at least one object;

[0202] A generation module 902 is configured to generate a multimodal image based on the multimodal conversation request using an image generation model to obtain a multimodal conversation image, wherein the image generation model is configured to generate the multimodal conversation image using a multimodal image generation method;

[0203] The feedback module 903 is configured to feedback the multimodal dialogue response, wherein the information carried in the multimodal dialogue response includes: the multimodal dialogue image.

[0204] Optionally, the above-mentioned receiving module 901 is also configured to: obtain multimodal conversation text information; generate position features of at least one object and image features of at least one object based on the multimodal conversation text information, and obtain multimodal conversation enhanced tag information; combine the multimodal conversation text information and the multimodal conversation enhanced tag information to obtain a multimodal conversation request.

[0205] According to the disclosed embodiments, by receiving a multimodal conversation request from a user, including multimodal conversation text information and multimodal conversation enhancement markup information, the system can determine the content of the multimodal conversation image to be generated based on the multimodal conversation text information and the positional features and image features of at least one object in the multimodal conversation image content to be generated based on the multimodal conversation enhancement markup information. An image generation model is then used to generate a multimodal image based on the multimodal conversation request, thereby obtaining a multimodal conversation image. A multimodal conversation reply containing the multimodal conversation image is then fed back to the user. This achieves the goal of accurately generating high-quality multimodal conversation images containing multiple objects based on the multimodal conversation request, and enables more precise control over the generated multimodal conversation image, resulting in a more accurate multimodal conversation image. Furthermore, the system supports the generation of multimodal conversation images for multiple objects, expanding the range of multimodal conversation image generation. This addresses the technical issues in related arts where generating corresponding images solely through text prompts results in low image accuracy and a limited image generation range.

[0206] It should be noted that the receiving module 901, generating module 902, and feedback module 903 correspond to steps S51 to S53 in Example 3. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 3. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the modules can also run in the server 10 provided in Example 1.

[0207] It should be noted that the preferred implementation scheme involved in the above embodiments of the present disclosure is the same as the solution provided in Example 1, as well as the application scenario and implementation process, but is not limited to the solution provided in Example 1.

[0208] Example 5

[0209] The embodiment of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.

[0210] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0211] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the image generation method: obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, and the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of at least one target object; using the image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.

[0212] Optionally, Figure 10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present disclosure. As shown in Figure 10, the computer terminal A may include: one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.

[0213] Among them, the memory can be configured to store software programs and modules, such as the program instructions / modules corresponding to the image generation method and device in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the stored software programs and modules, that is, realizing the above-mentioned image generation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0214] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, and the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of at least one target object; using the image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.

[0215] Optionally, the processor may also execute the program code of the following steps: obtaining text information; generating the position feature of at least one target object and the image feature of at least one target object based on the text information to obtain enhanced marking information; combining the text information and the enhanced marking information to obtain multimodal prompt information.

[0216] Optionally, the processor may also execute the program code for the following steps: performing part-of-speech analysis on the text information, selecting at least one target participle, wherein the at least one target participle satisfies a preset part-of-speech requirement, and the at least one target participle is used to determine at least one target object; generating a position feature of the at least one target object based on the text information and the at least one target participle, and generating an image feature of the at least one target object based on the text information and the position feature of the at least one target object; and determining enhanced tag information using the at least one target participle, the position feature of the at least one target object, and the image feature of the at least one target object.

[0217] Optionally, the processor may further execute program code of the following steps: using a position feature generation model to generate position features for text information and at least one target word segmentation to obtain a position feature of at least one target object.

[0218] Optionally, the processor may further execute program code of the following steps: using an image feature generation model to generate image features for the text information and the position features of at least one target object to obtain image features of the at least one target object.

[0219] Optionally, the processor may further execute program code of the following steps: performing text encoding on text information using a text encoder to obtain global text features, and performing text encoding on at least one target word using a text encoder to obtain object text features.

[0220] Optionally, the processor may further execute program code of the following steps: combining object text features, position features of at least one target object, and image features of at least one target object to obtain enhanced marking information.

[0221] Optionally, the processor may further execute program code of the following steps: combining global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.

[0222] Optionally, the processor may also execute the program code of the following steps: obtaining text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object; determining enhanced marking information based on the text information and the additional information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.

[0223] By adopting the embodiment of the present disclosure, by obtaining multimodal prompt information including text information and enhanced mark information, it is possible to determine the image content to be generated based on the text information, determine the position features and image features of at least one target object in the image content based on the enhanced mark information, and then use the image generation model to perform multimodal image generation on the multimodal prompt information to obtain the target image, thereby achieving the purpose of accurately generating high-quality target images including multiple target objects through multimodal prompt information. In addition, it is possible to achieve more precise control over the generated image, generate images with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby achieving more accurate, more diverse, and high-quality image generation that supports multiple target objects, so as to meet the technical effect of image generation needs in different fields, thereby solving the technical problem in the related art of generating corresponding images only through text prompts, resulting in low accuracy of the generated images and a limited image generation range.

[0224] Those skilled in the art will appreciate that the structure shown in FIG10 is merely illustrative, and the computer terminal A may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG10 does not limit the structure of the aforementioned electronic devices. For example, the computer terminal A may include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG10 , or may have a configuration different from that shown in FIG10 .

[0225] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0226] Example 6

[0227] The embodiment of the present disclosure further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the image generation method provided in the first embodiment.

[0228] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0229] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, and the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of at least one target object; using an image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.

[0230] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining text information; generating position features of at least one target object and image features of at least one target object based on the text information to obtain enhanced marking information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.

[0231] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing part-of-speech analysis on text information, selecting at least one target participle, wherein the at least one target participle meets preset part-of-speech requirements, and the at least one target participle is used to determine at least one target object; generating position features of at least one target object based on the text information and the at least one target participle, and generating image features of at least one target object based on the text information and the position features of the at least one target object; and determining enhanced tag information using the at least one target participle, the position features of the at least one target object, and the image features of the at least one target object.

[0232] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: using a position feature generation model to generate position features for text information and at least one target word segmentation to obtain position features of at least one target object.

[0233] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: using an image feature generation model to generate image features for text information and position features of at least one target object to obtain image features of at least one target object.

[0234] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: using a text encoder to perform text encoding on text information to obtain global text features, and using a text encoder to perform text encoding on at least one target word to obtain object text features.

[0235] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: combining object text features, position features of at least one target object, and image features of at least one target object to obtain enhanced tag information.

[0236] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: combining global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.

[0237] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object; determining enhanced marking information based on the text information and the additional information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.

[0238] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.

[0239] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0240] In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0241] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0242] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0243] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0244] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure.

Claims

1. A method for generating an image, comprising: Acquiring multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, the image content includes: at least one target object, and the enhanced mark information is used to determine the position feature and image feature of the at least one target object; An image generation model is used to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image in a multimodal image generation manner.

2. The image generation method according to claim 1, wherein: Acquiring the multimodal prompt information includes: Obtaining the text information; Based on the text information, respectively generate the position feature of the at least one target object and the image feature of the at least one target object to obtain the enhanced marking information; The text information is combined with the enhanced mark information to obtain the multimodal prompt information.

3. The image generation method according to claim 2, wherein: Generating the position feature and the image feature of the at least one target object based on the text information respectively, and obtaining the enhanced marking information includes: Performing part-of-speech analysis on the text information, selecting at least one target participle, wherein the at least one target participle meets a preset part-of-speech requirement, and the at least one target participle is used to determine the at least one target object; Generating a position feature of the at least one target object based on the text information and the at least one target word segmentation, and generating an image feature of the at least one target object based on the text information and the position feature of the at least one target object; The enhanced marking information is determined by using the at least one target word segment, the position feature of the at least one target object, and the image feature of the at least one target object.

4. The image generation method according to claim 3, wherein: Generating the location feature of the at least one target object based on the text information and the at least one target word segmentation includes: A position feature generation model is used to generate position features for the text information and the at least one target word segmentation to obtain the position features of the at least one target object.

5. The image generation method according to claim 3, wherein: Generating the image feature of the at least one target object based on the text information and the position feature of the at least one target object comprises: An image feature generation model is used to generate image features for the text information and the position features of the at least one target object to obtain image features of the at least one target object.

6. The image generation method according to claim 3, wherein: The image generation method further comprises: The text information is encoded by a text encoder to obtain global text features, and The text encoder is used to perform text encoding on the at least one target word segment to obtain object text features.

7. The image generation method according to claim 6, wherein: Determining the enhanced marking information by using the at least one target word segmentation, the position feature of the at least one target object, and the image feature of the at least one target object includes: The object text feature, the position feature of the at least one target object, and the image feature of the at least one target object are combined to obtain the enhanced marking information.

8. The image generation method according to claim 6, wherein: Combining the text information with the enhanced mark information to obtain the multimodal prompt information includes: The global text feature, the object text feature, the position feature of the at least one target object, and the image feature of the at least one target object are combined to obtain the multimodal prompt information.

9. The image generation method according to claim 1, wherein: Acquiring the multimodal prompt information includes: Acquire the text information and additional information, wherein the additional information includes at least one of the following: location information of the at least one target object, image information of the at least one target object; Determining the enhanced marking information based on the text information and the additional information; The text information is combined with the enhanced mark information to obtain the multimodal prompt information.

10. An image generation method, providing a graphical user interface through a terminal device, wherein the content displayed by the graphical user interface at least partially includes an image generation scene, the image generation method comprising: In response to a first control operation performed on the graphical user interface, inputting text information, wherein the text information is used to describe image content to be generated, the image content including: at least one target object; In response to a second control operation performed on the graphical user interface, based on the text information, respectively generate the position feature of the at least one target object and the image feature of the at least one target object to obtain enhanced mark information, and use an image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the enhanced mark information is used to determine the position feature and the image feature of the at least one target object, the image generation model is used to generate the target image in a multimodal image generation manner, and the multimodal prompt information includes: the text information and the enhanced mark information; The target image is displayed within the graphical user interface.

11. A method for generating an image, comprising: Receive a currently input multimodal dialogue request, wherein the information carried in the multimodal dialogue request includes: multimodal dialogue text information and multimodal dialogue enhancement mark information, the multimodal dialogue text information is used to describe the multimodal dialogue image content to be generated, the multimodal dialogue image content includes: at least one object, and the multimodal dialogue enhancement mark information is used to determine the position feature and image feature of the at least one object; The image generation model is used to generate a multimodal image for the multimodal dialogue request to obtain a multimodal dialogue image, wherein the image generation model is used to generate the multimodal dialogue image in a multimodal image generation manner. Dynamic dialogue image; Feedback is given of a multimodal dialogue response, wherein the information carried in the multimodal dialogue response includes: the multimodal dialogue image.

12. The image generation method according to claim 11, wherein: Receiving the currently input multimodal dialogue request includes: Acquiring the multimodal conversation text information; Based on the multimodal conversation text information, respectively generate the position feature of the at least one object and the image feature of the at least one object to obtain the multimodal conversation enhanced tag information; The multimodal dialogue text information is combined with the multimodal dialogue enhancement mark information to obtain the multimodal dialogue request.

13. An electronic device, comprising: A memory storing an executable program; A processor, configured to run the program, wherein the program, when run, executes the image generation method described in any one of claims 1 to 12.

14. A computer-readable storage medium comprising a stored executable program, wherein: When the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the image generation method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Power transmission inspection image generation method and device based on multi-modal data

    CN114937181A

  • Image generation method and server

    CN116597039A

  • Data processing method and device based on AIGC, electronic equipment and storage medium

    CN116704062A

  • Image generation method, electronic equipment and computer readable storage medium

    CN117689749A

  • Image segmentation using text embedding

    US20220156992A1