Visual Data Generation Method, Multimodal Model Training Method, Device, and Medium
Through fine-tuning of the literary and biographical graphics model and multimodal model training, visual data with fusion of multiple styles is generated, which solves the problem that multiple style images cannot be generated in the existing technology, and achieves high-quality multi-style image generation.
Patent Information
- Application Number
- CN202411619665.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-13
AI Technical Summary
Existing literary and biographical graphics models cannot generate visual content that integrates multiple styles, and it is difficult to meet users' needs for diversified image styles.
By fine-tuning the existing literary and biographical model, visual data that is fused with multiple styles is generated using multimodal model training method. By obtaining the style trigger words and image area content information in the description information, the pre-trained multimodal model generates output visual data, so that there are differences in the image styles of at least two areas in the output visual data.
It realizes high-quality generation of multi-style fusion images, improves the accuracy and artistic quality of image generation, meets users' needs for multi-style fusion images, and lowers the threshold for users to use AI technology.
Smart Images

Figure CN119312796B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image generation technology, and in particular to a method for generating visual data, a method for training a multi-modal model, an apparatus, and a medium. Background Art
[0002] In AIGC applications, the text-to-image model plays a key role. Users can generate images corresponding to the text in terms of style and content by inputting text. Further, the text-to-video model generates a video corresponding to a single style of the input text by the user. For example, in related technologies, in order to achieve the diversity of image styles, the specific style trigger words in the training data can be improved. However, in this method, the text-to-image model can only generate images of a single style and cannot meet the user's generation requirements for fusing multiple styles in one image.
[0003] Therefore, the large models in the prior art cannot generate visual content that fuses multiple image styles. Summary of the Invention
[0004] To overcome the problems existing in the related technologies, this application provides a method for generating visual data, a method for training a multi-modal model, an apparatus, and a medium. Based on the fine-tuning of the existing text-to-image model, it is possible to generate images that fuse multiple styles, so as to meet the user's expectations for generating images of multiple styles.
[0005] The first aspect of this application provides a method for generating visual data, including:
[0006] Obtain the description information of the visual data to be generated; wherein, the description information corresponds to style trigger words and the image region content information corresponding to the style trigger words;
[0007] Based on the description information and a pre-trained multi-modal model, obtain the output visual data;
[0008] The image styles of at least two regions in the output visual data are different.
[0009] Optionally, the obtaining of the description information of the visual data to be generated includes:
[0010] Obtain the prompt information input by the user; wherein, the prompt information includes at least text information;
[0011] Perform natural language processing on the text information in the prompt information to obtain a processing result;
[0012] According to the processing result, obtain the style trigger words indicating the visual data to be generated and the corresponding region content information;
[0013] Generate structured description information according to the style trigger word and the corresponding region content information.
[0014] Optionally, generating structured description information according to the style trigger word and the corresponding region content information includes:
[0015] Based on the processing result, match the corresponding style trigger word from the preset word library;
[0016] Determine the image region content information corresponding to each of the matched style trigger words;
[0017] Generate structured description information according to the matched style trigger word and the corresponding region content information.
[0018] Optionally, the matching the corresponding style trigger word from the preset word library includes:
[0019] Calculate the similarity between the natural language processing result and the preset word library, and obtain the style trigger words in the preset word library that meet the preset similarity rules.
[0020] The second aspect of this application provides a training method for a multimodal model as described in any one of the above, including:
[0021] Obtain an initial model for image generation;
[0022] Obtain an image data set, including image data of multiple styles;
[0023] Perform structured annotation on the image data set, including annotating the image region content information of the image and the style trigger word of the image style;
[0024] Train the initial model based on the structured annotation data to obtain a trained multimodal model.
[0025] Optionally, the image data only includes one image region, and the performing structured annotation on the image data set includes:
[0026] Annotate the image data to obtain the content information of the image region in the image data and the style trigger word corresponding to the image region.
[0027] Optionally, the image data includes two or more image regions, and the performing structured annotation on the image data set includes:
[0028] Perform image segmentation on the image data to obtain the at least two image regions;
[0029] Respectively annotate the at least two image regions obtained by image segmentation, including the image region content information of the image region and the style trigger word of the image style.
[0030] After annotating the image data with the image style, the method further includes: for the same image style, generating a set of style trigger words, where the set of style trigger words includes synonyms or near-synonyms.
[0031] The third aspect of the present application provides a visual data generation device, including:
[0032] An acquisition module, configured to obtain description information of the visual data to be generated; wherein, the description information corresponds to style trigger words and image region content information corresponding to the style trigger words;
[0033] A processing module, configured to obtain output visual data based on the description information and a pre-trained multimodal model; there are differences in the image styles of at least two regions in the output visual data.
[0034] Optionally, the acquisition module is configured to:
[0035] Obtain prompt information input by the user; wherein, the prompt information includes at least text information;
[0036] Perform natural language processing on the text information in the prompt information to obtain a processing result;
[0037] According to the processing result, obtain style trigger words indicating the visual data to be generated and corresponding region content information;
[0038] Generate structured description information according to the style trigger words and corresponding region content information.
[0039] Optionally, the acquisition module is further configured to:
[0040] Match corresponding style trigger words from a preset word library based on the processing result;
[0041] Determine image region content information corresponding to each of the matched style trigger words;
[0042] Generate structured description information according to the matched style trigger words and corresponding region content information.
[0043] Optionally, the acquisition module is further configured to:
[0044] Calculate the similarity between the processing result and a preset word library, and obtain style trigger words in the preset word library that meet the preset similarity rules.
[0045] The fourth aspect of the present application provides a device for training a visual data generation model, including:
[0046] An initial model acquisition model, configured to obtain an initial model for image generation;
[0047] An image dataset acquisition module for obtaining an image dataset, including image data of multiple styles;
[0048] An annotation module for performing structured annotation on the image dataset, including annotating the content information of the image region of the image and the style trigger words of the image style;
[0049] A training module for training the initial model based on the structured annotation data to obtain a trained multi-modal model.
[0050] Optionally, the annotation module is further configured to:
[0051] Annotate the image data to obtain the content information of the image region in the image data and the style trigger words corresponding to the image region.
[0052] Optionally, the annotation module is further configured to:
[0053] Perform image segmentation on the image data to obtain the at least two image regions;
[0054] Respectively annotate the at least two image regions obtained by image segmentation, including the content information of the image region and the style trigger words of the image style.
[0055] Optionally, the annotation module is further configured to: for the same image style, generate a set of style trigger words, and the set of style trigger words includes synonyms or near-synonyms.
[0056] The fifth aspect of the present application provides a computer-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described in any one of the above.
[0057] The sixth aspect of the present application provides an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading the executable instructions from the memory and executing the instructions to implement the above-described method for generating visual data or the method for generating a multi-modal model.
[0058] It can be seen that in the first aspect, a visual data generation method provided by the present application obtains description information of visual data to be generated; obtains a style trigger word and image region content information corresponding to the style trigger word according to the description information; and obtains output visual data based on the description information and a pre-trained multi-modal model, so that there are differences in the image styles of at least two regions in the output visual data, effectively improving the accuracy and artistry of image generation, and being able to generate images with multi-style fusion according to the needs of users, meeting the needs of users for generating images with multi-style fusion.
[0059] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious, where in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.
[0061] Figure 1 is a schematic flowchart of a visual data generation method shown in an embodiment of the present application.
[0062] Figure 2 is a schematic diagram of a target image generated according to description information shown in an embodiment of the present application.
[0063] Figure 3 is a schematic flowchart of a multi-modal model training method shown in an embodiment of the present application.
[0064] Figure 4 is a schematic diagram of an initial image input by a user shown in an embodiment of the present application.
[0065] Figure 5 is a schematic diagram of another target image generated according to description information shown in an embodiment of the present application.
[0066] Figure 6 is a schematic structural diagram of a visual data generation device shown in an embodiment of the present application.
[0067] Figure 7 is a schematic structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] Preferred embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0069] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0070] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0071] In the embodiments of the present application, it can be applied to AIGC (artificial intelligence generated content) applications for image generation or video generation. Among them, an image corresponding to the text can be generated through a text-to-image model, and a video corresponding to the text can be generated through a text-to-video model. The text-to-image technology refers to generating an image corresponding to the text by inputting a text description and using a text-to-image model. The text-to-image model realizes the ability to understand the content and intention in the text description and generate an image that conforms to the text description through training on text-image data. For example, the applications of the text-to-image model include generating landscape paintings, task portraits, cartoon images, etc., and can generate corresponding image content according to specific user requirements, such as specific scenes, task postures, color combinations, etc.
[0072] In the related art, the text-to-image model can generate high-quality single-style images, but there are obvious deficiencies in multi-style fusion. The existing text-to-image model can only generate images of a specific single style by fine-tuning the trained base model, so that images of a certain specific painting style can be generated. Therefore, it is difficult to meet the user's demand for diverse image styles.
[0073] SeeFigure 1 , the visual data generation method of the present application includes:
[0074] S100. Obtain the description information of the visual data to be generated; wherein, the description information corresponds to a style trigger word and the image area content information corresponding to the style trigger word.
[0075] Among them, the visual data may include static image data and dynamic video data. The video data generally refers to a continuous image sequence and is composed of a group of continuous images. The image data described below refers to static images or the image data constituting the video data.
[0076] The description information is used to indicate the content of the visual data to be generated. This description information can be the information directly input by the user or the information automatically generated according to the user's selection of operation instructions (such as instructions for automatically generating the text description of the visual data to be generated, instructions for changing the style of the visual data to be generated, instructions for selecting the reference image of the visual data to be generated, etc.).
[0077] The style trigger word is used to indicate the image style of the visual data to be generated. The image area content information is used to indicate the information of the area or element of the visual data to be generated, such as area position description, element name, element color, etc. Through this image area content information, one or more area ranges or element objects of the visual data to be generated can be located, as well as the image content (i.e., the image frame) within the area range or the image content of the element object.
[0078] Optionally, obtaining the description information of the visual data to be generated includes: obtaining the prompt information input by the user; wherein, the prompt information at least includes text information; performing natural language processing on the text information in the prompt information to obtain a processing result; according to the processing result, obtaining the style trigger word indicating the visual data to be generated and the corresponding area content style; generating structured description information according to the style trigger word and the corresponding area content style. By obtaining the structured description information in this embodiment, it can make the multi-modal model for generating visual data more easily understand the description information, assist in overcoming the technical difficulty of integrating multiple styles in a visual data, and thus realize the multi-modal model to generate higher-quality multi-style integrated visual data.
[0079] Among them, the structured description information has a predetermined format to facilitate the understanding of the multimodal model, and the predetermined format includes the correspondence and positional order between the style trigger words and the image region content information, and in the case of multiple style trigger words, the positional order between the multiple style trigger words, and so on. The predetermined format can be set based on specific business requirements. For example, for the prompt information "There is a blue cat in the cartoon style on the grassland, and the mountains in the distance are in the ink painting style." input by the user, the obtained structured description information is:
[0080] 1. Cartoon style: Blue cat
[0081] Supplementary information: The blue cat is on the grassland
[0082] 2. Ink painting style: Mountains in the distance
[0083] Supplementary information: None
[0084] In the embodiments of the present application, the prompt information input by the user can be the text information input by the user, such as a natural language description; the prompt information can also be a combination of text information and image information, that is, the user can also transmit image information while inputting text information. When applied to the scenario of multimedia content input, the present application obtains the prompt information input by the user, so as to draw image data with multi-style fusion according to the user's intention, or adjust the style of the image data.
[0085] Specifically, in the prompt information input by the user, there are usually two aspects of information:
[0086] One type of information is called image region content information, which refers to the name of a certain region or a certain object or element in the image / video that the user hopes to adjust or generate, and can be the main body, foreground, background, accompanying object, and / or blank space in the image, such as various elements included in the image such as people, trees, sky, rivers, etc.
[0087] Another type of information is image style information, that is, the style desired for a certain region or object in the above image, such as a realistic style, a comic style, a flat painting style, a nostalgic style, etc. This information referring to the image style is called the style trigger word of the image. The style trigger word can include but is not limited to art style, color matching, light and shadow effect, texture details, etc.
[0088] It can be understood that the prompt information input by the user may contain the standardized style trigger words pre-set in the thesaurus; it may also not contain the style trigger words pre-set in the thesaurus, but use other words to refer to a certain image style. In this case, in the technical solution of the present application, the word will be used to perform matching or similarity calculation in the style trigger word library, so as to clarify the image style desired by the user.
[0089] It should be noted that this application does not intend to define different image styles. On the contrary, in practical applications, by adding or defining new image styles, the selection of image styles becomes more diverse and meets the needs of the business. Similarly, when objects that have never been defined before appear in image or video data, new definitions can be added for these objects. Moreover, using the multi-modal model training method disclosed in this application, the image processing system can generate richer objects / elements and render images with richer styles.
[0090] It can be understood that the prompt information is the content input by the user. Since the operation habits of different users are different, the forms of the content input by the user may be diverse. As a preferred implementation, after this application obtains the prompt information input by the user, it further processes this content (such as natural language processing, speech information extraction processing, image information extraction processing, etc.). Then, the above-mentioned image region content information and image style information required for the technical solution of this application are extracted from it. Moreover, the description information and prompt information described in this application can be natural languages such as Chinese or English.
[0091] For example, in a case, the user hopes to obtain an image in which different regions have different styles. In this case, from the prompt information input by the user, at least two pieces of information or names of image styles can be obtained, and each piece of information or name of an image style corresponds to a region in the image to be generated. For example, the information input by the user can be, "There is a blue cat in a cartoon style on the grassland, and the mountains in the distance are in the ink painting style."
[0092] In another case, the user hopes to change the style of a part of an existing image. In this case, the prompt words input by the user indicate the regions for which the image style is hoped to be adjusted, which can be one region or multiple regions. When hoping to change the style of one image region, the image style after adjustment for this region is indicated; when hoping to change the styles of multiple image regions, the image styles after adjustment for each region are indicated respectively. For example, if the user already has an image of "There is a blue cat in a cartoon style on the grassland, and the mountains in the distance are in the ink painting style.", the user may input "Change the mountains in the distance to the cartoon style" or similar information to achieve the purpose of modifying the style of a part of the objects / elements in an image.
[0093] In the embodiments of the present application, the process of performing natural language word segmentation on the text content input by the user and decomposing the input user text content into basic units that can be analyzed subsequently, for example, decomposing the user text content into words, part-of-speech tagging, syntactic analysis, etc. For example, when the user input is "There is a blue cat in cartoon style on the prairie in realistic style", identifying the above text, which includes at least words such as "realistic style", "prairie", "on", "a", "blue cat", "cartoon style", etc. "Prairie" and "blue cat" correspond to the content information of the corresponding regions in the image to be generated, "on" is the spatial position relationship between the elements, and "realistic style" and "cartoon style" are the style information used to refer to the style of the object in the user input information. When the preset image style trigger word library of the present application contains "realistic" and "cartoon", the style trigger words are matched in the user input information.
[0094] In order to generate the visual data to be generated with corresponding styles and elements, in one embodiment, according to the description information, a plurality of style trigger words and the image region content information corresponding to each style trigger word are obtained.
[0095] In practical applications, after performing word segmentation on the user text content, at least the first image region content information, the second image region content information, the first style trigger word, and the second style trigger word are obtained. Among them, the first image region content information and the second image region content information are used to refer to any two combinations of the main body, foreground, background, accompanying body, and / or blank space in the image; the first image style information indicates the image style of the first image region, and the second image style information indicates the image style of the second image region.
[0096] In the embodiments of the present application, generating the structured description information according to the style trigger word and the corresponding region content information includes: based on the processing result, matching the corresponding style trigger word from the preset word library; determining the image region content information corresponding to each of the matched style trigger words; and generating the structured description information according to the matched style trigger word and the corresponding region content information.
[0097] In the embodiments of the present application, the preset word library can be a predefined vocabulary library, which includes natural language words or descriptions related to the style trigger words, so that the system can quickly and accurately identify the key information in the user text content. By calculating the similarity, the style trigger word that best matches the user input text content can be screened out, so as to ensure that the generated image can meet the user's expectations as much as possible.
[0098] In the embodiments of the present application, the similarity between the user text content and the style trigger words in the preset vocabulary can be evaluated by calculating the similarity. Multiple methods can be used for similarity calculation, such as cosine similarity, Jaccard similarity, edit distance, etc. By setting a threshold, only when the calculated similarity is higher than the threshold, the corresponding style trigger word will be recognized as a successful match.
[0099] In the application, the similarity between the processing result and the style trigger words in the style trigger word set can be calculated; the calculated similarity values are sorted, and the style trigger word used to indicate a part of the area in the initial image is determined based on the sorting result of the style trigger word similarity values.
[0100] In the embodiments of the present application, when calculating the similarity of the style trigger words, if there are multiple words in the word library whose similarity with the user input vocabulary is higher than the preset threshold, the multiple words in the word library higher than the threshold can be sorted in descending order of similarity, and the candidate image area content information or candidate style trigger word with the highest similarity value is used as the image area content information or style trigger word of the image to be generated.
[0101] For example, the processing result includes "a", "animation style", "white", "cat". The preset image area content information set is "cat", "dog", and the preset style trigger word set is "realistic style", "cartoon style", "abstract style". Calculate the similarity between each word segmentation result and each name in the style trigger word set. The similarity value between "realistic style" in the current word segmentation result and "realistic style" in the preset style trigger word set is 1.0, while the similarity values between "realistic style" in the current word segmentation result and other names in the style trigger word set are all less than the preset threshold. It can be determined that the image style corresponding to "realistic style" in the preset style trigger word set is the "realistic style" corresponding to the user text content.
[0102] Calculate the similarity between each word segmentation result and each style trigger word in the style trigger word set. The similarity value between "animation style" in the current word segmentation result and "cartoon style" in the style trigger word set is 0.9, and the similarity values with "realistic style" and "abstract style" are 0.7, both higher than the preset threshold. According to the sorting of the candidate style trigger words, "cartoon style" is determined as the style trigger word of the current image to be generated.
[0103] S101. Based on the description information and a pre-trained multi-modal model, obtain output visual data. There are differences in the image styles of at least two regions in the output visual data.
[0104] Multimodal models can generate images based on text information, generate videos based on text information, and in addition, can also generate images based on the input text information and image information, and can generate videos based on the input text information and image information, and so on. In the embodiments of the present application, the multimodal model can be based on a diffusion model as the basic model and trained with a pre-training dataset. The pre-trained multimodal model can generate output visual data with multi-style fusion according to multiple style trigger words and the image region content information corresponding to each style trigger word. Among them, the pre-training dataset is a dataset containing a large number of images and corresponding text descriptions. During the training process, the reconstruction loss is used to ensure the similarity between the generated image and the real image at the pixel level, and each image in the pre-training dataset corresponds to text description information to ensure that the text-to-image model can accurately understand the intention in the text description and generate an image that conforms to the text description.
[0105] That the image styles of at least two regions in the output visual data are different means that one output visual data output by the pre-trained multimodal model includes two or more image styles, and each image style and the region of each image style correspond to each style trigger word and the image region content information corresponding to the style trigger word.
[0106] In the embodiments of the present application, by obtaining the description information of the visual data to be generated; the description information corresponds to style trigger words and the image region content information corresponding to the style trigger words; based on the description information and the pre-trained multimodal model, output visual data with multi-style fusion is obtained, that is, the image styles of at least two regions in the output visual data are different, which not only effectively improves the accuracy and artistry of image generation, but also adopts an end-to-end method to directly output the multi-style fusion visual data expected by the user based on the description information through the pre-trained multimodal model, meeting the user's demand for generating images with multi-style fusion and also reducing the threshold for users to use AI technology.
[0107] In practical applications, as Figure 2 shown, obtain the prompt information of the image data to be generated input by the user. The prompt information is "There is a cartoon-style kitten on a realistic grassland, and there is an oil painting-style small stream in front of the kitten running through the image from left to right, and there is a cartoon small goldfish in the stream".
[0108] Perform natural language processing on the prompt information. The processing results include the correspondence between the image region content information and the style trigger words. That is, "a grassland in a realistic style" in the prompt information can be used as a group of correspondences between the image region content information and the style trigger words. The remaining word segmentation processing results are "a kitten in a cartoon style", "a brook in an oil painting style", and "a small goldfish in a cartoon style". Each group of correspondences between the image region content information and the style trigger words corresponds to at least one image region in the image to be generated.
[0109] Then, calculate the similarity between the obtained processing results and the preset set of image region content information and the similarity with the preset set of style trigger words. Among them, the words with the highest similarity values to the style trigger words in the set of style trigger words for "realistic style", "cartoon style", "oil painting style", and "cartoon" can be "realistic style", "cartoon style", "oil painting style". For "cartoon" in "a small goldfish in a cartoon style", according to the similarity calculation, the corresponding one is "cartoon style".
[0110] In addition, it is recognized that "brook" is a word that is not easily understood by the multi-modal model, and this word is replaced with "river".
[0111] The style trigger words for the image to be generated in the word segmentation processing results are "realistic style", "cartoon style", "oil painting style", "cartoon style", and the corresponding image region content information are "grassland", "kitten", "river", "small goldfish".
[0112] There is a kitten in a cartoon style on a grassland in a realistic style. In front of the kitten, there is a brook in an oil painting style running through the left and right of the image, and there is a small goldfish in a cartoon style in the brook.
[0113] Based on the above natural language processing, structured description information is obtained:
[0114] 1. Realistic style: Grassland
[0115] Supplementary information: There is a kitten on the grassland
[0116] 2. Cartoon style: Kitten, small goldfish
[0117] Supplementary information: There is a river in front of the kitten, and there is a small goldfish in the river
[0118] 3. Oil painting style: River
[0119] Supplementary information: The river runs through the left and right of the image
[0120] Finally, based on the structured description information, output visual data is obtained.
[0121] Based on the above embodiments, description information of the visual data to be generated is obtained, and based on the description information and a pre-trained multimodal model, output visual data is obtained to achieve the generation of images with a combination of multiple styles. The training method of the multimodal model is as follows Figure 3 shown, including:
[0122] S300. Obtain an initial model for image generation.
[0123] S301. Obtain an image data set, including image data of multiple styles.
[0124] S302. Perform structured annotation on the image data set, including annotating the image region content information of the image and the style trigger words of the image style.
[0125] S303. Train the initial model based on the structured annotation data to obtain a trained multimodal model.
[0126] In the embodiments of the present application, the image data in the image data set can be the historical generated images of the multimodal model, or image resources from various sources such as art works and photographic photos. Structured annotation is performed on the image data. The annotation of the image style can be performed manually or automatically through a deep learning model. For example, a pre-trained image style recognition model can be used to classify the image data by style, such as realistic, cartoon, abstract, oil painting, etc. The annotated image data will contain the image region content information and the corresponding style trigger words.
[0127] In the embodiments of the present application, the structured annotation can be to attach annotation information with a specific hierarchy or association relationship to the image data. First, all the image region content information included in each image data can be determined, and the image region content information can be classified by element category. For example, the element main body and the background included in the image data are identified. Among them, the element main body can be annotated in the way of character role, animal species, still life category, and main body attribute, and the background can include but is not limited to scene light, lens, composition, etc., which are respectively annotated.
[0128] For example, first, describe the element main body in the following order: character role, gender, body shape, facial features, expression, and limb movement. Then, describe the background, including the characteristics of the light (such as soft or bright), lens selection (such as close-up or telephoto), shooting angle (such as eye level or upward shooting), and the overall composition of the background. In addition, a style trigger word needs to be added before the description of each element main body and the background to mark the specific attributes of the style, so as to be used as a style guide during subsequent image generation. It should be emphasized that the style trigger word should avoid being the same as the common words in natural language to ensure its uniqueness.
[0129] It can be understood that the image data in the image dataset only includes one image region, and the image data only contains a single element in a single style. The image region in the image data is structurally annotated to obtain the content information of the image region and the style trigger word corresponding to the image region.
[0130] In addition, the image dataset also includes more than two image regions. Structurally annotating the image dataset includes:
[0131] Performing image segmentation on the image data to obtain the at least two image regions.
[0132] Annotating each of the at least two image regions obtained by image segmentation, including the image region content information of the image region and the style trigger word of the image style.
[0133] It can be understood that during the annotation process, the style can be defined according to the color, texture, shape, and artistic style presented by each image data, and the style can also be defined through the custom visual presentation features of the image data. Each image data is annotated as a specific style and trained and stored in the model in the form of a style trigger word set, so as to provide sufficient learning information to the text-to-image model during the training process.
[0134] In the embodiment of the present application, after annotating the image style of the image data, the method further includes:
[0135] Generating a set of style trigger words for the same image style, and the set of style trigger words includes synonyms or near-synonyms.
[0136] In the embodiment of the present application, for the convenience of subsequent model application, the set of style trigger words includes a set of synonyms or near-synonyms. For example, the image to be trained is "a cat in a cartoon style". Identifying the image region content information included in the image to be trained is "cat", and the style trigger word is "cartoon style". Therefore, during the annotation of this image, synonyms or near-synonyms other than "cartoon style" can be used as the style trigger words of the current image to be trained, such as "comic style", "animation style", "fairy tale style", etc., so as to provide diverse image style expressions for the model.
[0137] In another embodiment, it includes:
[0138] Obtain a prompt message input by the user, where the prompt message includes text information and initial visual data; perform natural language processing on the text information in the prompt message to obtain a processing result; according to the processing result, obtain a style trigger word indicating the visual data to be generated and corresponding region content information; generate structured description information according to the style trigger word and the corresponding region content information; the description information includes the corresponding relationship between the image region content information indicating a part of the region in the initial visual data and the style trigger word.
[0139] Based on the description information and a pre-trained multi-modal model, obtain output visual data; compared with the initial visual data, in the image region indicated by the image region content information, the image style of the output visual data changes according to the style trigger word.
[0140] In the embodiments of the present application, the initial visual data may be image data and video data previously generated by the model, or may be new visual data input by the user side simultaneously with the prompt message.
[0141] In the embodiments of the present application, the text content included in the prompt message may be a change or addition to the style trigger word or the image region content information in the initial visual data. For example, the initial visual data input by the user is a waterfall landscape picture taken by the user's camera, and the prompt message is "add a cartoon-style boy to the picture". Since the content in the initial visual data is not changed in the prompt message, the initial visual data can be used as the background picture, and natural language processing is performed on the description information to obtain that the style trigger word in the prompt message is "cartoon style", and the image region content information is "boy", which is added to the initial visual data as the foreground region of the image.
[0142] In another embodiment, the initial visual data is a waterfall landscape picture taken by the user's camera, and the prompt message may be to "paint the sky in the initial visual data in a cartoon style". The prompt message changes the style trigger word of a part of the initial visual data. First, identify the initial visual data, determine the image region corresponding to the "sky" in the initial image through an image segmentation algorithm and perform segmentation, and add the image that meets the image region content information of "sky" and the style trigger word of "cartoon style" to the initial visual data to generate output visual data.
[0143] In the embodiments of the present application, in practical applications, such as Figure 4 shown, the initial image may be an image input by the user, or may be an initial image previously generated by the multi-modal model according to relevant text. Assume that when the initial image is an image previously generated by the user, the prompt message input by the user simultaneously with the initial image is "a realistic-style forest, draw a fairy-tale-style deer and a comic-style castle in the forest".
[0144] Perform natural language processing on the prompt information to obtain image region content information indicating some regions in the initial image and style trigger words. For example, some regions may refer to the regions to be changed determined according to the text content of the prompt information, that is, the regions to be changed refer to the image regions of "deer" and "castle" in the initial image. The obtained processing result includes the correspondence between the image region content information and the style trigger words, that is, the "deer in fairy tale style" and "castle in comic style" in the prompt information are respectively used as a group of correspondences between the image region content information and the style trigger words.
[0145] Determine the regions to be changed in the initial image according to the image region content information in the prompt information, extract the style trigger words corresponding to the image region content information in each region to be changed, and calculate the similarity between the obtained processing result and the preset style trigger word set. Among them, calculate the words with the highest similarity between "fairy tale style" and "comic style" and the style trigger words in the style trigger word set. The style trigger words with the highest similarity value may be "fairy tale style" and "comic style".
[0146] It can be understood that in this embodiment, the content of the prompt information does not include words similar to image region change. Therefore, the region corresponding to the image region content information in the initial image can be used as the region to be changed corresponding to the current style trigger word.
[0147] Finally, as Figure 5 shown, send the output image in the pre-trained text-to-image model to the user side. Compared with the initial image, in the image regions indicated by "deer" and "castle", the image style of "deer" is changed to "fairy tale style", and the image style of "castle" is changed to "comic style".
[0148] As Figure 6 shown, this application provides a visual data generation device, including:
[0149] An acquisition module 61, configured to obtain a description information of the visual data to be generated; wherein, the description information corresponds to style trigger words and image region content information corresponding to the style trigger words;
[0150] A processing module 62, configured to obtain output visual data based on the description information and a pre-trained multimodal model; there are differences in the image styles of at least two regions in the output visual data.
[0151] In the embodiment of this application, the acquisition module 61 is used for:
[0152] Obtain the prompt information input by the user; wherein, the prompt information includes at least text information;
[0153] Perform natural language processing on the text information in the prompt message to obtain a processing result;
[0154] According to the processing result, obtain a style trigger word indicating the style of the visual data to be generated and the corresponding area content information;
[0155] Generate structured description information according to the style trigger word and the corresponding area content information.
[0156] In the embodiment of the present application, the obtaining module 61 is further configured to:
[0157] Based on the processing result, match the corresponding style trigger word from a preset word library;
[0158] Determine the image area content information corresponding to each of the matched style trigger words;
[0159] Generate structured description information according to the matched style trigger word and the corresponding area content information.
[0160] Optionally, the obtaining module 61 is further configured to:
[0161] Calculate the similarity between the processing result and a preset word library, and obtain a style trigger word in the preset word library that meets the preset similarity rule.
[0162] The above-mentioned training method for a multi-modal model includes:
[0163] Obtain an initial model for image generation;
[0164] Obtain an image data set, including image data of multiple styles;
[0165] Perform structured annotation on the image data set, including annotating the image area content information of the image and the style trigger word of the image style;
[0166] Train the initial model based on the structured annotation data to obtain a trained multi-modal model.
[0167] Optionally, the image data only includes one image area, and the performing structured annotation on the image data set includes:
[0168] Annotate the image data to obtain the content information of the image area in the image data and the style trigger word corresponding to the image area.
[0169] Optionally, the image data includes two or more image areas, and the performing structured annotation on the image data set includes:
[0170] Perform image segmentation on the image data to obtain the at least two image areas;
[0171] Perform annotation on at least two image regions obtained by image segmentation, including the image region content information of the image region and the style trigger words of the image style.
[0172] Optionally, after annotating the image style for the image data, the method further includes: for the same image style, generating a set of style trigger words, where the set of style trigger words includes synonyms or near-synonyms.
[0173] See Figure 7 , Figure 7 FIG. is a schematic structural diagram of an electronic device shown in an embodiment of the present application. The electronic device 700 includes a memory 710 and a processor 720.
[0174] The present application can be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium), on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or an electronic device, a server, etc.), the processor is caused to execute some or all of the steps of the above method according to the present application.
[0175] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the applications herein can be implemented as electronic hardware, computer software, or a combination of both.
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems and methods according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0177] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating visual data, characterized in that, Including: Obtaining description information of visual data to be generated; wherein, the description information is structured description information generated based on natural language processing, and the description information includes style trigger words and image region content information corresponding to the style trigger words; Based on the description information and a pre-trained multimodal model, respectively applying corresponding style trigger words to different regions indicated by the image region content information to obtain output visual data; There are differences in the image styles of at least two regions in the output visual data; The training method of the multimodal model includes: Obtaining an initial model for image generation; Obtaining an image data set, including image data of multiple styles; Performing structured annotation on the image data set, including annotating the image region content information of the image and the style trigger words of the image style; Training the initial model based on the structured annotation data to obtain a trained multimodal model.
2. The method according to claim 1, wherein The obtaining of the description information of the visual data to be generated includes: Obtaining prompt information input by a user; wherein, the prompt information includes at least text information; Performing natural language processing on the text information in the prompt information to obtain a processing result; According to the processing result, obtaining style trigger words indicating the visual data to be generated and corresponding region content information; Generating structured description information according to the style trigger words and the corresponding region content information.
3. The method according to claim 2, wherein The generating of the structured description information according to the style trigger words and the corresponding region content information includes: Based on the processing result, matching corresponding style trigger words from a preset word library; Determining the image region content information corresponding to each of the matched style trigger words; Generating structured description information according to the matched style trigger words and the corresponding region content information.
4. The method according to claim 3, characterized in that, The matching of the corresponding style trigger words from the preset word library includes: Calculating the similarity between the processing result and the preset word library to obtain style trigger words in the preset word library that meet the preset similarity rules.
5. The method according to claim 1, characterized in that The image data only includes one image region, and the performing of structured annotation on the image data set includes: Annotating the image data to obtain the content information of the image region in the image data and the style trigger words corresponding to the image region.
6. The method according to claim 1, wherein The image data includes two or more image regions, and the performing of structured annotation on the image data set includes: Performing image segmentation on the image data to obtain the at least two image regions; Respectively annotating the at least two image regions obtained by image segmentation, including the image region content information of the image region and the style trigger words of the image style.
7. The method according to claim 1, wherein After annotating the image style of the image data, the method further includes: for the same image style, generating a set of style trigger words, and the set of style trigger words includes synonyms or near-synonyms.
8. A visual data generation device, characterized in that, Including: An obtaining module, configured to obtain description information of visual data to be generated; wherein, the description information is structured description information generated based on natural language processing, and the description information includes style trigger words and image region content information corresponding to the style trigger words; A processing module, configured to apply corresponding style trigger words to different regions indicated by the image region content information respectively based on the description information and a pre-trained multimodal model, so as to obtain output visual data; there are differences in the image styles of at least two regions in the output visual data; the training method of the multimodal model includes: Obtaining an initial model for image generation; Obtaining an image data set, including image data of multiple styles; Performing structured annotation on the image data set, including annotating the image region content information of the image and the style trigger words of the image style; Training the initial model based on the structured annotation data to obtain a trained multimodal model.
9. A computer-readable storage medium, characterized in that, It stores executable code, which when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Image generation method and device of multi-modal model based on double-contrast learning
CN116612205A
Image redrawing model training method, image redrawing method and device
CN116664719A
Image generation method and device, computer equipment and storage medium
CN117078790A
Image generation method and device and storage medium
CN117475031A